Journal Article

·2025

Automatic Turkish Image Captioning Using Non-Native Deep Caption Generator Models and Neural Machine Translators

Serdar Yıldız YTU , Abbas Memiş YTU

Abstract

This paper presents a study of Turkish image captioning that leverages a combination of leading non-native deep caption generator models and neural machine translators. Essentially, a non-native image captioning model is employed alongside a language translation module to generate Turkish caption texts for the input query images. First, we generate English captions for the input images using a method from a set of advanced deep models, including CLIPCap, BLIP, BLIP2, FUSE-CAP, OFA, PromptCap, Kosmos2, MiniGPT4, LlaVA, BakLlaVA and GIT. The output image captions are then translated into Turkish using NLLB and OPUS-MT deep language translation models. Extensive experimental analysis was performed on a recent benchmark dataset proposed for general-purpose image captioning tasks, Tiny TR-CAP, and the observed results are reported and discussed. In the performance evaluation tests, best image captioning success rates of 0.4124 BLEU-1, 0.2221 BLEU-2, 0.1155 BLEU-3, 0.0591 BLEU-4, 0.1176 METEOR, 0.2993 ROUGE-L, 0.1446 CIDEr and 0.0168 SPICE were achieved for the OFA caption generation model with OPUS-MT.

Keywords

Closed captioning Generator (circuit theory) Image (mathematics) Machine translation Turkish Benchmark (surveying) Set (abstract data type) Deep learning Computer science Artificial intelligence Natural language processing

Subject Areas

Multimodal Machine Learning Applications ·Computer Vision and Pattern Recognition ·Physical Sciences
Generative Adversarial Networks and Image Synthesis ·Computer Vision and Pattern Recognition ·Physical Sciences
Topic Modeling ·Artificial Intelligence ·Physical Sciences

Related YTU Activities

Sustainability activities and projects connected to this article.