Abstract
This paper presents a study of Turkish image captioning that leverages a combination of leading non-native deep caption generator models and neural machine translators. Essentially, a non-native image captioning model is employed alongside a language translation module to generate Turkish caption texts for the input query images. First, we generate English captions for the input images using a method from a set of advanced deep models, including CLIPCap, BLIP, BLIP2, FUSE-CAP, OFA, PromptCap, Kosmos2, MiniGPT4, LlaVA, BakLlaVA and GIT. The output image captions are then translated into Turkish using NLLB and OPUS-MT deep language translation models. Extensive experimental analysis was performed on a recent benchmark dataset proposed for general-purpose image captioning tasks, Tiny TR-CAP, and the observed results are reported and discussed. In the performance evaluation tests, best image captioning success rates of 0.4124 BLEU-1, 0.2221 BLEU-2, 0.1155 BLEU-3, 0.0591 BLEU-4, 0.1176 METEOR, 0.2993 ROUGE-L, 0.1446 CIDEr and 0.0168 SPICE were achieved for the OFA caption generation model with OPUS-MT.
Keywords
Subject Areas
Related YTU Activities
Sustainability activities and projects connected to this article.