(2) * Joseph Emerson Raja
(3) Md. Jakir Hossein
*corresponding author
AbstractAutomatic Video Captions are increasingly important to use in applications in the accessibility, search, education, and content moderation areas. Generating accurate and context-aware captions is difficult for lengthy and complex videos. Recent surveys focus more on language models than on the vision side. This means that the contribution of visual feature extraction towards the performance of the model is not studied properly. Here we attempt to fill this gap by conducting reviews of video captioning literature from the vision viewpoint, comparing CNNs-based and transformer-based feature extraction, and detecting trends, pros and cons. Approximately 100 papers published between 2017 - 2025 were reviewed and analyzed based on base vision, dataset and evaluation metrics. The review classified the selected papers into CNN-based pipelines and transformer-based techniques. Moreover, it compared their rank measured by application popularity, dataset usage, and performance scores across popular datasets/benchmarks. The review shows that CNNs were dominated before 2021, but transformer-based approaches now lead due to their ability to capture long-range temporal dependencies and multimodal interactions. However, CNNs continue to perform well for short videos. As for Dataset usage, patterns show that ActivityNet Captions and YouCook2 are suited for long-form reasoning because of timely manner captioning, while MSR-VTT and MSVD are effective for short clips (one sentence caption). In terms of Special semantic quality metrics, CIDEr and METEOR are more semantically appropriate compared to n-gram quality metrics like BLEU. This review emphasizes transformer-based architecture, particularly multimodal models, which will be the future of video captioning. However, challenges remain in handling lengthy videos, adapting across domains and languages, and ensuring efficient deployment (small, optimized model). Addressing these gaps will add value by making video captioning systems more reliable, interpretable, and broadly usable in real-world contexts.
KeywordsVideo Captioning, Video feature extraction, Vision-Language Models
|
DOIhttps://doi.org/10.26555/ijain.v12i3.2280 |
Article metricsAbstract views : 38 |
Cite |

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
___________________________________________________________
International Journal of Advances in Intelligent Informatics
ISSN 2442-6571 (print) | 2548-3161 (online)
Organized by UAD and ASCEE Computer Society
Published by Universitas Ahmad Dahlan
W: http://ijain.org
E: info@ijain.org (paper handling issues)
andri.pranolo.id@ieee.org (publication issues)
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0
























