A comprehensive review of CNN and transformer based visual feature extraction for automatic video captioning

(1) Hemel Sharker Akash Mail (Graduate Research Assistant, Faculty of Engineering & Technology (FET) , Multimedia University, Melaka, Malaysia)
(2) * Joseph Emerson Raja Mail (Assistant Professor, FACULTY OF ENGINEERING AND TECHNOLOGY (FET), Multimedia University, Malaysia)
(3) Md. Jakir Hossein Mail (Associate Professor, FACULTY OF ENGINEERING AND TECHNOLOGY (FET), Multimedia University, Malaysia)
*corresponding author

Abstract


Automatic Video Captions are increasingly important to use in applications in the accessibility, search, education, and content moderation areas. Generating accurate and context-aware captions is difficult for lengthy and complex videos. Recent surveys focus more on language models than on the vision side. This means that the contribution of visual feature extraction towards the performance of the model is not studied properly. Here we attempt to fill this gap by conducting reviews of video captioning literature from the vision viewpoint, comparing CNNs-based and transformer-based feature extraction, and detecting trends, pros and cons. Approximately 100 papers published between 2017 - 2025 were reviewed and analyzed based on base vision, dataset and evaluation metrics. The review classified the selected papers into CNN-based pipelines and transformer-based techniques. Moreover, it compared their rank measured by application popularity, dataset usage, and performance scores across popular datasets/benchmarks. The review shows that CNNs were dominated before 2021, but transformer-based approaches now lead due to their ability to capture long-range temporal dependencies and multimodal interactions. However, CNNs continue to perform well for short videos. As for Dataset usage, patterns show that ActivityNet Captions and YouCook2 are suited for long-form reasoning because of timely manner captioning, while MSR-VTT and MSVD are effective for short clips (one sentence caption). In terms of Special semantic quality metrics, CIDEr and METEOR are more semantically appropriate compared to n-gram quality metrics like BLEU. This review emphasizes transformer-based architecture, particularly multimodal models, which will be the future of video captioning. However, challenges remain in handling lengthy videos, adapting across domains and languages, and ensuring efficient deployment (small, optimized model). Addressing these gaps will add value by making video captioning systems more reliable, interpretable, and broadly usable in real-world contexts.

Keywords


Video Captioning, Video feature extraction, Vision-Language Models

   

DOI

https://doi.org/10.26555/ijain.v12i3.2280
      

Article metrics

Abstract views : 38

   

Cite

   


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

___________________________________________________________
International Journal of Advances in Intelligent Informatics
ISSN 2442-6571  (print) | 2548-3161 (online)
Organized by UAD and ASCEE Computer Society
Published by Universitas Ahmad Dahlan
W: http://ijain.org
E: info@ijain.org (paper handling issues)
 andri.pranolo.id@ieee.org (publication issues)

View IJAIN Stats

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0