Fetching the paper…
Reading the bibliography…
Despite the significant progress of fully-supervised video captioning, zero-shot methods remain much less explored.
D. L. Chen, W. B. Dolan, Collecting highly parallel data for paraphrase evaluation, in: ACL, 2011, pp. 190–200
2011
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Xu, T. Mei, T. Yao, Y. Rui, MSR-VTT: A large video description dataset for bridging video and language, in: CVPR, 2016, pp. 5288–5296
2016
Earlier work this paper cites.
K. Lee, M. Chang, K. Toutanova, Latent retrieval for weakly supervised open domain question answering, in: A. Korhonen, D. R. Traum, L. Màrquez (Eds.), ACL, Association for Computational Linguistics, 2019, pp. 6086–6096
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al., Language models are unsupervised multitask learners, OpenAI blog 1 (8) (2019) 9
2019
Earlier work this paper cites.
X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, W. Y. Wang, Vatex: A large-scale, high-quality multilingual dataset for video-and-language research, in: CVPR, 2019, pp. 4581–4591
2019
Earlier work this paper cites.
S. Zhang, Z. Tan, Z. Zhao, J. Yu, K. Kuang, T. Jiang, J. Zhou, H. Yang, F. Wu, Comprehensive information integration modeling framework for video titling, in: KDD, ACM, 2020, pp. 2744–2754
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
V. Karpukhin, B. Oguz, S. Min, P. S. H. Lewis, L. Wu, S. Edunov, D. Chen, W. Yih, Dense passage retrieval for open-domain question answering, in: EMNLP, 2020, pp. 6769–6781
2020
Earlier work this paper cites.
B. Y. Lin, W. Zhou, M. Shen, P. Zhou, C. Bhagavatula, Y. Choi, X. Ren, Commongen: A constrained text generation challenge for generative commonsense reasoning, in: T. Cohn, Y. He, Y. Liu (Eds.), EMNLP, Vol. EMNLP 2020 of Findings of ACL, Association for Computational Linguistics, 2020, pp. 1823–1840
2020
Earlier work this paper cites.
H. Chen, J. Li, X. Hu, Delving deeper into the decoder for video captioning, in: G. D. Giacomo, A. Catalá, B. Dilkina, M. Milano, S. Barro, A. Bugarín, J. Lang (Eds.), ECAI, Vol. 325 of Frontiers in Artificial Intelligence and Applications, IOS Press, 2020, pp. 1079–1086
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
Z. Zhang, Z. Qi, C. Yuan, Y. Shan, B. Li, Y. Deng, W. Hu, Open-book video captioning with retrieve-copy-generate network, in: CVPR, Computer Vision Foundation / IEEE, 2021, pp. 9837–9846
2021
Earlier work this paper cites.
Y. Tu, C. Zhou, J. Guo, S. Gao, Z. Yu, Enhancing the alignment between target words and corresponding frames for video captioning, Pattern Recognition 111 (2021) 107702
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, I. Sutskever, Learning transferable visual models from natural language supervision, in: ICML, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
J. Perez-Martin, B. Bustos, J. Pérez, Improving video captioning with temporal composition of a visual-syntactic embedding * {}^{\mbox{*}} , in: WACV, IEEE, 2021, pp. 3038–3048
2021
Earlier work this paper cites.
L. Nie, L. Qu, D. Meng, M. Zhang, Q. Tian, A. D. Bimbo, Search-oriented micro-video captioning, in: ACM MM, 2022, pp. 3234–3243
2022
Earlier work this paper cites.
Y. Lu, M. Bartolo, A. Moore, S. Riedel, P. Stenetorp, Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity, in: S. Muresan, P. Nakov, A. Villavicencio (Eds.), ACL, Association for Computational Linguistics, 2022, pp. 8086–8098
2022
Cited alongside, same era.
A. Webson, E. Pavlick, Do prompt-based models really understand the meaning of their prompts?, in: M. Carpuat, M. de Marneffe, I. V. M. Ruíz (Eds.), NAACL, Association for Computational Linguistics, 2022, pp. 2300–2344
2022
Cited alongside, same era.
M. Shu, W. Nie, D. Huang, Z. Yu, T. Goldstein, A. Anandkumar, C. Xiao, Test-time prompt tuning for zero-shot generalization in vision-language models, in: S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, A. Oh (Eds.), NeurIPS, 2022
2022
Cited alongside, same era.
Y. Tewel, Y. Shalev, I. Schwartz, L. Wolf, Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic, in: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, IEEE, 2022, pp. 17897–17907
2023
Later among the works it cites.
H. Zhang, X. Li, L. Bing, Video-llama: An instruction-tuned audio-visual language model for video understanding, in: Y. Feng, E. Lefever (Eds.), EMNLP, Association for Computational Linguistics, 2023, pp. 543–553
2023
Later among the works it cites.
Y. Shi, H. Xu, C. Yuan, B. Li, W. Hu, Z. Zha, Learning video-text aligned representations for video captioning, TOMM 19 (2) (2023) 63:1–63:21
2023
Later among the works it cites.
X. Luo, X. Luo, D. Wang, J. Liu, B. Wan, L. Zhao, Global semantic enhancement network for video captioning, Pattern Recognition 145 (2024) 109906
2024
Closest in time.
Z. Zeng, Y. Xie, H. Zhang, C. Chen, B. Chen, Z. Wang, Meacap: Memory-augmented zero-shot image captioning, in: CVPR, IEEE, 2024, pp. 14100–14110
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
H. Luo, L. Ji, M. Zhong, Y. Chen, W. Lei, N. Duan, T. Li, Clip4clip: An empirical study of CLIP for end to end video clip retrieval and captioning, Neurocomputing 508 (2022) 293–304
2022
Cited alongside, same era.
B. Ni, H. Peng, M. Chen, S. Zhang, G. Meng, J. Fu, S. Xiang, H. Ling, Expanding language-image pretrained models for general video recognition, in: S. Avidan, G. J. Brostow, M. Cissé, G. M. Farinella, T. Hassner (Eds.), ECCV, Vol. 13664 of Lecture Notes in Computer Science, Springer, 2022, pp. 1–18
2022
Cited alongside, same era.
H. Ye, G. Li, Y. Qi, S. Wang, Q. Huang, M. Yang, Hierarchical modular network for video captioning, in: CVPR, 2022, pp. 17918–17927
2022
Cited alongside, same era.
T.-Z. Niu, S.-S. Dong, Z.-D. Chen, X. Luo, Z. Huang, S. Guo, X.-S. Xu, A multi-layer memory sharing network for video captioning, Pattern Recognition 136 (2023) 109202
2023
Cited alongside, same era.
B. Yang, F. Liu, X. Wu, Y. Wang, X. Sun, Y. Zou, Multicapclip: Auto-encoding prompts for zero-shot multilingual visual captioning, in: A. Rogers, J. L. Boyd-Graber, N. Okazaki (Eds.), ACL, Association for Computational Linguistics, 2023, pp. 11908–11922
2023
Cited alongside, same era.
W. Li, L. Zhu, L. Wen, Y. Yang, Decap: Decoding CLIP latents for zero-shot captioning via text-only training, in: ICLR, OpenReview.net, 2023
2023
Cited alongside, same era.
2024
Closest in time.
M. Patidar, R. Sawhney, A. K. Singh, B. Chatterjee, Mausam, I. Bhattacharya, Few-shot transfer learning for knowledge base question answering: Fusing supervised models with in-context learning, in: L. Ku, A. Martins, V. Srikumar (Eds.), ACL, Association for Computational Linguistics, 2024, pp. 9147–9165
2024
Closest in time.
X. Jin, B. Zhang, W. Gong, K. Xu, X. Deng, P. Wang, Z. Zhang, X. Shen, J. Feng, Mv-adapter: Multimodal video transfer learning for video text retrieval, in: CVPR, IEEE, 2024, pp. 27134–27143
2024
Closest in time.
X. Li, J. Li, Aoe: Angle-optimized embeddings for semantic textual similarity, in: L. Ku, A. Martins, V. Srikumar (Eds.), ACL, Association for Computational Linguistics, 2024, pp. 1825–1839
2024
Closest in time.
F. Yuan, S. Gu, X. Zhang, Z. Fang, Fully exploring object relation interaction and hidden state attention for video captioning, Pattern Recognition 159 (2025) 111138
2025
Closest in time.
D. Zeng, Y. Shen, M. Lin, Z. Yi, J. Ouyang, Zero-shot image captioning with multi-type entity representations, in: T. Walsh, J. Shah, Z. Kolter (Eds.), AAAI, AAAI Press, 2025, pp. 22308–22316
2025
Closest in time.
Z. Xi, G. Shi, H. Sun, B. Zhang, S. Li, L. Wu, EIKA: explicit & implicit knowledge-augmented network for entity-aware sports video captioning, Expert Syst. Appl. 274 (2025) 126906
2025
Closest in time.
Z. Xi, G. Shi, X. Li, J. Yan, Z. Li, L. Wu, Z. Liu, L. Wang, A simple yet effective knowledge guided method for entity-aware video captioning on a basketball benchmark, Neurocomputing 619 (2025) 129177
2025
Closest in time.
H. Jiang, J. Zhang, R. Huang, C. Ge, Z. Ni, S. Song, G. Huang, Cross-modal adapter for vision–language retrieval, Pattern Recognition 159 (2025) 111144
2025
Closest in time.
P. Li, T. Wang, X. Zhao, X. Xu, M. Song, Pseudo-labeling with keyword refining for few-supervised video captioning, Pattern Recognition 159 (2025) 111176
2025
Closest in time.
Z. Dai, K. Cheng, F. Shao, Z. Dong, S. Zhu, Text–video retrieval re-ranking via multi-grained cross attention and frozen image encoders, Pattern Recognition 159 (2025) 111099
2025
Closest in time.
W. Xu, Y. Xu, Z. Miao, Y. Cen, L. Wan, X. Ma, Crocaps: A clip-assisted cross-domain video captioner, Expert Syst. Appl. 268 (2025) 126296
2025
Closest in time.