Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have taken a great step towards AGI.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli · 2004
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Earlier work this paper cites.
Automatic understanding of image and video advertisements
Zaeem Hussain, Mingda Zhang, Xiaozhong Zhang, Keren Ye, Christopher Thomas, Zuha Agha, Nathan Ong, and Adriana Kovashka · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
TVQA: localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L. Berg · 2018
Earlier work this paper cites.
How2: a large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze · 2018
Earlier work this paper cites.
Open-ended long-form video question answering via adaptive hierarchical reinforced networks
Zhou Zhao, Zhu Zhang, Shuwen Xiao, Zhou Yu, Jun Yu, Deng Cai, Fei Wu, and Yueting Zhuang · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason Corso · 2018
Earlier work this paper cites.
Beyond rnns: Positional self-attention with co-attention for video question answering
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan · 2019
Earlier work this paper cites.
Interpreting the rhetoric of visual advertisements
Keren Ye, Narges Honarvar Nazari, James Hahn, Zaeem Hussain, Mingda Zhang, and Adriana Kovashka · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao · 2019
Earlier work this paper cites.
Reasoning with heterogeneous graph alignment for video question answering
Pin Jiang and Yahong Han · 2020
Earlier work this paper cites.
Intentonomy: a dataset and study towards human intent understanding
Menglin Jia, Zuxuan Wu, Austin Reiter, Claire Cardie, Serge Belongie, and Ser-Nam Lim · 2021
Earlier work this paper cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Earlier work this paper cites.
Look before you speak: Visually contextualized utterances
Paul Hongsuck Seo, Arsha Nagrani, and Cordelia Schmid · 2021
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Earlier work this paper cites.
Multimet: A multimodal dataset for metaphor understanding
Dongyu Zhang, Minghao Zhang, Heting Zhang, Liang Yang, and Hongfei Lin · 2021
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Metaclue: Towards comprehensive visual metaphors research
Arjun R Akula, Brendan Driscoll, Pradyumna Narayana, Soravit Changpinyo, Zhiwei Jia, Suyash Damle, Garima Pruthi, Sugato Basu, Leonidas Guibas, William T Freeman, et al · 2023
Earlier work this paper cites.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee · 2023
Earlier work this paper cites.
Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering
Difei Gao, Luowei Zhou, Lei Ji, Linchao Zhu, Yi Yang, and Mike Zheng Shou · 2023
Earlier work this paper cites.
Zhiwei Jia, Pradyumna Narayana, Arjun R Akula, Garima Pruthi, Hao Su, Sugato Basu, and Varun Jampani · 2023
Cited alongside, same era.
Persuasion strategies in advertisements
Yaman Kumar, Rajat Jha, Arunim Gupta, Milan Aggarwal, Aditya Garg, Tushar Malyan, Ayush Bhardwaj, Rajiv Ratn Shah, Balaji Krishnamurthy, and Changyou Chen · 2023
Cited alongside, same era.
Video-llava: Learning united visual representation by alignment before projection
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan · 2023
Cited alongside, same era.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2023
Cited alongside, same era.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik · 2023
Large language models as biomedical hypothesis generators: a comprehensive evaluation
Biqing Qi, Kaiyan Zhang, Kai Tian, Haoxiang Li, Zhang-Ren Chen, Sihang Zeng, Ermo Hua, Hu Jinfang, and Bowen Zhou · 2024
Later among the works it cites.
Seeing the unseen: Visual metaphor captioning for videos
Abisek Rajakumar Kalarani, Pushpak Bhattacharyya, and Sumit Shekhar · 2024
Later among the works it cites.
Unlocking video-llm via agent-of-thoughts distillation
Yudi Shi, Shangzhe Di, Qirui Chen, and Weidi Xie · 2024
Later among the works it cites.
Moviechat+: Question-aware sparse memory for long video question answering
Enxin Song, Wenhao Chai, Tian Ye, Jenq-Neng Hwang, Xi Li, and Gaoang Wang · 2024
Later among the works it cites.
Can i trust your answer? visually grounded video question answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models
Munan Ning, Bin Zhu, Yujia Xie, Bin Lin, Jiaxi Cui, Lu Yuan, Dongdong Chen, and Li Yuan · 2023
Cited alongside, same era.
An eye-tracking study on how the popularity and gender of the endorsers affected the audience’s attention on the advertisement
Majid Zahmati, Seyed Morteza Azimzadeh, Mohammad Saber Sotoodeh, and Omid Asgari · 2023
Cited alongside, same era.
Multicmet: A novel chinese benchmark for understanding multimodal metaphor
Dongyu Zhang, Jingwei Yu, Senyuan Jin, Liang Yang, and Hongfei Lin · 2023
Cited alongside, same era.
Temporalbench: Benchmarking fine-grained temporal understanding for multimodal video models
Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou, Kai Zhang, Feng Yao, Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang Shang, Yao Dou, Jaden Park, Jianfeng Gao, Yong Jae Lee, and Jianwei Yang · 2024
Cited alongside, same era.
Grounded multi-hop videoqa in long-form egocentric videos
Qirui Chen, Shangzhe Di, and Weidi Xie · 2024
Cited alongside, same era.
Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing · 2024
Cited alongside, same era.
Grounded question-answering in long egocentric videos
Shangzhe Di and Weidi Xie · 2024
Cited alongside, same era.
Later among the works it cites.
Funqa: Towards surprising video comprehension
Binzhu Xie, Sicheng Zhang, Zitang Zhou, Bo Li, Yuanhan Zhang, Jack Hessel, Jingkang Yang, and Ziwei Liu · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al · 2024
Later among the works it cites.
Evoagent: Towards automatic multi-agent generation via evolutionary algorithms
Siyu Yuan, Kaitao Song, Jiangjie Chen, Xu Tan, Dongsheng Li, and Deqing Yang · 2024
Later among the works it cites.
Enhancing baidu multimodal advertisement with chinese text-to-image generation via bilingual alignment and caption synthesis
Kang Zhao, Xinyu Zhao, Zhipeng Jin, Yi Yang, Wen Tao, Cong Han, Shuanglong Li, and Lin Liu · 2024
Later among the works it cites.
Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation
Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin, Marinka Zitnik, and Pan Zhou · 2024
Later among the works it cites.
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al · 2025
Closest in time.
Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al · 2025
Closest in time.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al · 2025
Closest in time.
Agent4edu: Generating learner response data by generative agents for intelligent education systems
Weibo Gao, Qi Liu, Linan Yue, Fangzhou Yao, Rui Lv, Zheng Zhang, Hao Wang, and Zhenya Huang · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Cos: Chain-of-shot prompting for long video understanding
Jian Hu, Zixu Cheng, Chenyang Si, Wei Li, and Shaogang Gong · 2025
Closest in time.
Visual-rft: Visual reinforcement fine-tuning
Ziyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong, Yuhang Cao, Haodong Duan, Dahua Lin, and Jiaqi Wang · 2025
Closest in time.
Retrieval-augmented visual question answering via built-in autoregressive search engines
Xinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang, Biqing Qi, and Bowen Zhou · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.
Eventvad: Training-free event-aware video anomaly detection
Yihua Shao, Haojin He, Sijie Li, Siyu Chen, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma, et al · 2025
Closest in time.
Streaming video understanding and multi-round interaction with memory-enhanced knowledge
Haomiao Xiong, Zongxin Yang, Jiazuo Yu, Yunzhi Zhuge, Lu Zhang, Jiawen Zhu, and Huchuan Lu · 2025
Closest in time.
Lens: Multi-level evaluation of multimodal reasoning with large language models
Ruilin Yao, Bo Zhang, Jirui Huang, Xinwei Long, Yifang Zhang, Tianyu Zou, Yufei Wu, Shichao Su, Yifan Xu, Wenxi Zeng, et al · 2025
Closest in time.
Internlm-xcomposer2. 5-reward: A simple yet effective multi-modal reward model
Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenxi Wu, Yubo Ma, Haodong Duan, Wenwei Zhang, et al · 2025
Closest in time.
Futuresightdrive: Thinking visually with spatio-temporal cot for autonomous driving
Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu, Yifan Bai, Zheng Pan, Mu Xu, and Xing Wei · 2025
Closest in time.