Fetching the paper…
Reading the bibliography…
The advancement of Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs) and large vision-language models (LVLMs).
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Y. Tang, D. Ding, Y. Rao, Y. Zheng, D. Zhang, L. Zhao, J. Lu, and J. Zhou · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao · 2019
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
J. Xiao, X. Shang, A. Yao, and T.-S. Chua · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Seed-bench: Benchmarking multimodal llms with generative comprehension
B. Li, R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan · 2023
Earlier work this paper cites.
Video-llava: Learning united visual representation by alignment before projection
B. Lin, Y. Ye, B. Zhu, J. Cui, M. Ning, P. Jin, and L. Yuan · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2023
Earlier work this paper cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
M. Maaz, H. Rasheed, S. Khan, and F. S. Khan · 2023
Earlier work this paper cites.
Mmbench: Benchmarking end-to-end multi-modal dnns and understanding their hardware-software implications
C. Xu, X. Hou, J. Liu, C. Li, T. Huang, X. Zhu, M. Niu, L. Sun, P. Tang, T. Xu, et al · 2023
Earlier work this paper cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
H. Zhang, X. Li, and L. Bing · 2023
Earlier work this paper cites.
Llama 3 model card, 2024
AI@Meta · 2024
Earlier work this paper cites.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Earlier work this paper cites.
Sharegpt4v: Improving large multi-modal models with better captions
L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin · 2024
Earlier work this paper cites.
Are we on the right way for evaluating large vision-language models?
L. Chen, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, J. Wang, Y. Qiao, D. Lin, et al · 2024
Earlier work this paper cites.
Sharegpt4video: Improving video understanding and generation with better captions
L. Chen, X. Wei, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, Z. Tang, L. Yuan, et al · 2024
Earlier work this paper cites.
Z. Chen, W. Wang, Y. Cao, Y. Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al · 2024
Earlier work this paper cites.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma, et al · 2024
Earlier work this paper cites.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al · 2024
Cited alongside, same era.
Tvbench: Redesigning video-language evaluation
D. Cores, M. Dorkenwald, M. Mucientes, C. G. Snoek, and Y. M. Asano · 2024
Cited alongside, same era.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
C. Fu, Y. Dai, Y. Luo, L. Li, S. Ren, R. Zhang, Z. Wang, C. Zhou, Y. Shen, M. Zhang, et al · 2024
Cited alongside, same era.
Sciverse
Z. Guo, R. Zhang, H. Chen, J. Gao, P. Gao, H. Li, and P.-A. Heng · 2024
Cited alongside, same era.
S. Han, W. Huang, H. Shi, L. Zhuo, X. Su, S. Zhang, X. Zhou, X. Qi, Y. Liao, and S. Liu · 2024
Internvideo2: Scaling foundation models for multimodal video understanding
Y. Wang, K. Li, X. Li, J. Yu, Y. He, G. Chen, B. Pei, R. Zheng, Z. Wang, Y. Shi, et al · 2024
Later among the works it cites.
Longvideobench: A benchmark for long-context interleaved video-language understanding
H. Wu, D. Li, B. Chen, and J. Li · 2024
Later among the works it cites.
Pllava: Parameter-free llava extension from images to videos for video dense captioning
L. Xu, Y. Zhao, D. Zhou, Z. Lin, S. K. Ng, and J. Feng · 2024
Later among the works it cites.
Visa: Reasoning video object segmentation via large language models
C. Yan, H. Wang, S. Yan, X. Jiang, Y. Hu, G. Kang, W. Xie, and E. Gavves · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Y. Yao, T. Yu, A. Zhang, C. Wang, J. Cui, H. Zhu, T. Cai, H. Li, W. Zhao, Z. He, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024
C. He, R. Luo, Y. Bai, S. Hu, Z. L. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, J. Liu, L. Qi, Z. Liu, and M. Sun · 2024
Cited alongside, same era.
Cogvlm2: Visual language models for image and video understanding
W. Hong, W. Wang, M. Ding, W. Yu, Q. Lv, Y. Wang, Y. Cheng, S. Huang, J. Ji, Z. Xue, et al · 2024
Cited alongside, same era.
Llava-onevision: Easy visual task transfer
B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y. Li, Z. Liu, et al · 2024
Cited alongside, same era.
Aria: An open multimodal native mixture-of-experts model
D. Li, Y. Liu, H. Wu, Y. Wang, Z. Shen, B. Qu, X. Niu, F. Zhou, C. Huang, Y. Li, et al · 2024
Cited alongside, same era.
Mvbench: A comprehensive multi-modal video understanding benchmark
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo, et al · 2024
Cited alongside, same era.
Mmbench: Is your multi-modal model an all-around player?
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al · 2024
Cited alongside, same era.
Videogpt+: Integrating image and video encoders for enhanced video understanding
M. Maaz, H. Rasheed, S. Khan, and F. S. Khan · 2024
Cited alongside, same era.
Later among the works it cites.
mplug-owl3: Towards long image-sequence understanding in multi-modal large language models
J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou · 2024
Later among the works it cites.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
R. Zhang, D. Jiang, Y. Zhang, H. Lin, Z. Guo, P. Qiu, A. Zhou, P. Lu, K.-W. Chang, P. Gao, et al · 2024
Later among the works it cites.
Video instruction tuning with synthetic data
Y. Zhang, J. Wu, W. Li, B. Li, Z. Ma, Z. Liu, and C. Li · 2024
Later among the works it cites.
Mlvu: A comprehensive benchmark for multi-task long video understanding
J. Zhou, Y. Shu, B. Zhao, B. Wu, S. Xiao, X. Yang, Y. Xiong, B. Zhang, T. Huang, and Z. Liu · 2024
Later among the works it cites.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos
K. Hu, P. Wu, F. Pu, W. Xiao, Y. Zhang, X. Yue, B. Li, and Z. Liu · 2025
Closest in time.
D. Jiang, R. Zhang, Z. Guo, Y. Li, Y. Qi, X. Chen, L. Wang, J. Jin, C. Guo, S. Yan, et al · 2025
Closest in time.
Llamav-o1: Rethinking step-by-step visual reasoning in llms
O. Thawakar, D. Dissanayake, K. More, R. Thawkar, A. Heakl, N. Ahsan, Y. Li, M. Zumri, J. Lahoud, R. M. Anwer, et al · 2025
Closest in time.
Internvideo2.5: Empowering video mllms with long and rich context modeling
Y. Wang, X. Li, Z. Yan, Y. He, J. Yu, X. Zeng, C. Wang, C. Ma, H. Huang, J. Gao, M. Dou, K. Chen, W. Wang, Y. Qiao, Y. Wang, and L. Wang · 2025
Closest in time.
Videollama 3: Frontier multimodal foundation models for image and video understanding
B. Zhang, K. Li, Z. Cheng, Z. Hu, Y. Yuan, G. Chen, S. Leng, Y. Jiang, H. Zhang, X. Li, et al · 2025
Closest in time.
Mmvu: Measuring expert-level multi-discipline video understanding
Y. Zhao, L. Xie, H. Zhang, G. Gan, Y. Long, Z. Hu, T. Hu, W. Chen, C. Li, J. Song, et al · 2025
Closest in time.
Y. Zhao, Y. Zeng, Y. Qi, Y. Liu, L. Chen, Z. Chen, X. Bao, J. Zhao, and F. Zhao · 2025
Closest in time.