Fetching the paper…
Reading the bibliography…
Visual reasoning is a core component of human intelligence and a critical capability for advanced multimodal models.
Artificial general intelligence
B. Goertzel and C. Pennachin · 2007
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server, 2015
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
T.-Y. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Dollár · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
The kinetics human action video dataset, 2017
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Earlier work this paper cites.
Reflective decoding network for image captioning
L. Ke, W. Pei, R. Li, X. Shen, and Y.-W. Tai · 2019
Earlier work this paper cites.
Raven: A dataset for relational and analogical visual reasoning
C. Zhang, F. Gao, B. Jia, Y. Zhu, and S.-C. Zhu · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
M. Mathew, D. Karatzas, and C. Jawahar · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Pointgpt: Auto-regressively generative pre-training from point clouds
G. Chen, M. Wang, Y. Yang, K. Yu, L. Yuan, and Y. Yue · 2023
Earlier work this paper cites.
Sharegpt4v: Improving large multi-modal models with better captions
L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin · 2023
Earlier work this paper cites.
Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset
S. Chen, H. Li, Q. Wang, Z. Zhao, M. Sun, X. Zhu, and J. Liu · 2023
Earlier work this paper cites.
Language models can be logical solvers
J. Feng, R. Xu, J. Hao, H. Sharma, Y. Shen, D. Zhao, and W. Chen · 2023
Earlier work this paper cites.
Z. Guo, R. Zhang, X. Zhu, Y. Tang, X. Ma, J. Han, K. Chen, P. Gao, X. Li, H. Li, et al · 2023
Earlier work this paper cites.
Infimm-eval: Complex open-ended reasoning evaluation for multi-modal large language models
X. Han, Q. You, Y. Liu, W. Chen, H. Zheng, K. Mrini, X. Lin, Y. Wang, B. Zhai, J. Yuan, et al · 2023
Earlier work this paper cites.
Codecot: Tackling code syntax errors in cot reasoning for code generation
D. Huang, Q. Bu, Y. Qing, and H. Cui · 2023
Earlier work this paper cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Earlier work this paper cites.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Earlier work this paper cites.
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Earlier work this paper cites.
Logicot: Logical chain-of-thought instruction-tuning
H. Liu, Z. Teng, L. Cui, C. Zhang, Q. Zhou, and Y. Zhang · 2023
Earlier work this paper cites.
Agentbench: Evaluating llms as agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, et al · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2023
Earlier work this paper cites.
Egoschema: A diagnostic benchmark for very long-form video language understanding
K. Mangalam, R. Akshulakov, and J. Malik · 2023
Earlier work this paper cites.
Improving multimodal datasets with image captioning
T. Nguyen, S. Y. Gadre, G. Ilharco, S. Oh, and L. Schmidt · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, et al · 2023
Cited alongside, same era.
Symbol-llm: Towards foundational symbol-centric interface for large language models
F. Xu, Z. Wu, Q. Sun, S. Ren, F. Yuan, S. Yuan, Q. Lin, Y. Qiao, and J. Liu · 2023
Cited alongside, same era.
Antgpt: Can large language models help long-term action anticipation from videos?
Q. Zhao, S. Wang, C. Zhang, C. Fu, M. Q. Do, N. Agarwal, K. Lee, and C. Sun · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Cited alongside, same era.
Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning, 2024
Y. Wang, W. Chen, X. Han, X. Lin, H. Zhao, Y. Liu, B. Zhai, J. Yuan, Q. You, and H. Yang · 2024
Later among the works it cites.
Logicvista: Multimodal llm logical reasoning benchmark in visual contexts
Y. Xiao, E. Sun, T. Liu, and W. Wang · 2024
Later among the works it cites.
Mini-omni: Language models can hear, talk while thinking in streaming, 2024
Z. Xie and C. Wu · 2024
Later among the works it cites.
Vlm-grounder: A vlm agent for zero-shot 3d visual grounding
R. Xu, Z. Huang, T. Wang, Y. Chen, J. Pang, and D. Lin · 2024
Later among the works it cites.
Chatglm-math: Improving math problem-solving in large language models with a self-critique pipeline
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024
A. Ahmadian, C. Cremer, M. Gallé, M. Fadaee, J. Kreutzer, O. Pietquin, A. Üstün, and S. Hooker · 2024
Cited alongside, same era.
Visreas: Complex visual reasoning with unanswerable questions, 2024
S. N. Akter, S. Lee, Y. Chang, Y. Bisk, and E. Nyberg · 2024
Cited alongside, same era.
Z. Chen, W. Wang, Y. Cao, Y. Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al · 2024
Cited alongside, same era.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma, et al · 2024
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al · 2024
Cited alongside, same era.
Egothink: Evaluating first-person perspective thinking capability of vision-language models
S. Cheng, Z. Guo, J. Wu, K. Fang, P. Li, H. Liu, and Y. Liu · 2024
Cited alongside, same era.
Moshi: a speech-text foundation model for real-time dialogue
A. Défossez, L. Mazaré, M. Orsini, A. Royer, P. Pérez, H. Jégou, E. Grave, and N. Zeghidour · 2024
Cited alongside, same era.
Benchmarking and improving detail image caption
H. Dong, J. Li, B. Wu, J. Wang, Y. Zhang, and H. Guo · 2024
Cited alongside, same era.
Y. Xu, X. Liu, X. Liu, Z. Hou, Y. Li, X. Zhang, Z. Wang, A. Zeng, Z. Du, W. Zhao, et al · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al · 2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Y. Yao, T. Yu, A. Zhang, C. Wang, J. Cui, H. Zhu, T. Cai, H. Li, W. Zhao, Z. He, et al · 2024
Later among the works it cites.
J. Ye, G. Li, S. Gao, C. Huang, Y. Wu, S. Li, X. Fan, S. Dou, Q. Zhang, T. Gui, et al · 2024
Later among the works it cites.
mplug-owl3: Towards long image-sequence understanding in multi-modal large language models, 2024
J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou · 2024
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen · 2024
Later among the works it cites.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
R. Zhang, D. Jiang, Y. Zhang, H. Lin, Z. Guo, P. Qiu, A. Zhou, P. Lu, K.-W. Chang, Y. Qiao, et al · 2024
Later among the works it cites.
Opencodereasoning: Advancing data distillation for competitive coding
W. U. Ahmad, S. Narenthiran, S. Majumdar, A. Ficek, S. Jain, J. Huang, V. Noroozi, and B. Ginsburg · 2025
Closest in time.
Digi-q: Learning q-value functions for training device-control agents
H. Bai, Y. Zhou, L. E. Li, S. Levine, and A. Kumar · 2025
Closest in time.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin · 2025
Closest in time.
R1-v: Reinforcing super generalization ability in vision-language models with less than $3
L. Chen, L. Li, H. Zhao, Y. Song, and Vinci · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark
Y. Hao, J. Gu, H. W. Wang, L. Li, Z. Yang, L. Wang, and Y. Cheng · 2025
Closest in time.
Structured chain-of-thought prompting for code generation
J. Li, G. Li, Y. Li, and Z. Jin · 2025
Closest in time.
Visual-rft: Visual reinforcement fine-tuning
Z. Liu, Z. Sun, Y. Zang, X. Dong, Y. Cao, H. Duan, D. Lin, and J. Wang · 2025
Closest in time.
Z. Liu, Y. Zhang, F. Liu, C. Zhang, Y. Sun, and J. Wang · 2025
Closest in time.
Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning
F. Meng, L. Du, Z. Liu, Z. Zhou, Q. Lu, D. Fu, B. Shi, W. Wang, J. He, K. Zhang, et al · 2025
Closest in time.
Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl
Y. Peng, G. Zhang, M. Zhang, Z. You, J. Liu, Q. Zhu, K. Yang, X. Xu, X. Geng, and X. Yang · 2025
Closest in time.
Vlm-r1: A stable and generalizable r1-style large vision-language model
H. Shen, Z. Zhang, K. Zhao, Q. Zhang, R. Xu, and T. Zhao · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning, March 2025
Q. Team · 2025
Closest in time.
Learning visual grounding from generative vision and language model
S. Wang, D. Kim, A. Taalimi, C. Sun, and W. Kuo · 2025
Closest in time.
Visualprm: An effective process reward model for multimodal reasoning
W. Wang, Z. Gao, L. Chen, Z. Chen, J. Zhu, X. Zhao, Y. Liu, Y. Cao, S. Ye, X. Zhu, et al · 2025
Closest in time.
Open-r1-video
X. Wang and P. Peng · 2025
Closest in time.
R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization
Y. Yang, X. He, H. Pan, X. Jiang, Y. Deng, X. Yang, H. Lu, D. Yin, F. Rao, M. Zhu, B. Zhang, and W. Chen · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025
J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, Y. Duan, H. Tian, W. Su, J. Shao, Z. Gao, E. Cui, Y. Cao, Y. Liu, X. Wei, H. Zhang, H. Wang, W. Xu, H. Li, J. Wang, D. Chen, S. Li, Y. He, T. Jiang, J. Luo, Y. Wang, C. He, B. Shi, X. Zhang, W. Shao, J. He, Y. Xiong, W. Qu, P. Sun, P. Jiao, H. Lv, L. Wu, K. Zhang, H. Deng, J. Ge, K. Chen, L. Wang, M. Dou, L. Lu, X. Zhu, T. Lu, D. Lin, Y. Qiao, J. Dai, and W. Wang · 2025
Closest in time.