Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (LVLMs) have achieved significant success in multimodal tasks, with multimodal chain-of-thought (MCoT) further enhancing performance and interpretability.
Is visual thinking “imageless thought”?
Robert G Kunzendorf, Kara Young, Tamara Beecy, and Karen Beals · 2000
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Earlier work this paper cites.
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al · 2023
Earlier work this paper cites.
See and learn more: Dense caption-aware representation for visual question answering
Yandong Bi, Huajie Jiang, Yongli Hu, Yanfeng Sun, and Baocai Yin · 2023
Earlier work this paper cites.
Measuring and improving chain-of-thought reasoning in vision-language models
Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran · 2023
Earlier work this paper cites.
Deepanway Ghosal, Navonil Majumder, Roy Ka-Wei Lee, Rada Mihalcea, and Soujanya Poria · 2023
Earlier work this paper cites.
Visual programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi · 2023
Earlier work this paper cites.
Visual instruction tuning, 2023
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Earlier work this paper cites.
Chain of images for intuitively reasoning
Fanxu Meng, Haotong Yang, Yiding Wang, and Muhan Zhang · 2023
Earlier work this paper cites.
Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages
Libo Qin, Qiguang Chen, Fuxuan Wei, Shijue Huang, and Wanxiang Che · 2023
Earlier work this paper cites.
Vipergpt: Visual inference via python execution for reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick · 2023
Earlier work this paper cites.
Label words are anchors: An information flow perspective for understanding in-context learning
Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun · 2023
Earlier work this paper cites.
Thinking like an expert: Multimodal hypergraph-of-thought (hot) reasoning to boost foundation modals
Fanglong Yao, Changyuan Tian, Jintao Liu, Zequn Zhang, Qing Liu, Li Jin, Shuchao Li, Xiaoyu Li, and Xian Sun · 2023
Cited alongside, same era.
Multimodal chain-of-thought reasoning in language models
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola · 2023
Cited alongside, same era.
Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models
Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Cited alongside, same era.
Through the lens of core competency: Survey on evaluation of large language models
Whiteboard-of-thought: Thinking step-by-step across modalities
Sachit Menon, Richard Zemel, and Carl Vondrick · 2024
Later among the works it cites.
Compositional chain-of-thought prompting for large multimodal models
Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig · 2024
Later among the works it cites.
Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning
Debjyoti Mondal, Suraj Modi, Subhadarshi Panda, Rituraj Singh, and Godawari Sudhakar Rao · 2024
Later among the works it cites.
Openai gpt-4o, 2024
Openai · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos, 2024
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feichtenhofer · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ziyu Zhuang, Qiguang Chen, Longxuan Ma, Mingda Li, Yi Han, Yushan Qian, Haopeng Bai, Zixian Feng, Weinan Zhang, and Ting Liu · 2023
Cited alongside, same era.
Qiguang Chen, Libo Qin, Jiaqi Wang, Jinxuan Zhou, and Wanxiang Che · 2024
Cited alongside, same era.
M 3 CoT: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought
Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che · 2024
Cited alongside, same era.
Comt: A novel benchmark for chain of multi-modal thought on large vision-language models
Zihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei, Xiaocheng Feng, Wanxiang Che, Min Li, and Libo Qin · 2024
Cited alongside, same era.
Anole: An open, autoregressive, native large multimodal models for interleaved image-text generation
Ethan Chern, Jiadi Su, Yan Ma, and Pengfei Liu · 2024
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Cited alongside, same era.
IsoBench: Benchmarking multimodal foundation models on isomorphic representations
Deqing Fu, Ruohao Guo, Ghazal Khalighinejad, Ollie Liu, Bhuwan Dhingra, Dani Yogatama, Robin Jia, and Willie Neiswanger · 2024
Cited alongside, same era.
Interleaved-modal chain-of-thought
Jun Gao, Yongqi Li, Ziqiang Cao, and Wenjie Li · 2024
Cited alongside, same era.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie · 2024
Later among the works it cites.
T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering
Lei Wang, Yi Hu, Jiabang He, Xing Xu, Ning Liu, Hui Liu, and Heng Tao Shen · 2024
Later among the works it cites.
Enhancing human-like multimodal reasoning: a new challenging dataset and comprehensive framework
Jingxuan Wei, Cheng Tan, Zhangyang Gao, Linzhuang Sun, Siyuan Li, Bihui Yu, Ruifeng Guo, and Stan Z Li · 2024
Later among the works it cites.
V?: Guided visual search as a core mechanism in multimodal llms
Penghao Wu and Saining Xie · 2024
Later among the works it cites.
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Later among the works it cites.
Cocot: Contrastive chain-of-thought prompting for large multimodal models with multiple image inputs
Daoan Zhang, Junming Yang, Hanjia Lyu, Zijian Jin, Yuan Yao, Mingkai Chen, and Jiebo Luo · 2024
Later among the works it cites.
Image-of-thought prompting for visual reasoning refinement in multimodal large language models
Qiji Zhou, Ruochen Zhou, Zike Hu, Panzhong Lu, Siyang Gao, and Yue Zhang · 2024
Later among the works it cites.
Imagine while reasoning in space: Multimodal visualization-of-thought
Chengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia, Shaoguang Mao, Li Dong, Ivan Vulić, and Furu Wei · 2025
Closest in time.
Openai o3-mini system card, 2025
Openai · 2025
Closest in time.
Openthinkimg: Learning to think with images via visual tool reinforcement learning
Zhaochen Su, Linjie Li, Mingyang Song, Yunzhuo Hao, Zhengyuan Yang, Jun Zhang, Guanjie Chen, Jiawei Gu, Juntao Li, Xiaoye Qu, et al · 2025
Closest in time.
Lifting the veil on visual information flow in mllms: Unlocking pathways to faster inference
Hao Yin, Guangzong Si, and Zilei Wang · 2025
Closest in time.
Yongheng Zhang, Xu Liu, Ruihan Tao, Qiguang Chen, Hao Fei, Wanxiang Che, and Libo Qin · 2025
Closest in time.