Fetching the paper…
Reading the bibliography…
Large Vision-Language Models (LVLMs) have recently demonstrated amazing success in multi-modal tasks, including advancements in Multi-modal Chain-of-Thought (MCoT) reasoning.
The measurement of observer agreement for categorical data
Landis, J. R.; and Koch, G. G. 1977 · 1977
Earlier work this paper cites.
Core networks for visual-concrete and abstract thought content: a brain electric microstate analysis
Lehmann, D.; Pascual-Marqui, R. D.; Strik, W. K.; and Koenig, T. 2010 · 2010
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
Zellers, R.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 2019
Earlier work this paper cites.
Jhu-crowd++: Large-scale crowd counting dataset and a benchmark method
Sindagi, V. A.; Yasarla, R.; and Patel, V. M. 2020 · 2020
Earlier work this paper cites.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Le Bras, R.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Cao, J.; and Xiao, J. 2022 · 2022
Earlier work this paper cites.
Abstract visual reasoning with tangram shapes
Ji, A.; Kojima, N.; Rush, N.; Suhr, A.; Vong, W. K.; Hawkins, R. D.; and Artzi, Y. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022 · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022 · 2022
Earlier work this paper cites.
A-okvqa: A benchmark for visual question answering using world knowledge
Schwenk, D.; Khandelwal, A.; Clark, C.; Marino, K.; and Mottaghi, R. 2022 · 2022
Earlier work this paper cites.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Earlier work this paper cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Cited alongside, same era.
Reasoning Implicit Sentiment with Chain-of-Thought Prompting
Fei, H.; Li, B.; Liu, Q.; Bing, L.; Li, F.; and Chua, T.-S. 2023 · 2023
Cited alongside, same era.
ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
Golovneva, O.; Chen, M. P.; Poff, S.; Corredor, M.; Zettlemoyer, L.; Fazel-Zarandi, M.; and Celikyilmaz, A. 2023 · 2023
Cited alongside, same era.
Generating Images with Multimodal Language Models
Koh, J. Y.; Fried, D.; and Salakhutdinov, R. 2023 · 2023
Cited alongside, same era.
Retrieval-augmented multi-modal chain-of-thoughts reasoning for large language models
Liu, B.; Lyu, C.; Min, Z.; Wang, Z.; Su, J.; and Wang, L. 2023 · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
Multi-modal latent space learning for chain-of-thought reasoning in language models
He, L.; Li, Z.; Cai, X.; and Wang, P. 2024 · 2024
Closest in time.
What matters when building vision-language models?
Laurencon, H.; Tronchon, L.; Cord, M.; and Sanh, V. 2024 · 2024
Closest in time.
Multimodal Reasoning with Multimodal Knowledge Graph
Lee, J.; Wang, Y.; Li, J.; and Zhang, M. 2024 · 2024
Closest in time.
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chain of Images for Intuitively Reasoning
Meng, F.; Yang, H.; Wang, Y.; and Zhang, M. 2023 · 2023
Cited alongside, same era.
Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages
Qin, L.; Chen, Q.; Wei, F.; Huang, S.; and Che, W. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Cited alongside, same era.
The role of chain-of-thought in complex vision-language reasoning task
Wu, Y.; Zhang, P.; Xiong, W.; Oguz, B.; Gee, J. C.; and Nie, Y. 2023 · 2023
Cited alongside, same era.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; et al. 2023 · 2023
Cited alongside, same era.
Multimodal chain-of-thought reasoning in language models
Zhang, Z.; Zhang, A.; Li, M.; Zhao, H.; Karypis, G.; and Smola, A. 2023 · 2023
Cited alongside, same era.
Unlocking the Boundaries of Thought: A Reasoning Granularity Framework to Quantify and Optimize Chain-of-Thought
Chen, Q.; Qin, L.; Jiaqi, W.; Jinxuan, Z.; and Che, W. 2024a
Cited in the paper.
Lin, W.; Wei, X.; An, R.; Gao, P.; Zou, B.; Luo, Y.; Huang, S.; Zhang, S.; and Li, H. 2024 · 2024
Closest in time.
DeepSeek-VL: Towards Real-World Vision-Language Understanding
Lu, H.; Liu, W.; Zhang, B.; Wang, B.; Dong, K.; Liu, B.; Sun, J.; Ren, T.; Li, Z.; Yang, H.; Sun, Y.; Deng, C.; Xu, H.; Xie, Z.; and Ruan, C. 2024 · 2024
Closest in time.
KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning
Mondal, D.; Modi, S.; Panda, S.; Singh, R.; and Rao, G. S. 2024 · 2024
Closest in time.
Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning
Tan, C.; Wei, J.; Sun, L.; Gao, Z.; Li, S.; Yu, B.; Guo, R.; and Li, S. Z. 2024 · 2024
Closest in time.
Faithful Logical Reasoning via Symbolic Chain-of-Thought
Xu, J.; Fei, H.; Pan, L.; Liu, Q.; Lee, M.-L.; and Hsu, W. 2024 · 2024
Closest in time.
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Zhan, J.; Dai, J.; Ye, J.; Zhou, Y.; Zhang, D.; Liu, Z.; Zhang, X.; Yuan, R.; Zhang, G.; Li, L.; et al. 2024 · 2024
Closest in time.
Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models
Zheng, G.; Yang, B.; Tang, J.; Zhou, H.-Y.; and Yang, S. 2024 · 2024
Closest in time.