Fetching the paper…
Reading the bibliography…
Large multimodal models extend the impressive capabilities of large language models by integrating multimodal understanding abilities.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, et al. 2020 · 1901
Earlier work this paper cites.
Theory of fluid and crystallized intelligence: A critical experiment
Raymond Bernard Cattell. 1963 · 1963
Earlier work this paper cites.
Piaget’s theory
Jean Piaget. 1976 · 1976
Earlier work this paper cites.
Fluid concepts and creative analogies: Computer models of the fundamental mechanisms of thought
Charles Cole. 1996 · 1996
Earlier work this paper cites.
The origin of concepts
Susan Carey. 2000 · 2000
Earlier work this paper cites.
Superior pattern processing is the essence of the evolved human brain
Mark P. Mattson. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
Building machines that learn and think like people
Joshua B. Tenenbaum. 2018 · 2018
Earlier work this paper cites.
Human few-shot learning of compositional instructions
Brenden M. Lake, Tal Linzen, and Marco Baroni. 2019 · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019 · 2019
Cited alongside, same era.
Raven: A dataset for relational and analogical visual reasoning
Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, and Song-Chun Zhu. 2019 · 2019
Cited alongside, same era.
Is gpt-3 a good data annotator?
Bosheng Ding, Chengwei Qin, Linlin Liu, Lidong Bing, Shafiq R. Joty, and Boyang Albert Li. 2022 · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022 · 2022
Cited alongside, same era.
Emergent analogical reasoning in large language models
The conceptarc benchmark: Evaluating understanding and generalization in the arc domain
Arseny Moskvichev, Victor Vikram Odouard, and Melanie Mitchell. 2023 · 2023
Later among the works it cites.
Gpt-4v(ision) system card
OpenAI. 2023 · 2023
Later among the works it cites.
Cognitive architectures for language agents
Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L. Griffiths. 2023 · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v(ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. 2023 · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taylor W. Webb, Keith J. Holyoak, and Hongjing Lu. 2022 · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, John A. Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuan-Fang Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023 · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Google Gemini Team. 2023 · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a
Cited in the paper.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin R. Stone, et al. 2023b
Cited in the paper.
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2023 · 2023
Later among the works it cites.
Qwen-VL: A versatile vision-language model for understanding, localization, text reading, and beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2024 · 2024
Closest in time.
Phenomenal yet puzzling: Testing inductive reasoning capabilities of language models with hypothesis refinement
Linlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar, Valentina Pyatkin, Chandra Bhagavatula, Bailin Wang, Yoon Kim, Yejin Choi, Nouha Dziri, and Xiang Ren. 2024 · 2024
Closest in time.
A vision check-up for language models
Pratyusha Sharma, Tamar Rott Shaham, Manel Baradad, Stephanie Fu, Adrian Rodriguez-Munoz, Shivam Duggal, Phillip Isola, and Antonio Torralba. 2024 · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. 2024 · 2024
Closest in time.