Fetching the paper…
Reading the bibliography…
This paper introduces the novel task of multimodal puzzle solving, framed within the context of visual question-answering.
Nouvelles applications des paramètres continus à la théorie des formes quadratiques. deuxième mémoire. recherches sur les parallélloèdres primitifs
Georges Voronoi. 1908 · 1908
Earlier work this paper cites.
Critical thinking. an introduction to logic and scientific method
Max Black. 1948 · 1948
Earlier work this paper cites.
Reentrant polygon clipping
Ivan E Sutherland and Gary W Hodgman. 1974 · 1974
Earlier work this paper cites.
The solution of the four-color-map problem
Kenneth Appel and Wolfgang Haken. 1977 · 1977
Earlier work this paper cites.
The Moscow Puzzles: 359 Mathematical Recreations
Boris A Kordemsky. 1992 · 1992
Earlier work this paper cites.
Modular elliptic curves and fermat’s last theorem
Andrew Wiles. 1995 · 1995
Earlier work this paper cites.
Dancing links
Donald E Knuth. 2000 · 2000
Earlier work this paper cites.
Tiling with dominoes
Nathan S Mendelsohn. 2004 · 2004
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
Kvqa: Knowledge-aware visual question answering
Naganand Yadati Sanket Shah, Anand Mishra and Partha Pratim Talukdar. 2019 · 2019
Cited alongside, same era.
A-okvqa: A benchmark for visual question answering using world knowledge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022 · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Albert Li, Pascale Fung, and Steven C. H. Hoi. 2023 · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Google Gemini Team. 2023 · 2023
Later among the works it cites.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Curtis G Northcutt, Anish Athalye, and Jonas Mueller. 2021 · 2021
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022 · 2022
Cited alongside, same era.
Solutio problematis ad geometriam situs pertinentis
Leonhard Euler. 1741
Cited in the paper.
Gpt-4v(ision) system card
OpenAI. 2023 · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. 2023 · 2023
Later among the works it cites.