Fetching the paper…
Reading the bibliography…
Combining Large Language Models (LLMs) with external specialized tools (LLMs+tools) is a recent paradigm to solve multimodal tasks such as Visual Question Answering (VQA).
S. Antol
2015
Earlier work this paper cites.
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,” in
2017
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “LXMERT: Learning cross-modality encoder representations from transformers,” in
2019
Earlier work this paper cites.
M. Acharya, K. Kafle, and C. Kanan, “Tallyqa: Answering complex counting questions,” in
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “GQA: A new dataset for real-world visual reasoning and compositional question answering,” in
2019
Earlier work this paper cites.
A. Singh
2019
Earlier work this paper cites.
A. Suhr, S. Zhou, I. Zhang, H. Bai, and Y. Artzi, “A corpus for reasoning about natural language grounded in photographs,” in
2019
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “OK-VQA: A visual question answering benchmark requiring external knowledge,” in
2019
Earlier work this paper cites.
S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar, “KVQA: Knowledge-aware visual question answering,” in
2019
Earlier work this paper cites.
A. Mani, N. Yoo, W. Hinthorn, and O. Russakovsky, “Point and ask: Incorporating pointing into visual question answering,” arXiv, Tech. Rep., 2020
2020
Earlier work this paper cites.
T. B. Brown
2020
Earlier work this paper cites.
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.-W. Chang, “REALM: Retrieval-Augmented Language Model Pre-Training,” in
2020
Earlier work this paper cites.
K. Marino, X. Chen, D. Parikh, A. Gupta, and M. Rohrbach, “KRISP: Integrating Implicit and Symbolic Knowledge for Open-Domain Knowledge-Based VQA,” in
2021
Cited alongside, same era.
P. Vickers, N. Aletras, E. Monti, and L. Barrault, “In Factuality: Efficient Integration of Relevant Facts for Visual Question Answering,” in
2021
Cited alongside, same era.
P. Zhang
2021
Cited alongside, same era.
W. Kim, B. Son, and I. Kim, “Vilt: Vision-and-language transformer without convolution or region supervision,” 2021
2021
Cited alongside, same era.
J.-B. Alayrac
2022
Cited alongside, same era.
A. Chowdhery
2022
Cited alongside, same era.
T. Gupta and A. Kembhavi, “Visual programming: Compositional visual reasoning without training,” in
2023
Later among the works it cites.
D. Surís, S. Menon, and C. Vondrick, “Vipergpt: Visual inference via python execution for reasoning,” in
2023
Later among the works it cites.
2023
Later among the works it cites.
OpenAI, “GPT-4 Technical Report,”
2023
Later among the works it cites.
M. Besta
2023
Later among the works it cites.
O. Khattab
2023
Later among the works it cites.
T. Schick
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” in
2022
Cited alongside, same era.
S. Borgeaud
2022
Cited alongside, same era.
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-OKVQA: A benchmark for visual question answering using world knowledge,” in
2022
Cited alongside, same era.
M. Minderer
2022
Cited alongside, same era.
2022
Cited alongside, same era.
L. Ouyang
2022
Cited alongside, same era.
2023
Later among the works it cites.
A. Parisi, Y. Zhao, and N. Fiedel, “Talm: Tool augmented language models,” in
2023
Later among the works it cites.
B. Paranjape
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in
2023
Later among the works it cites.
M. Minderer, A. Gritsenko, and N. Houlsby, “Scaling open-vocabulary object detection,”
2023
Later among the works it cites.
N. H. Matthias Minderer, Alexey Gritsenko, “Scaling open-vocabulary object detection,”
2023
Later among the works it cites.
Gemini Team Google, “Gemini: A family of highly capable multimodal models,” in
2023
Later among the works it cites.