Fetching the paper…
Reading the bibliography…
Answering visual queries is a complex task that requires both visual processing and reasoning.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering, May 2019
Drew A. Hudson and Christopher D. Manning · 1902
Earlier work this paper cites.
Information streams sharing a finite buffer
E.W. Dijkstra · 1972
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Obtaining Faithful Interpretations from Compositional Neural Networks, Sept. 2020
Sanjay Subramanian, Ben Bogin, Nitish Gupta, Tomer Wolfson, Sameer Singh, Jonathan Berant, and Matt Gardner · 2005
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, Apr. 2016
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2016
Earlier work this paper cites.
Learning to Reason: End-to-End Module Networks for Visual Question Answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko · 2017
Earlier work this paper cites.
Inferring and Executing Programs for Visual Reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Compositional Attention Networks for Machine Reasoning
Drew A. Hudson and Christopher D. Manning · 2018
Earlier work this paper cites.
Multimodal Explanations: Justifying Decisions and Pointing to the Evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach · 2018
Earlier work this paper cites.
Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding
Kexin Yi, Jiajun Wu, Chuang Gan, A. Torralba, Pushmeet Kohli, and J. Tenenbaum · 2018
Earlier work this paper cites.
Yundong Zhang, Juan Carlos Niebles, and Alvaro Soto · 2018
Earlier work this paper cites.
Systematic Generalization: What Is Required and Can It Be Learned?, Apr. 2019
Dzmitry Bahdanau, Shikhar Murty, Michael Noukhovitch, Thien Huu Nguyen, Harm de Vries, and Aaron Courville · 2019
Earlier work this paper cites.
The Consciousness Prior, Dec. 2019
Yoshua Bengio · 2019
Earlier work this paper cites.
Language-conditioned graph networks for relational reasoning
Ronghang Hu, Anna Rohrbach, Trevor Darrell, and Kate Saenko · 2019
Earlier work this paper cites.
Learning by abstraction: The neural state machine
Drew Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
Visual reasoning by progressive module networks
Seung Wook Kim, Makarand Tapaswi, and Sanja Fidler · 2019
Earlier work this paper cites.
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2020
Cited alongside, same era.
Learning to learn words from visual scenes
Dídac Surís, Dave Epstein, Heng Ji, Shih-Fu Chang, and Carl. Vondrick · 2020
Cited alongside, same era.
Interpretable Visual Question Answering by Reasoning on Dependency Trees
Qingxing Cao, Xiaodan Liang, Bailin Li, and Liang Lin · 2021
Cited alongside, same era.
Interpretable visual reasoning: A survey
Feijuan He, Yaxian Wang, Xianglin Miao, and Xia Sun · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Separating skills and concepts for novel visual question answering
Language models of code are few-shot commonsense learners
Aman Madaan, Shuyan Zhou, Uri Alon, Yiming Yang, and Graham Neubig · 2022
Later among the works it cites.
Doubly Right Object Recognition: A Why Prompt for Visual Rationales, Dec. 2022
Chengzhi Mao, Revant Teotia, Amrutha Sundar, Sachit Menon, Junfeng Yang, Xin Wang, and Carl Vondrick · 2022
Later among the works it cites.
Visual Classification via Description from Large Language Models, Dec. 2022
Sachit Menon and Carl Vondrick · 2022
Later among the works it cites.
Simple open-vocabulary object detection with vision transformers
Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, Xiao Wang, Xiaohua Zhai, Thomas Kipf, and Neil Houlsby · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spencer Whitehead, Hui Wu, Heng Ji, Rogerio Feris, and Kate Saenko · 2021
Cited alongside, same era.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Yan Zeng, Xinsong Zhang, and Hang Li · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan · 2022
Cited alongside, same era.
Open-vocabulary attribute detection
Maria A. Bravo, Sudhanshu Mittal, Simon Ging, and Thomas Brox · 2022
Cited alongside, same era.
Revisiting the" video" in video-language understanding
Shyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles · 2022
Cited alongside, same era.
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen · 2022
Cited alongside, same era.
Transform-retrieve-generate: Natural language-centric outside-knowledge visual question answering
Feng Gao, Qing Ping, Govind Thattai, Aishwarya Reganti, Ying Nian Wu, and Prem Natarajan · 2022
Cited alongside, same era.
Coarse-to-fine reasoning for visual question answering
Binh X Nguyen, Tuong Do, Huy Tran, Erman Tjiputra, Quang D Tran, and Anh Nguyen · 2022
Later among the works it cites.
Talm: Tool augmented language models
Aaron Parisi, Yao Zhao, and Noah Fiedel · 2022
Later among the works it cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun · 2022
Later among the works it cites.
Mumuqa: Multimedia multi-hop news question answering via cross-media knowledge extraction and grounding
Revant Gangi Reddy, Xilin Rui, Manling Li, Xudong Lin, Haoyang Wen, Jaemin Cho, Lifu Huang, Mohit Bansal, Avirup Sil, Shih-Fu Chang, et al · 2022
Later among the works it cites.
Reclip: A strong zero-shot baseline for referring expression comprehension
Sanjay Subramanian, Will Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, and Anna Rohrbach · 2022
Later among the works it cites.
Plug-and-play VQA: Zero-shot VQA by conjoining large pretrained models with zero training
Anthony Meng Huat Tiong, Junnan Li, Boyang Li, Silvio Savarese, and Steven C.H. Hoi · 2022
Later among the works it cites.
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang · 2022
Later among the works it cites.
Code4struct: Code generation for few-shot structured prediction from natural language
Xingyao Wang, Sha Li, and Heng Ji · 2022
Later among the works it cites.
Language models with image descriptors are strong few-shot video-language learners
Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei, Xudong Lin, Shuohang Wang, Ziyi Yang, Chenguang Zhu, Derek Hoiem, et al · 2022
Later among the works it cites.
Chain of Thought Prompting Elicits Reasoning in Large Language Models, Oct. 2022
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2022
Later among the works it cites.
Video graph transformer for video question answering
Junbin Xiao, Pan Zhou, Tat-Seng Chua, and Shuicheng Yan · 2022
Later among the works it cites.
An empirical study of gpt-3 for few-shot knowledge-based vqa
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang · 2022
Later among the works it cites.
Hitea: Hierarchical temporal-aware video-language pre-training
Qinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu, Qi Qian, Ji Zhang, and Fei Huang · 2022
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence · 2022
Later among the works it cites.
Language is not all you need: Aligning perception with language models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, et al · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Closest in time.