Fetching the paper…
Reading the bibliography…
Large vision-language models (LVLMs) offer a novel capability for performing in-context learning (ICL) in Visual QA.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning · 2018
Earlier work this paper cites.
Multilingual constituency parsing with self-attention and pre-training
Nikita Kitaev, Steven Cao, and Dan Klein · 2018
Earlier work this paper cites.
Learning by abstraction: The neural state machine
Drew Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Earlier work this paper cites.
Dynamic language binding in relational visual reasoning
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Earlier work this paper cites.
Roses are red, violets are blue… but should vqa expect them to?
Corentin Kervadec, Grigory Antipov, Moez Baccouche, and Christian Wolf · 2021
Earlier work this paper cites.
Coarse-to-fine reasoning for visual question answering
Binh X. Nguyen, Tuong Khanh Long Do, Huy Tran, Erman Tjiputra, Quang D. Tran, and A. Nguyen · 2021
Earlier work this paper cites.
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
Successive prompting for decomposing complex questions
Dheeru Dua, Shivanshu Gupta, Sameer Singh, and Matt Gardner · 2022
Earlier work this paper cites.
Demystifying prompts in language models via perplexity estimation, 2022
Hila Gonen, Srini Iyer, Terra Blevins, Noah A. Smith, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Instruction induction: From few examples to natural language task descriptions
Or Honovich, Uri Shaham, Samuel R Bowman, and Omer Levy · 2022
Earlier work this paper cites.
What makes good in-context examples for GPT-3?
Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen · 2022
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Measuring and narrowing the compositionality gap in language models
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis · 2022
Cited alongside, same era.
Learning to retrieve prompts for in-context learning
Ohad Rubin, Jonathan Herzig, and Jonathan Berant · 2022
Cited alongside, same era.
An information-theoretic approach to prompt engineering without ground truth labels
Taylor Sorensen, Joshua Robinson, Christopher Rytting, Alexander Shaw, Kyle Rogers, Alexia Delorey, Mahmoud Khalil, Nancy Fulda, and David Wingate · 2022
Cited alongside, same era.
Cric: A vqa dataset for compositional reasoning on vision and commonsense
Difei Gao, Ruiping Wang, Shiguang Shan, and Xilin Chen · 2023
Later among the works it cites.
Decomposed prompting: A modular approach for solving complex tasks
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal · 2023
Later among the works it cites.
Diverse demonstrations improve in-context compositional generalization
Itay Levy, Ben Bogin, and Jonathan Berant · 2023
Later among the works it cites.
Mimic-it: Multi-modal in-context instruction tuning
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li, and Ziwei Liu · 2023
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Iteratively prompt pre-trained language models for chain of thought
Boshi Wang, Xiang Deng, and Huan Sun · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Cited alongside, same era.
Knowledge-based visual question generation
Jiayuan Xie, Wenhao Fang, Yi Cai, Qingbao Huang, and Qing Li · 2022
Cited alongside, same era.
An Explanation of In-context Learning as Implicit Bayesian Inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Cited alongside, same era.
An empirical study of gpt-3 for few-shot knowledge-based vqa
Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang · 2022
Cited alongside, same era.
Automatic chain of thought prompting in large language models
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola · 2022
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Later among the works it cites.
Large language models as general pattern machines
Suvir Mirchandani, Fei Xia, Pete Florence, Brian Ichter, Danny Driess, Montserrat Gonzalez Arenas, Kanishka Rao, Dorsa Sadigh, and Andy Zeng · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Scene graph refinement network for visual question answering
Tianwen Qian, Jingjing Chen, Shaoxiang Chen, Bo Wu, and Yu-Gang Jiang · 2023
Later among the works it cites.
Multilingual LLMs are better cross-lingual in-context learners with alignment
Eshaan Tanwar, Subhabrata Dutta, Manish Borthakur, and Tanmoy Chakraborty · 2023
Later among the works it cites.
Self-adaptive in-context learning: An information compression perspective for in-context example selection and ordering, 2023
Zhiyong Wu, Yaoxiang Wang, Jiacheng Ye, and Lingpeng Kong · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi · 2023
Later among the works it cites.
Fair attention network for robust visual question answering
Yandong Bi, Huajie Jiang, Yongli Hu, Yanfeng Sun, and Baocai Yin · 2024
Closest in time.
See and learn more: Dense caption-aware representation for visual question answering
Yandong Bi, Huajie Jiang, Yongli Hu, Yanfeng Sun, and Baocai Yin · 2024
Closest in time.
Question type-aware debiasing for test-time visual question answering model adaptation
Jin Liu, Jialong Xie, Fengyu Zhou, and Shengfeng He · 2024
Closest in time.
Pro-tuning: Unified prompt tuning for vision tasks
Xing Nie, Bolin Ni, Jianlong Chang, Gaofeng Meng, Chunlei Huo, Shiming Xiang, and Qi Tian · 2024
Closest in time.