Fetching the paper…
Reading the bibliography…
Compositional Reasoning (CR) entails grasping the significance of attributes, relations, and word order.
“Microsoft coco: Common objects in context”
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár and C Zitnick · 2014
Earlier work this paper cites.
“Lxmert: Learning Cross-modality Encoder Representations from Transformers”
Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
“Uniter: Universal Image-text Representation Learning”
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El, Faisal Ahmed, Zhe Gan, Yu Cheng and Jingjing Liu · 2020
Earlier work this paper cites.
“Oscar: Object-semantics Aligned Pre-training for Vision-language Tasks”
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong and Furu Wei · 2020
Earlier work this paper cites.
“Learning Transferable Visual Models from Natural Language Supervision”
Alec Radford, Jong Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin and Jack Clark · 2021
Earlier work this paper cites.
“Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm”
Yangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui, Wanli Ouyang, Jing Shao, Fengwei Yu and Junjie Yan · 2021
Earlier work this paper cites.
“Slip: Self-supervision Meets Language-image Pre-training”
Norman Mu, Alexander Kirillov, David Wagner and Saining Xie · 2021
Earlier work this paper cites.
“FILIP: Fine-grained Interactive Language-Image Pre-Training”
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang and Chunjing Xu · 2021
Earlier work this paper cites.
“FLAVA: A Foundational Language And Vision Alignment Model”
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach and Douwe Kiela · 2021
Earlier work this paper cites.
“Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision”
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc. Le, Yunhsuan Sung, Zhen Li and Tom Duerig · 2021
Earlier work this paper cites.
“Vilt: Vision-and-language Transformer without Convolution or Region Supervision”
Wonjae Kim, Bokyung Son and Ildoo Kim · 2021
Earlier work this paper cites.
“Align before Fuse: Vision and Language Representation Learning with Momentum Distillation”
Junnan Li, Ramprasaath. Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong and Steven Hoi · 2021
Earlier work this paper cites.
Junnan Li, Dongxu Li, Caiming Xiong and Steven Hoi · 2022
Earlier work this paper cites.
“CoCa: Contrastive Captioners are Image-Text Foundation Models”
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini and Yonghui Wu · 2022
Earlier work this paper cites.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar and Nando de Freitas · 2022
Earlier work this paper cites.
“Winoground: Probing vision and language models for visio-linguistic compositionality”
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela and Candace Ross · 2022
Earlier work this paper cites.
“When and why vision-language models behave like bag-of-words models, and what to do about it?”
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky and James Zou · 2022
Earlier work this paper cites.
“Scaling instruction-finetuned language models”
Hyung Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani and Siddhartha Brahma · 2022
Earlier work this paper cites.
“Flamingo: a visual language model for few-shot learning”
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican and Malcolm Reynolds · 2022
Earlier work this paper cites.
“Cyclip: Cyclic contrastive language-image pretraining”
Shashank Goel, Hritik Bansal, Sumit Bhatia, Ryan Rossi, Vishwa Vinay and Aditya Grover · 2022
Earlier work this paper cites.
“Vision-Language Pre-Training with Triple Contrastive Learning”
Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi and Junzhou Huang · 2022
Earlier work this paper cites.
“Learning to Prompt for Vision-Language Models”
Kaiyang Zhou, Jingkang Yang, Chen Loy and Ziwei Liu · 2022
Cited alongside, same era.
“Conditional Prompt Learning for Vision-Language Models”
Kaiyang Zhou, Jingkang Yang, Chen Loy and Ziwei Liu · 2022
Cited alongside, same era.
“Demystifying prompts in language models via perplexity estimation”
Hila Gonen, Srini Iyer, Terra Blevins, Noah Smith and Luke Zettlemoyer · 2022
Cited alongside, same era.
“Minigpt-4: Enhancing vision-language understanding with advanced large language models”
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li and Mohamed Elhoseiny · 2023
Cited alongside, same era.
“Improved Baselines with Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Cited alongside, same era.
“Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality”, 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph. Gonzalez, Ion Stoica and Eric. Xing · 2023
Later among the works it cites.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava and Shruti Bhosale · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Later among the works it cites.
“Improved baselines with visual instruction tuning”
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Later among the works it cites.
“Visionllm: Large language model is also an open-ended decoder for vision-centric tasks”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning”
Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechu Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong and Mohamed Elhoseiny · 2023
Cited alongside, same era.
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Cited alongside, same era.
“VLC-BERT: visual question answering with contextualized commonsense knowledge”
Sahithya Ravi, Aditya Chinchure, Leonid Sigal, Renjie Liao and Vered Shwartz · 2023
Cited alongside, same era.
“CREPE: Can Vision-Language Foundation Models Reason Compositionally?”
Zixian Ma, Jerry Hong, Mustafa Gul, Mona Gandhi, Irena Gao and Ranjay Krishna · 2023
Cited alongside, same era.
“SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality”
Cheng-Yu Hsieh, Jieyu Zhang, Zixian Ma, Aniruddha Kembhavi and Ranjay Krishna · 2023
Cited alongside, same era.
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Cited alongside, same era.
“The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)”, 2023
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu and Lijuan Wang · 2023
Cited alongside, same era.
Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou and Yu Qiao · 2023
Later among the works it cites.
“Kosmos-2: Grounding Multimodal Large Language Models to the World”
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma and Furu Wei · 2023
Later among the works it cites.
“Shikra: Unleashing Multimodal LLM’s Referential Dialogue Magic”
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu and Rui Zhao · 2023
Later among the works it cites.
“Qwen-vl: A frontier large vision-language model with versatile abilities”
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou and Jingren Zhou · 2023
Later among the works it cites.
“COLA: How to adapt vision-language models to Compose Objects Localized with Attributes?”
Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan Plummer, Ranjay Krishna and Kate Saenko · 2023
Later among the works it cites.
“Seed-bench: Benchmarking multimodal llms with generative comprehension”
Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge and Ying Shan · 2023
Later among the works it cites.
“Seed-bench-2: Benchmarking multimodal large language models”
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang and Ying Shan · 2023
Later among the works it cites.
“MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models”
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu and Rongrong Ji · 2023
Later among the works it cites.
“Image Captioners Are Scalable Vision Learners Too”
Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai, Neil Houlsby and Lucas Beyer · 2023
Later among the works it cites.
“OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents”, 2023
Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander. Rush, Douwe Kiela, Matthieu Cord and Victor Sanh · 2023
Later among the works it cites.
“Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs”
M Mirza, Leonid Karlinsky, Wei Lin, Sivan Doveh, Jakub Micorek, Mateusz Kozinski, Hilde Kuhene and Horst Possegger · 2024
Closest in time.
“Towards Multimodal In-Context Learning for Vision & Language Models”
Sivan Doveh, Shaked Perek, M Mirza, Amit Alfassy, Assaf Arbelle, Shimon Ullman and Leonid Karlinsky · 2024
Closest in time.
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2024
Closest in time.
“LLaVA-NeXT: Improved reasoning, OCR, and world knowledge”, 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen and Yong Lee · 2024
Closest in time.
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang, Haodong Duan, Maosong Cao, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin and Jiaqi Wang · 2024
Closest in time.
“Evaluating Text-to-Visual Generation with Image-to-Text Generation”
Zhiqiu Lin, Deepak Pathak, Baiqi Li, Jiayao Li, Xide Xia, Graham Neubig, Pengchuan Zhang and Deva Ramanan · 2024
Closest in time.
“Llama 3 Model Card”, 2024
AI@Meta · 2024
Closest in time.