Fetching the paper…
Reading the bibliography…
Comparing two images in terms of Commonalities and Differences (CaD) is a fundamental human capability that forms the basis of advanced visual reasoning and interpretation.
“The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale”
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci and Alexander Kolesnikov · 1981
Earlier work this paper cites.
“Microsoft coco: Common objects in context”
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár and C Zitnick · 2014
Earlier work this paper cites.
“From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions”
Peter Young, Alice Lai, Micah Hodosh and Julia Hockenmaier · 2014
Earlier work this paper cites.
“The language of actions: Recovering the syntax and semantics of goal-directed human activities”
Hilde Kuehne, Ali Arslan and Thomas Serre · 2014
Earlier work this paper cites.
“Microsoft coco captions: Data collection and evaluation server”
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár and C Zitnick · 2015
Earlier work this paper cites.
“What are the Gestalt Principles?”, 2016
Interaction- IxDF · 2016
Earlier work this paper cites.
“Making the v in vqa matter: Elevating the role of image understanding in visual question answering”
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra and Devi Parikh · 2017
Earlier work this paper cites.
“Visual genome: Connecting language and vision using crowdsourced dense image annotations”
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li and David Shamma · 2017
Earlier work this paper cites.
“The" something something" video database for learning and evaluating visual common sense”
Raghav Goyal, Samira Ebrahimi, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos and Moritz Mueller-Freitag · 2017
Earlier work this paper cites.
“Learning to Describe Differences Between Pairs of Similar Images”
Harsh Jhamtani and Taylor Berg-Kirkpatrick · 2018
Earlier work this paper cites.
“Learning to Describe Differences Between Pairs of Similar Images”
Harsh Jhamtani and Taylor Berg-Kirkpatrick · 2018
Earlier work this paper cites.
“Towards automatic learning of procedures from web instructional videos”
Luowei Zhou, Chenliang Xu and Jason Corso · 2018
Earlier work this paper cites.
“Evaluating text-to-image matching using binary image selection (bison)”
Hexiang Hu, Ishan Misra and Laurens Van · 2019
Earlier work this paper cites.
“A Corpus for Reasoning about Natural Language Grounded in Photographs”
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai and Yoav Artzi · 2019
Earlier work this paper cites.
“Gqa: A new dataset for real-world visual reasoning and compositional question answering”
Drew Hudson and Christopher Manning · 2019
Earlier work this paper cites.
“Semantic understanding of scenes through the ade20k dataset”
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso and Antonio Torralba · 2019
Earlier work this paper cites.
“Connecting vision and language with localized narratives”
Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut and Vittorio Ferrari · 2020
Earlier work this paper cites.
“Action genome: Actions as compositions of spatio-temporal scene graphs”
Jingwei Ji, Ranjay Krishna, Li Fei-Fei and Juan Niebles · 2020
Earlier work this paper cites.
“Probing Image-Language Transformers for Verb Understanding”
Lisa Hendricks and Aida Nematzadeh · 2021
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision”
Alec Radford, Jong Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin and Jack Clark · 2021
Earlier work this paper cites.
“Lora: Low-rank adaptation of large language models”
Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang and Weizhu Chen · 2021
Earlier work this paper cites.
“On the Opportunities and Risks of Foundation Models”, 2022
Rishi Bommasani, Drew. Hudson, Ehsan Adeli, Percy Liang and et al · 2022
Earlier work this paper cites.
“Geb+: A benchmark for generic event boundary captioning, grounding and retrieval”
Yuxuan Wang, Difei Gao, Licheng Yu, Weixian Lei, Matt Feiszli and Mike Shou · 2022
Cited alongside, same era.
“Kubric: A scalable dataset generator”
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David Fleet, Dan Gnanapragasam, Florian Golemo and Charles Herrmann · 2022
Cited alongside, same era.
“High-resolution image synthesis with latent diffusion models”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer · 2022
Cited alongside, same era.
“Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100”
Dima Damen, Hazel Doughty, Giovanni Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett and Will Price · 2022
Cited alongside, same era.
“Visual instruction tuning”
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Cited alongside, same era.
“M3IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning”
Lei Li, Yuwei Yin, Shicheng Li, Liang Chen, Peiyi Wang, Shuhuai Ren, Mukai Li, Yazheng Yang, Jingjing Xu and Xu Sun · 2023
Later among the works it cites.
“Llavar: Enhanced visual instruction tuning for text-rich image understanding”
Yanzhe Zhang, Ruiyi Zhang, Jiuxiang Gu, Yufan Zhou, Nedim Lipka, Diyi Yang and Tong Sun · 2023
Later among the works it cites.
“Svit: Scaling up visual instruction tuning”
Bo Zhao, Boya Wu and Tiejun Huang · 2023
Later among the works it cites.
“Judging llm-as-a-judge with mt-bench and chatbot arena”
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li and Eric Xing · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu and Lijuan Wang · 2023
Cited alongside, same era.
“Sparkles: Unlocking chats across multiple images for multimodal instruction-following models”
Yupan Huang, Zaiqiao Meng, Fangyu Liu, Yixuan Su, Nigel Collier and Yutong Lu · 2023
Cited alongside, same era.
“Otter: A Multi-Modal Model with In-Context Instruction Tuning”
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang and Ziwei Liu · 2023
Cited alongside, same era.
“Generative Multimodal Models are In-Context Learners”
Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Zhengxiong Luo, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang and Xinlong Wang · 2023
Cited alongside, same era.
“MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models”
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li and Mohamed Elhoseiny · 2023
Cited alongside, same era.
“Openflamingo: An open-source framework for training large autoregressive vision-language models”
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre and Shiori Sagawa · 2023
Cited alongside, same era.
“Mimic-it: Multi-modal in-context instruction tuning”
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li and Ziwei Liu · 2023
Cited alongside, same era.
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava and Shruti Bhosale · 2023
Later among the works it cites.
“Seed-bench: Benchmarking multimodal llms with generative comprehension”
Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge and Ying Shan · 2023
Later among the works it cites.
“Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality”
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang and Joseph Gonzalez · 2023
Later among the works it cites.
“Eva: Exploring the limits of masked visual representation learning at scale”
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang and Yue Cao · 2023
Later among the works it cites.
“Eva-clip: Improved training techniques for clip at scale”
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang and Yue Cao · 2023
Later among the works it cites.
“Llama: Open and efficient foundation language models”
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro and Faisal Azhar · 2023
Later among the works it cites.
“Internlm: A multilingual language model with progressively enhanced capabilities”
InternLM Team · 2023
Later among the works it cites.
“Seed-bench-2: Benchmarking multimodal large language models”
Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang and Ying Shan · 2023
Later among the works it cites.
“GPT-4 Technical Report”, 2024
OpenAI et al · 2024
Closest in time.
“Gemini: A Family of Highly Capable Multimodal Models”, 2024
Gemini et al · 2024
Closest in time.
“The Claude 3 Model Family: Opus, Sonnet, Haiku”, 2024
Anthropic · 2024
Closest in time.
“Llama 3 Model Card”, 2024
AI@Meta · 2024
Closest in time.
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang, Haodong Duan, Maosong Cao, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin and Jiaqi Wang · 2024
Closest in time.
“LLaVA-NeXT: Improved reasoning, OCR, and world knowledge”, 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen and Yong Lee · 2024
Closest in time.
“Generative multimodal models are in-context learners”
Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Zhengxiong Luo, Yueze Wang, Yongming Rao, Jingjing Liu and Tiejun Huang · 2024
Closest in time.
“MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning”
Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma, Kaikai An, Liang Chen, Zixuan Liu, Sheng Wang, Wenjuan Han and Baobao Chang · 2024
Closest in time.
Albert Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Chaplot, Diego Casas, Emma Hanna and Florian Bressand · 2024
Closest in time.