Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) have demonstrated promising results in a variety of tasks that combine vision and language.
Cortical analysis of visual context
Moshe Bar and Elissa Aminoff. 2003 · 2003
Earlier work this paper cites.
Cortical areas involved in object, background, and object-background processing revealed with functional magnetic resonance adaptation
Joshua OS Goh, Soon Chun Siong, Denise Park, Angela Gutchess, Andy Hebrank, and Michael WL Chee. 2004 · 2004
Earlier work this paper cites.
VQA: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Microsoft COCO Captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Two distinct scene-processing networks connecting vision and memory
Christopher Baldassano, Andre Esteva, Li Fei-Fei, and Diane M Beck. 2016 · 2016
Earlier work this paper cites.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017 · 2017
Earlier work this paper cites.
Humans incorporate attention-dependent uncertainty into perceptual decisions and confidence
Rachel N Denison, William T Adler, Marisa Carrasco, and Wei Ji Ma. 2018 · 2018
Earlier work this paper cites.
Graph r-cnn for scene graph generation
Jianwei Yang, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. 2018 · 2018
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Audio-visual floorplan reconstruction
Senthil Purushwalkam, Sebastia Vicenc Amengual Gari, Vamsi Krishna Ithapu, Carl Schissler, Philip Robinson, Abhinav Gupta, and Kristen Grauman. 2021 · 2021
Earlier work this paper cites.
The meaning and structure of scenes
Melissa Le-Hoa Vo. 2021 · 2021
Earlier work this paper cites.
MMDialog: A large-scale multi-turn dialogue dataset towards multi-modal open-domain conversation
Jiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao, Yaming Yang, Chongyang Tao, Dongyan Zhao, and Qingwei Lin. 2022 · 2022
Earlier work this paper cites.
Multi-label iterated learning for image classification with label ambiguity
Sai Rajeswar, Pau Rodriguez, Soumye Singhal, David Vazquez, and Aaron Courville. 2022 · 2022
Earlier work this paper cites.
Ambiguous images with human judgments for robust visual event classification
Kate Sanders, Reno Kriz, Anqi Liu, and Benjamin Van Durme. 2022 · 2022
Cited alongside, same era.
Is one annotation enough?-a data-centric image classification benchmark for noisy and ambiguous label estimation
Lars Schmarje, Vasco Grossmann, Claudius Zelenka, Sabine Dippel, Rainer Kiko, Mariusz Oszust, Matti Pastell, Jenny Stracke, Anna Valros, Nina Volkmann, et al. 2022 · 2022
Cited alongside, same era.
Scene graph generation: A comprehensive survey
Guangming Zhu, Liang Zhang, Youliang Jiang, Yixuan Dang, Haoran Hou, Peiyi Shen, Mingtao Feng, Xia Zhao, Qiguang Miao, Syed Afaq Ali Shah, et al. 2022 · 2022
Cited alongside, same era.
OpenFlamingo: An open-source framework for training large autoregressive vision-language models
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al. 2023 · 2023
Cited alongside, same era.
MI-GAN: A simple baseline for image inpainting on mobile devices
Andranik Sargsyan, Shant Navasardyan, Xingqian Xu, and Humphrey Shi. 2023 · 2023
Later among the works it cites.
Prompting large language models with answer heuristics for knowledge-based visual question answering
Zhenwei Shao, Zhou Yu, Meng Wang, and Jun Yu. 2023 · 2023
Later among the works it cites.
Generative multimodal models are in-context learners
Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Zhengxiong Luo, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, et al. 2023 · 2023
Later among the works it cites.
Controllable image captioning via prompting
Ning Wang, Jiahao Xie, Jihao Wu, Mingbo Jia, and Linlin Li. 2023 · 2023
Later among the works it cites.
Context understanding in computer vision: A survey
Xuan Wang and Zhigang Zhu. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 · 2023
Cited alongside, same era.
RelTR: Relation transformer for scene graph generation
Yuren Cong, Michael Ying Yang, and Bodo Rosenhahn. 2023 · 2023
Cited alongside, same era.
Holistic analysis of hallucination in GPT-4V(ision): Bias and interference challenges
Chenhang Cui, Yiyang Zhou, Xinyu Yang, Shirley Wu, Linjun Zhang, James Zou, and Huaxiu Yao. 2023 · 2023
Cited alongside, same era.
InstructBLIP: Towards general-purpose vision-language odels with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023 · 2023
Cited alongside, same era.
Object detection using YOLO: Challenges, architectural successors, datasets and applications
Tausif Diwan, G Anirudh, and Jitendra V Tembhurne. 2023 · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
G Gemini Team, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023 · 2023
Cited alongside, same era.
HallusionBench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models
Tianrui Guan, Fuxiao Liu, Xiyang Wu Ruiqi Xian Zongxia Li, Xiaoyu Liu Xijun Wang, Lichang Chen Furong Huang Yaser Yacoob, and Dinesh Manocha Tianyi Zhou. 2023 · 2023
Cited alongside, same era.
Visual Programming: Compositional visual reasoning without training
Tanmay Gupta and Aniruddha Kembhavi. 2023 · 2023
Cited alongside, same era.
Hanyu Xiang, Qin Zou, Muhammad Ali Nawaz, Xianfeng Huang, Fan Zhang, and Hongkai Yu. 2023 · 2023
Later among the works it cites.
Tiny object detection with context enhancement and feature purification
Jinsheng Xiao, Haowen Guo, Jian Zhou, Tao Zhao, Qiuze Yu, Yunhua Chen, and Zhongyuan Wang. 2023 · 2023
Later among the works it cites.
mPLUG-Owl2: Revolutionizing multi-modal large language model with modality collaboration
Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Haowei Liu, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. 2023 · 2023
Later among the works it cites.
Image inpainting based on deep learning: A review
Xiaobo Zhang, Donghai Zhai, Tianrui Li, Yuxin Zhou, and Yang Lin. 2023 · 2023
Later among the works it cites.
MMICL: Empowering vision-language model with multi-modal in-context learning
Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma, Kaikai An, Liang Chen, Zixuan Liu, Sheng Wang, Wenjuan Han, and Baobao Chang. 2023 · 2023
Later among the works it cites.
Prototype-based embedding network for scene graph generation
Chaofan Zheng, Xinyu Lyu, Lianli Gao, Bo Dai, and Jingkuan Song. 2023 · 2023
Later among the works it cites.
MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Later among the works it cites.
Object detection in 20 years: A survey
Zhengxia Zou, Keyan Chen, Zhenwei Shi, Yuhong Guo, and Jieping Ye. 2023 · 2023
Later among the works it cites.