Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have shown impressive results on various multimodal tasks.
Building a test collection for complex document information processing. In ACM Int. Conf. Multimedia , Efthimis N. Efthimiadis, Susan T. Dumais, David Hawking, and Kalervo Järvelin (Eds.). 665–666
David D. Lewis, Gady Agam, Shlomo Argamon, Ophir Frieder, David A. Grossman, and Jefferson Heard. 2006 · 2006
Earlier work this paper cites.
Im2Text: Describing Images Using 1 Million Captioned Photographs. In Adv. Neural Inform. Process. Syst. , John Shawe-Taylor, Richard S. Zemel, Peter L. Bartlett, Fernando C. N. Pereira, and Kilian Q. Weinberger (Eds.). 1143–1151
Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. 2011 · 2011
Earlier work this paper cites.
ReferItGame: Referring to Objects in Photographs of Natural Scenes. In Proc. EMNLP , Alessandro Moschitti, Bo Pang, and Walter Daelemans (Eds.). 787–798
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara L. Berg. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Compositional Semantic Parsing on Semi-Structured Tables. In Proc. ACL . 1470–1480
Panupong Pasupat and Percy Liang. 2015 · 2015
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Feature Pyramid Networks for Object Detection. In IEEE Conf. Comput. Vis. Pattern Recog. 936–944
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. 2017 · 2017
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In Proc. ACL , Iryna Gurevych and Yusuke Miyao (Eds.). 2556–2565
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Aishwarya Agrawal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2019 · 2019
Earlier work this paper cites.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering. In IEEE Conf. Comput. Vis. Pattern Recog. Computer Vision Foundation / IEEE, 6700–6709
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge. In IEEE Conf. Comput. Vis. Pattern Recog. 3195–3204
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019 · 2019
Earlier work this paper cites.
OCR-VQA: Visual Question Answering by Reading Text in Images. In Proc. ICDAR . 947–952
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. 2019 · 2019
Earlier work this paper cites.
End-to-End Object Detection with Transformers. In Eur. Conf. Comput. Vis. , Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.), Vol. 12346. 213–229
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Earlier work this paper cites.
TabFact: A Large-scale Dataset for Table-based Fact Verification. In Int. Conf. Learn. Represent
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020 · 2020
Earlier work this paper cites.
Point and Ask: Incorporating Pointing into Visual Question Answering
Arjun Mani, Will Hinthorn, Nobline Yoo, and Olga Russakovsky. 2020 · 2020
Earlier work this paper cites.
TextCaps: A Dataset for Image Captioning with Reading Comprehension. In Eur. Conf. Comput. Vis. , Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.), Vol. 12347. 742–758
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh. 2020 · 2020
Earlier work this paper cites.
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts. In IEEE Conf. Comput. Vis. Pattern Recog. 3558–3568
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. 2021 · 2021
Earlier work this paper cites.
DocVQA: A Dataset for VQA on Document Images. In Proc. WACV . 2199–2208
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision. In Proc. ICML , Marina Meila and Tong Zhang (Eds.), Vol. 139. 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
VisualMRC: Machine Reading Comprehension on Document Images. In AAAI . AAAI Press, 13878–13888
Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. 2021 · 2021
Cited alongside, same era.
Improved Baselines with Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023b · 2023
Later among the works it cites.
MMBench: Is Your Multi-modal Model an All-around Player?
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. 2023a · 2023
Later among the works it cites.
Kosmos-2: Grounding Multimodal Large Language Models to the World
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023 · 2023
Later among the works it cites.
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images. In AAAI , Brian Williams, Yiling Chen, and Jennifer Neville (Eds.). 13636–13645
Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa, Itsumi Saito, and Kuniko Saito. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq R. Joty, and Enamul Hoque. 2022 · 2022
Cited alongside, same era.
InfographicVQA. In Proc. WACV . 2582–2591
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and C. V. Jawahar. 2022 · 2022
Cited alongside, same era.
A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge. In Eur. Conf. Comput. Vis. , Shai Avidan, Gabriel J. Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner (Eds.), Vol. 13668. 146–162
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022 · 2022
Cited alongside, same era.
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 · 2023
Cited alongside, same era.
Nougat: Neural Optical Understanding for Academic Documents
Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic. 2023 · 2023
Cited alongside, same era.
Shikra: Unleashing Multimodal LLM’s Referential Dialogue Magic
Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. 2023b · 2023
Cited alongside, same era.
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2023a · 2023
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023 · 2023
Cited alongside, same era.
InternLM: A Multilingual Language Model with Progressively Enhanced Capabilities
InternLM Team. 2023 · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Later among the works it cites.
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Yuhao Dan, Chenlin Zhao, Guohai Xu, Chenliang Li, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. 2023a · 2023
Later among the works it cites.
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. 2023c · 2023
Later among the works it cites.
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Anwen Hu, Haowei Liu, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. 2023d · 2023
Later among the works it cites.
Sigmoid Loss for Language Image Pre-Training. In Int. Conf. Comput. Vis. 11941–11952
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023 · 2023
Later among the works it cites.
Pan Zhang, Xiaoyi Dong, Bin Wang, Yuhang Cao, Chao Xu, Linke Ouyang, Zhiyuan Zhao, Shuangrui Ding, Songyang Zhang, Haodong Duan, Wenwei Zhang, Hang Yan, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, and Jiaqi Wang. 2023 · 2023
Later among the works it cites.
SVIT: Scaling up Visual Instruction Tuning
Bo Zhao, Boya Wu, and Tiejun Huang. 2023 · 2023
Later among the works it cites.
Abdelrahman Abdallah, Daniel Eberharter, Zoe Pfister, and Adam Jatowt. 2024 · 2024
Closest in time.
ALLaVA: Harnessing GPT4V-synthesized Data for A Lite Vision-Language Model
Guiming Hardy Chen, Shunian Chen, Ruifei Zhang, Junying Chen, Xiangbo Wu, Zhiyi Zhang, Zhihong Chen, Jianquan Li, Xiang Wan, and Benyou Wang. 2024 · 2024
Closest in time.
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai. 2024 · 2024
Closest in time.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. In Int. Conf. Learn. Represent
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2024 · 2024
Closest in time.