Fetching the paper…
Reading the bibliography…
By harnessing the capabilities of large language models (LLMs), recent large multimodal models (LMMs) have shown remarkable versatility in open-world multimodal understanding.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
K. Kafle, B. Price, S. Cohen, and C. Kanan · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Towards vqa models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Textcaps: a dataset for image captioning with reading comprehension
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
M. Mathew, D. Karatzas, and C. Jawahar · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque · 2022
Cited alongside, same era.
Infographicvqa
M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
G. Team · 2023
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, X. Song, et al · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan · 2023
Later among the works it cites.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Q. Ye, H. Xu, J. Ye, M. Yan, H. Liu, Q. Qian, J. Zhang, F. Huang, and J. Zhou · 2023
Later among the works it cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2023
Cited alongside, same era.
Mobilevlm : A fast, strong and open vision language assistant for mobile devices, 2023
X. Chu, L. Qiao, X. Lin, S. Xu, Y. Yang, Y. Hu, F. Wei, X. Zhang, B. Zhang, X. Wei, and C. Shen · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. A. Li, P. Fung, and S. C. H. Hoi · 2023
Cited alongside, same era.
Pengi: An audio language model for audio tasks
S. Deshmukh, B. Elizalde, R. Singh, and H. Wang · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, Z. Qiu, W. Lin, J. Yang, X. Zheng, et al · 2023
Cited alongside, same era.
Imagebind: One embedding space to bind them all
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra · 2023
Cited alongside, same era.
S. Gunasekar, Y. Zhang, J. Aneja, C. C. T. Mendes, A. Del Giorno, S. Gopi, M. Javaheripi, P. Kauffmann, G. de Rosa, O. Saarikivi, et al · 2023
Cited alongside, same era.
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang · 2023
Later among the works it cites.
Tinygpt-v: Efficient multimodal large language model via small backbones
Z. Yuan, Z. Li, and L. Sun · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Allava: Harnessing gpt4v-synthesized data for a lite vision-language model
G. H. Chen, S. Chen, R. Zhang, J. Chen, X. Wu, Z. Zhang, Z. Chen, J. Li, X. Wan, and B. Wang · 2024
Closest in time.
Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
P. Gao, R. Zhang, C. Liu, L. Qiu, S. Huang, W. Lin, S. Zhao, S. Geng, Z. Lin, P. Jin, et al · 2024
Closest in time.
Gemma: Open models based on gemini research and technology, 2024
G. Gemma team · 2024
Closest in time.
Efficient multimodal learning from data-centric perspective
M. He, Y. Liu, B. Wu, J. Yuan, Y. Wang, T. Huang, and B. Zhao · 2024
Closest in time.
Prismatic vlms: Investigating the design space of visually-conditioned language models
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Closest in time.
Mini-gemini: Mining the potential of multi-modality vision language models
Y. Li, Y. Zhang, C. Wang, Z. Zhong, Y. Chen, R. Chu, S. Liu, and J. Jia · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Deepseek-vl: towards real-world vision-language understanding
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, Y. Sun, et al · 2024
Closest in time.
Phi-3 technical report: A highly capable language model locally on your phone, 2024
Microsoft · 2024
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang · 2024
Closest in time.
Small language model meets with reinforced vision vocabulary
H. Wei, L. Kong, J. Chen, L. Zhao, Z. Ge, E. Yu, J. Sun, C. Han, and X. Zhang · 2024
Closest in time.
Tinyllama: An open-source small language model, 2024
P. Zhang, G. Zeng, T. Wang, and W. Lu · 2024
Closest in time.