Fetching the paper…
Reading the bibliography…
The field of vision-language models (VLMs), which take images and texts as inputs and output texts, is rapidly evolving and has yet to reach consensus on several key aspects of the development pipeline, including data, architecture, and training methods.
Language models are few-shot learners
Brown, T., B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei (2020) · 1901
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick (2017) · 1997
Earlier work this paper cites.
The iam-database: An english sentence database for offline handwriting recognition
Marti, U.-V. and H. Bunke (2002, 11) · 2002
Earlier work this paper cites.
Pathvqa: 30000+ questions for medical visual question answering
He, X., Y. Zhang, L. Mou, E. Xing, and P. Xie (2020) · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei (2009) · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick (2014) · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., A. Lai, M. Hodosh, and J. Hockenmaier (2014) · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S., A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh (2015) · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and J. Ba (2015) · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Pasupat, P. and P. Liang (2015, July) · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Ren, M., R. Kiros, and R. Zemel (2015) · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi (2016) · 2016
Earlier work this paper cites.
Visual7W: Grounded Question Answering in Images
Zhu, Y., O. Groth, M. Bernstein, and L. Fei-Fei (2016) · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., T. Khot, D. Summers-Stay, D. Batra, and D. Parikh (2017) · 2017
Earlier work this paper cites.
Search-based neural structured learning for sequential question answering
Iyyer, M., W.-t. Yih, and M.-W. Chang (2017, July) · 2017
Earlier work this paper cites.
Figureqa: An annotated figure dataset for visual reasoning
Kahou, S. E., V. Michalski, A. Atkinson, Á. Kádár, A. Trischler, and Y. Bengio (2017) · 2017
Earlier work this paper cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Kembhavi, A., M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi (2017) · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Xie, S., R. Girshick, P. Dollár, Z. Tu, and K. He (2017) · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Zhong, V., C. Xiong, and R. Socher (2017) · 2017
Earlier work this paper cites.
Learning to describe differences between pairs of similar images
Jhamtani, H. et al. (2018, October-November) · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
Kafle, K., S. Cohen, B. Price, and C. Kanan (2018) · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Lau, J., S. Gayen, A. Ben Abacha, and D. Demner-Fushman (2018, 11) · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., N. Ding, S. Goodman, and R. Soricut (2018) · 2018
Earlier work this paper cites.
Tallyqa: Answering complex counting questions
Acharya, M., K. Kafle, and C. Kanan (2019) · 2019
Earlier work this paper cites.
nocaps: novel object captioning at scale
Agrawal, H., K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson (2019, October) · 2019
Earlier work this paper cites.
Scene text visual question answering
Biten, A. F., R. Tito, A. Mafla, L. Gomez, M. Rusiñol, C. Jawahar, E. Valveny, and D. Karatzas (2019) · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and C. D. Manning (2019) · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and F. Hutter (2019) · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., M. Rastegari, A. Farhadi, and R. Mottaghi (2019) · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A., S. Shekhar, A. K. Singh, and A. Chakraborty (2019) · 2019
Earlier work this paper cites.
Kvqa: Knowledge-aware visual question answering
Shah, S., A. Mishra, N. Yadati, and P. P. Talukdar (2019) · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., V. Natarjan, M. Shah, Y. Jiang, X. Chen, D. Parikh, and M. Rohrbach (2019) · 2019
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
Suhr, A., S. Zhou, A. Zhang, I. Zhang, H. Bai, and Y. Artzi (2019, July) · 2019
Earlier work this paper cites.
Raven: A dataset for relational and analogical visual reasoning
Zhang, C., F. Gao, B. Jia, Y. Zhu, and S.-C. Zhu (2019) · 2019
Earlier work this paper cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D., H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, and D. Testuggine (2020) · 2020
Earlier work this paper cites.
Plotqa: Reasoning over scientific plots
Methani, N., P. Ganguly, M. M. Khapra, and P. Kumar (2020, March) · 2020
Earlier work this paper cites.
Chart-to-text: Generating natural language descriptions for charts by adapting the transformer model
Obeid, J. and E. Hoque (2020, December) · 2020
Earlier work this paper cites.
Connecting vision and language with localized narratives
Pont-Tuset, J., J. Uijlings, S. Changpinyo, R. Soricut, and V. Ferrari (2020) · 2020
Earlier work this paper cites.
Textcaps: A dataset for image captioning with reading comprehension
Sidorov, O., R. Hu, M. Rohrbach, and A. Singh (2020) · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush (2020, October) · 2020
Earlier work this paper cites.
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., P. Sharma, N. Ding, and R. Soricut (2021) · 2021
Earlier work this paper cites.
FinQA: A dataset of numerical reasoning over financial data
Chen, Z., W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y. Wang (2021, November) · 2021
Earlier work this paper cites.
Redcaps: Web-curated image-text data created by the people, for the people
Desai, K., G. Kaul, Z. Aysola, and J. Johnson (2021) · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt (2021) · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention
Jaegle, A., F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira (2021, 18–24 Jul) · 2021
Earlier work this paper cites.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Lu, P., R. Gong, S. Jiang, L. Qiu, S. Huang, X. Liang, and S.-C. Zhu (2021) · 2021
Earlier work this paper cites.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
Lu, P., L. Qiu, J. Chen, T. Xia, Y. Zhao, W. Zhang, Z. Yu, X. Liang, and S.-C. Zhu (2021) · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Mathew, M., D. Karatzas, and C. V. Jawahar (2021) · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021) · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki (2021) · 2021
Earlier work this paper cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Srinivasan, K., K. Raman, J. Chen, M. Bendersky, and M. Najork (2021) · 2021
Earlier work this paper cites.
Visualmrc: Machine reading comprehension on document images
Tanaka, R., K. Nishida, and S. Yoshida (2021) · 2021
Earlier work this paper cites.
Multimodal few-shot learning with frozen language models
Tsimpoukelli, M., J. L. Menick, S. Cabi, S. M. A. Eslami, O. Vinyals, and F. Hill (2021) · 2021
Earlier work this paper cites.
Screen2words: Automatic mobile ui summarization with multimodal learning
Wang, B., G. Li, X. Zhou, Z. Chen, T. Grossman, and Y. Li (2021) · 2021
Earlier work this paper cites.
TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance
Zhu, F., W. Lei, Y. Huang, C. Wang, S. Zhang, J. Lv, F. Feng, and T.-S. Chua (2021, August) · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. a. Bińkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan (2022) · 2022
Earlier work this paper cites.
PromptSource: An integrated development environment and repository for natural language prompts
Bach, S., V. Sanh, Z. X. Yong, A. Webson, C. Raffel, N. V. Nayak, A. Sharma, T. Kim, M. S. Bari, T. Fevry, Z. Alyafeai, M. Dey, A. Santilli, Z. Sun, S. Ben-david, C. Xu, G. Chhablani, H. Wang, J. Fries, M. Al-shaibani, S. Sharma, U. Thakker, K. Almubarak, X. Tang, D. Radev, M. T.-j. Jiang, and A. Rush (2022, May) · 2022
Earlier work this paper cites.
Ocr-idl: Ocr annotations for industry document library dataset
Biten, A. F., R. Tito, L. Gomez, E. Valveny, and D. Karatzas (2022) · 2022
Earlier work this paper cites.
Coyo-700m: Image-text pair dataset
Byeon, M., B. Park, H. Kim, S. Lee, W. Baek, and S. Kim (2022) · 2022
Earlier work this paper cites.
MapQA: A dataset for question answering on choropleth maps
Chang, S., D. Palzer, J. Li, E. Fosler-Lussier, and N. Xiao (2022) · 2022
Earlier work this paper cites.
Pali: Scaling language-image learning in 100+ languages
Chen, X. and X. Wang (2022) · 2022
Earlier work this paper cites.
HiTab: A hierarchical table dataset for question answering and natural language generation
Cheng, Z., H. Dong, Z. Wang, R. Jia, J. Guo, Y. Gao, S. Han, J.-G. Lou, and D. Zhang (2022, May) · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) · 2022
Earlier work this paper cites.
The bigscience roots corpus: A 1.6tb composite multilingual dataset
Laurençon, H., L. Saulnier, T. Wang, C. Akiki, A. Villanova del Moral, T. Le Scao, L. Von Werra, C. Mou, E. González Ponferrada, H. Nguyen, J. Frohberg, M. Šaško, Q. Lhoest, A. McMillan-Major, G. Dupont, S. Biderman, A. Rogers, L. Ben allal, F. De Toni, G. Pistilli, O. Nguyen, S. Nikpoor, M. Masoud, P. Colombo, J. de la Rosa, P. Villegas, T. Thrush, S. Longpre, S. Nagel, L. Weber, M. Muñoz, J. Zhu, D. Van Strien, Z. Alyafeai, K. Almubarak, M. C. Vu, I. Gonzalez-Dios, A. Soroa, K. Lo, M. Dey, P. Ortiz Suarez, A. Gokaslan, S. Bose, D. Adelani, L. Phan, H. Tran, I. Yu, S. Pai, J. Chim, V. Lepercq, S. Ilic, M. Mitchell, S. A. Luccioni, and Y. Jernite (2022) · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., D. Li, C. Xiong, and S. Hoi (2022) · 2022
Earlier work this paper cites.
Clevr-math: A dataset for compositional language, visual and mathematical reasoning
Lindström, A. D. and al (2022) · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan (2022) · 2022
Earlier work this paper cites.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., D. Long, J. Q. Tan, S. Joty, and E. Hoque (2022, May) · 2022
Earlier work this paper cites.
Infographicvqa
Mathew, M., V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. V. Jawahar (2022) · 2022
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., A. Webson, C. Raffel, S. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. Nayak, D. Datta, J. Chang, M. T.-J. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Fevry, J. A. Fries, R. Teehan, T. L. Scao, S. Biderman, L. Gao, T. Wolf, and A. M. Rush (2022) · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev (2022) · 2022
Earlier work this paper cites.
Laion coco: 600m synthetic captions from laion2b-en
Schuhmann, C., A. Köpf, R. Vencu, T. Coombes, and R. Beaumont (2022) · 2022
Cited alongside, same era.
A-okvqa: A benchmark for visual question answering using world knowledge
Schwenk, D., A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi (2022) · 2022
Cited alongside, same era.
Flava: A foundational language and vision alignment model
Singh, A., R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela (2022) · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le (2022) · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al. (2022) · 2022
Cited alongside, same era.
MultiHiertt: Numerical reasoning over multi hierarchical tabular and textual data
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M., S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behl, et al. (2024) · 2024
Closest in time.
Visfocus: Prompt-guided vision encoders for ocr-free dense document understanding
Abramovich, O., N. Nayman, S. Fogel, I. Lavi, R. Litman, S. Tsiper, R. Tichauer, S. Appalaraju, S. Mazor, and R. Manmatha (2024) · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic, A. (2024) · 2024
Closest in time.
Mint-1t: Scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens
Awadalla, A., L. Xue, O. Lo, M. Shu, H. Lee, E. K. Guha, M. Jordan, S. Shen, M. Awadalla, S. Savarese, et al. (2024) · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhao, Y., Y. Li, C. Li, and R. Zhang (2022, May) · 2022
Cited alongside, same era.
Semdedup: Data-efficient learning at web-scale through semantic deduplication
Abbas, A., K. Tirumala, D. Simig, S. Ganguli, and A. S. Morcos (2023) · 2023
Cited alongside, same era.
Achiam, J., S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. (2023) · 2023
Cited alongside, same era.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Awadalla, A., I. Gao, J. Gardner, J. Hessel, Y. Hanafy, W. Zhu, K. Marathe, Y. Bitton, S. Gadre, S. Sagawa, et al. (2023) · 2023
Cited alongside, same era.
Bai, J., S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al. (2023) · 2023
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Bai, J., S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou (2023) · 2023
Cited alongside, same era.
Introducing our multimodal models
Bavishi, R., E. Elsen, C. Hawthorne, M. Nye, A. Odena, A. Somani, and S. Taşırlar (2023) · 2023
Cited alongside, same era.
Beyer, L., A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, et al. (2024) · 2024
Closest in time.
Cai, Z., M. Cao, H. Chen, K. Chen, K. Chen, X. Chen, X. Chen, Z. Chen, Z. Chen, P. Chu, et al. (2024) · 2024
Closest in time.
Chart-based reasoning: Transferring capabilities from llms to vlms
Carbune, V., H. Mansoor, F. Liu, R. Aralikatte, G. Baechler, J. Chen, and A. Sharma (2024) · 2024
Closest in time.
Honeybee: Locality-enhanced projector for multimodal llm
Cha, J., W. Kang, J. Mun, and B. Roh (2024) · 2024
Closest in time.
Are we on the right way for evaluating large vision-language models?
Chen, L., J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, J. Wang, Y. Qiao, D. Lin, et al. (2024) · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Chen, Z., W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma, et al. (2024) · 2024
Closest in time.
Mobilevlm v2: Faster and stronger baseline for vision language model
Chu, X., L. Qiao, X. Zhang, S. Xu, F. Wei, Y. Yang, X. Sun, Y. Hu, X. Lin, B. Zhang, et al. (2024) · 2024
Closest in time.
Vision transformers need registers
Darcet, T., M. Oquab, J. Mairal, and P. Bojanowski (2024) · 2024
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
DeepSeek-AI, A. Liu, B. Feng, B. Wang, B. Wang, B. Liu, C. Zhao, C. Dengr, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Yang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Chen, J. Yuan, J. Qiu, J. Song, K. Dong, K. Gao, K. Guan, L. Wang, L. Zhang, L. Xu, L. Xia, L. Zhao, L. Zhang, M. Li, M. Wang, M. Zhang, M. Zhang, M. Tang, M. Li, N. Tian, P. Huang, P. Wang, P. Zhang, Q. Zhu, Q. Chen, Q. Du, R. J. Chen, R. L. Jin, R. Ge, R. Pan, R. Xu, R. Chen, S. S. Li, S. Lu, S. Zhou, S. Chen, S. Wu, S. Ye, S. Ma, S. Wang, S. Zhou, S. Yu, S. Zhou, S. Zheng, T. Wang, T. Pei, T. Yuan, T. Sun, W. L. Xiao, W. Zeng, W. An, W. Liu, W. Liang, W. Gao, W. Zhang, X. Q. Li, X. Jin, X. Wang, X. Bi, X. Liu, X. Wang, X. Shen, X. Chen, X. Chen, X. Nie, X. Sun, X. Wang, X. Liu, X. Xie, X. Yu, X. Song, X. Zhou, X. Yang, X. Lu, X. Su, Y. Wu, Y. K. Li, Y. X. Wei, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zhao, Y. Sun, Y. Li, Y. Wang, Y. Zheng, Y. Zhang, Y. Xiong, Y. Zhao, Y. He, Y. Tang, Y. Piao, Y. Dong, Y. Tan, Y. Liu, Y. Wang, Y. Guo, Y. Zhu, Y. Wang, Y. Zou, Y. Zha, Y. Ma, Y. Yan, Y. You, Y. Liu, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Huang, Z. Zhang, Z. Xie, Z. Hao, Z. Shao, Z. Wen, Z. Xu, Z. Zhang, Z. Li, Z. Wang, Z. Gu, Z. Li, and Z. Xie (2024) · 2024
Closest in time.
Dong, X., P. Zhang, Y. Zang, Y. Cao, B. Wang, L. Ouyang, S. Zhang, H. Duan, W. Zhang, Y. Li, et al. (2024) · 2024
Closest in time.
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Duan, H., J. Yang, Y. Qiao, X. Fang, L. Chen, Y. Liu, X. Dong, Y. Zang, P. Zhang, J. Wang, et al. (2024) · 2024
Closest in time.
Dubey, A., A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al. (2024) · 2024
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
Gadre, S. Y., G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al. (2024) · 2024
Closest in time.
Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
Gao, P., R. Zhang, C. Liu, L. Qiu, S. Huang, W. Lin, S. Zhao, S. Geng, Z. Lin, P. Jin, et al. (2024) · 2024
Closest in time.
Imageinwords: Unlocking hyper-detailed image descriptions
Garg, R., A. Burns, B. K. Ayan, Y. Bitton, C. Montgomery, Y. Onoe, A. Bunner, R. Krishna, J. Baldridge, and R. Soricut (2024) · 2024
Closest in time.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
Guo, D., Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. Li, et al. (2024) · 2024
Closest in time.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Hu, A., H. Xu, J. Ye, M. Yan, L. Zhang, B. Zhang, C. Li, J. Zhang, Q. Jin, F. Huang, et al. (2024) · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Hu, S., Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, et al. (2024) · 2024
Closest in time.
NEFTune: Noisy embeddings improve instruction finetuning
Jain, N., P. yeh Chiang, Y. Wen, J. Kirchenbauer, H.-M. Chu, G. Somepalli, B. R. Bartoldson, B. Kailkhura, A. Schwarzschild, A. Saha, M. Goldblum, J. Geiping, and T. Goldstein (2024) · 2024
Closest in time.
Mantis: Interleaved multi-image instruction tuning
Jiang, D., X. He, H. Zeng, C. Wei, M. Ku, Q. Liu, and W. Chen (2024) · 2024
Closest in time.
Prismatic vlms: Investigating the design space of visually-conditioned language models
Karamcheti, S., S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh (2024) · 2024
Closest in time.
Geomverse: A systematic evaluation of large models for geometric reasoning
Kazemi, M., H. Alvari, A. Anand, J. Wu, X. Chen, and R. Soricut (2024) · 2024
Closest in time.
What matters when building vision-language models?
Laurençon, H., L. Tronchon, M. Cord, and V. Sanh (2024) · 2024
Closest in time.
Unlocking the conversion of web screenshots into html code with the websight dataset
Laurençon, H., L. Tronchon, and V. Sanh (2024) · 2024
Closest in time.
Meteor: Mamba-based traversal of rationale for large language and vision models
Lee, B.-K., C. W. Kim, B. Park, and Y. M. Ro (2024) · 2024
Closest in time.
Moai: Mixture of all intelligence for large language and vision models
Lee, B.-K., B. Park, C. W. Kim, and Y. M. Ro (2024) · 2024
Closest in time.
Llava-onevision: Easy visual task transfer
Li, B., Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, Y. Li, Z. Liu, and C. Li (2024) · 2024
Closest in time.
Omnicorpus: An unified multimodal corpus of 10 billion-level images interleaved with text
Li, Q., Z. Chen, W. Wang, W. Wang, S. Ye, Z. Jin, G. Chen, Y. He, Z. Gao, E. Cui, et al. (2024) · 2024
Closest in time.
Mini-gemini: Mining the potential of multi-modality vision language models
Li, Y., Y. Zhang, C. Wang, Z. Zhong, Y. Chen, R. Chu, S. Liu, and J. Jia (2024) · 2024
Closest in time.
Moe-llava: Mixture of experts for large vision-language models
Lin, B., Z. Tang, Y. Ye, J. Cui, B. Zhu, P. Jin, J. Zhang, M. Ning, and L. Yuan (2024) · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge
Liu, H., C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee (2024, January) · 2024
Closest in time.
Infimm-hd: A leap forward in high-resolution multimodal understanding
Liu, H., Q. You, X. Han, Y. Wang, B. Zhai, Y. Liu, Y. Tao, H. Huang, R. He, and H. Yang (2024) · 2024
Closest in time.
Best practices and lessons learned on synthetic data for language models
Liu, R., J. Wei, F. Liu, C. Si, Y. Zhang, J. Rao, S. Zheng, D. Peng, D. Yang, D. Zhou, and A. M. Dai (2024) · 2024
Closest in time.
Dora: Weight-decomposed low-rank adaptation
Liu, S.-Y., C.-Y. Wang, H. Yin, P. Molchanov, Y.-C. F. Wang, K.-T. Cheng, and M.-H. Chen (2024) · 2024
Closest in time.
Starcoder 2 and the stack v2: The next generation
Lozhkov, A., R. Li, L. B. Allal, F. Cassano, J. Lamy-Poirier, N. Tazi, A. Tang, D. Pykhtar, J. Liu, Y. Wei, et al. (2024) · 2024
Closest in time.
Deepseek-vl: towards real-world vision-language understanding
Lu, H., W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, Y. Sun, et al. (2024) · 2024
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao (2024) · 2024
Closest in time.
Improving automatic vqa evaluation using large language models
Mañas, O., B. Krojer, and A. Agrawal (2024) · 2024
Closest in time.
Chartgemma: Visual instruction-tuning for chart reasoning in the wild
Masry, A., M. Thakkar, A. Bajaj, A. Kartha, E. Hoque, and S. Joty (2024) · 2024
Closest in time.
Mm1: Methods, analysis & insights from multimodal llm pre-training
McKinzie, B., Z. Gan, J.-P. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, F. Weers, et al. (2024) · 2024
Closest in time.
Openelm: An efficient language model family with open-source training and inference framework
Mehta, S., M. H. Sekhavat, Q. Cao, M. Horton, Y. Jin, C. Sun, I. Mirzadeh, M. Najibi, D. Belenko, P. Zatloukal, et al. (2024) · 2024
Closest in time.
Orca-math: Unlocking the potential of slms in grade school math
Mitra, A., H. Khanpour, C. Rosset, and A. Awadallah (2024) · 2024
Closest in time.
The fineweb datasets: Decanting the web for the finest text data at scale
Penedo, G., H. Kydlíček, A. Lozhkov, M. Mitchell, C. Raffel, L. Von Werra, T. Wolf, et al. (2024) · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn (2024) · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al. (2024) · 2024
Closest in time.
Synth2: Boosting visual-language models with synthetic captions and image embeddings
Sharifzadeh, S., C. Kaplanis, S. Pathak, D. Kumaran, A. Ilic, J. Mitrovic, C. Blundell, and A. Banino (2024) · 2024
Closest in time.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, E., Y. Dong, and N. Abu-Ghazaleh (2024) · 2024
Closest in time.
From pixels to prose: A large dataset of dense image captions
Singla, V., K. Yue, S. Paul, R. Shirkavand, M. Jayawardhana, A. Ganjdanesh, H. Huang, A. Bhatele, G. Somepalli, and T. Goldstein (2024) · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Team, G., T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al. (2024) · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Y. Wu, Q. V. Le, H. He, and T. Luong (2024) · 2024
Closest in time.
Improved baselines for data-efficient perceptual augmentation of llms
Vallaeys, T., M. Shukor, M. Cord, and J. Verbeek (2024) · 2024
Closest in time.
The evolution of multimodal model architectures
Wadekar, S. N., A. Chaurasia, A. Chadha, and E. Culurciello (2024) · 2024
Closest in time.
Florence-2: Advancing a unified representation for a variety of vision tasks
Xiao, B., H. Wu, W. Xu, X. Dai, H. Hu, Y. Lu, M. Zeng, C. Liu, and L. Yuan (2024) · 2024
Closest in time.
Palm2-vadapter: Progressively aligned language model makes a strong vision-language adapter
Xiao, J., Z. Xu, A. Yuille, S. Yan, and B. Wang (2024) · 2024
Closest in time.
xgen-mm (blip-3): A family of open large multimodal models
Xue, L., M. Shu, A. Awadalla, J. Wang, A. Yan, S. Purushwalkam, H. Zhou, V. Prabhu, Y. Dai, M. S. Ryoo, S. Kendre, J. Zhang, C. Qin, S. Zhang, C.-C. Chen, N. Yu, J. Tan, T. M. Awalgaonkar, S. Heinecke, H. Wang, Y. Choi, L. Schmidt, Z. Chen, S. Savarese, J. C. Niebles, C. Xiong, and R. Xu (2024) · 2024
Closest in time.
Minicpm-v: A gpt-4v level mllm on your phone
Yao, Y., T. Yu, A. Zhang, C. Wang, J. Cui, H. Zhu, T. Cai, H. Li, W. Zhao, Z. He, Q. Chen, H. Zhou, Z. Zou, H. Zhang, S. Hu, Z. Zheng, J. Zhou, J. Cai, X. Han, G. Zeng, D. Li, Z. Liu, and M. Sun (2024) · 2024
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., W. Jiang, H. Shi, J. YU, Z. Liu, Y. Zhang, J. Kwok, Z. Li, A. Weller, and W. Liu (2024) · 2024
Closest in time.
Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Yu, T., Y. Yao, H. Zhang, T. He, Y. Han, G. Cui, J. Hu, Z. Liu, H.-T. Zheng, M. Sun, et al. (2024) · 2024
Closest in time.
Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness
Yu, T., H. Zhang, Y. Yao, Y. Dang, D. Chen, X. Lu, G. Cui, T. He, Z. Liu, T.-S. Chua, et al. (2024) · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen (2024) · 2024
Closest in time.
MAmmoTH: Building math generalist models through hybrid instruction tuning
Yue, X., X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen (2024) · 2024
Closest in time.
Zhang, P., X. Dong, Y. Zang, Y. Cao, R. Qian, L. Chen, Q. Guo, H. Duan, B. Wang, L. Ouyang, et al. (2024) · 2024
Closest in time.
Mavis: Mathematical visual instruction tuning
Zhang, R., X. Wei, D. Jiang, Y. Zhang, Z. Guo, C. Tong, J. Liu, A. Zhou, B. Wei, S. Zhang, et al. (2024) · 2024
Closest in time.
Spa-vl: A comprehensive safety preference alignment dataset for vision language model
Zhang, Y., L. Chen, G. Zheng, Y. Gao, R. Zheng, J. Fu, Z. Yin, S. Jin, Y. Qiao, X. Huang, et al. (2024) · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. (2024) · 2024
Closest in time.
Tinyllava: A framework of small-scale large multimodal models
Zhou, B., Y. Hu, X. Weng, J. Jia, J. Luo, X. Liu, J. Wu, and L. Huang (2024) · 2024
Closest in time.