Fetching the paper…
Reading the bibliography…
The growing interest in vision-language models (VLMs) has been driven by improvements in large language models and vision transformers.
Language models are few-shot learners
Brown, T., B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei (2020) · 1901
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick (2017) · 1997
Earlier work this paper cites.
The iam-database: An english sentence database for offline handwriting recognition
Marti, U.-V. and H. Bunke (2002, 11) · 2002
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei (2009) · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick (2014) · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., A. Lai, M. Hodosh, and J. Hockenmaier (2014) · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S., A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh (2015) · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and J. Ba (2015) · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Pasupat, P. and P. Liang (2015, July) · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Ren, M., R. Kiros, and R. Zemel (2015) · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi (2016) · 2016
Earlier work this paper cites.
Visual7W: Grounded Question Answering in Images
Zhu, Y., O. Groth, M. Bernstein, and L. Fei-Fei (2016) · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., T. Khot, D. Summers-Stay, D. Batra, and D. Parikh (2017) · 2017
Earlier work this paper cites.
Search-based neural structured learning for sequential question answering
Iyyer, M., W.-t. Yih, and M.-W. Chang (2017, July) · 2017
Earlier work this paper cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Kembhavi, A., M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi (2017) · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Zhong, V., C. Xiong, and R. Socher (2017) · 2017
Earlier work this paper cites.
Learning to describe differences between pairs of similar images
Jhamtani, H. et al. (2018, October-November) · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
Kafle, K., S. Cohen, B. Price, and C. Kanan (2018) · 2018
Earlier work this paper cites.
Figureqa: An annotated figure dataset for visual reasoning
Kahou, S. E., V. Michalski, A. Atkinson, A. Kadar, A. Trischler, and Y. Bengio (2018) · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Lau, J., S. Gayen, A. Ben Abacha, and D. Demner-Fushman (2018, 11) · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., N. Ding, S. Goodman, and R. Soricut (2018) · 2018
Earlier work this paper cites.
Tallyqa: Answering complex counting questions
Acharya, M., K. Kafle, and C. Kanan (2019) · 2019
Earlier work this paper cites.
nocaps: novel object captioning at scale
Agrawal, H., K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson (2019, October) · 2019
Earlier work this paper cites.
Scene text visual question answering
Biten, A. F., R. Tito, A. Mafla, L. Gomez, M. Rusiñol, C. Jawahar, E. Valveny, and D. Karatzas (2019) · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and C. D. Manning (2019) · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., M. Rastegari, A. Farhadi, and R. Mottaghi (2019) · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A., S. Shekhar, A. K. Singh, and A. Chakraborty (2019) · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., V. Natarjan, M. Shah, Y. Jiang, X. Chen, D. Parikh, and M. Rohrbach (2019) · 2019
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
Suhr, A., S. Zhou, A. Zhang, I. Zhang, H. Bai, and Y. Artzi (2019, July) · 2019
Earlier work this paper cites.
Raven: A dataset for relational and analogical visual reasoning
Zhang, C., F. Gao, B. Jia, Y. Zhu, and S.-C. Zhu (2019) · 2019
Earlier work this paper cites.
Pathvqa: 30000+ questions for medical visual question answering
He, X., Y. Zhang, L. Mou, E. Xing, and P. Xie (2020) · 2020
Earlier work this paper cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D., H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, and D. Testuggine (2020) · 2020
Earlier work this paper cites.
Plotqa: Reasoning over scientific plots
Methani, N., P. Ganguly, M. M. Khapra, and P. Kumar (2020, March) · 2020
Earlier work this paper cites.
Chart-to-text: Generating natural language descriptions for charts by adapting the transformer model
Obeid, J. and E. Hoque (2020, December) · 2020
Earlier work this paper cites.
Connecting vision and language with localized narratives
Pont-Tuset, J., J. Uijlings, S. Changpinyo, R. Soricut, and V. Ferrari (2020) · 2020
Earlier work this paper cites.
Textcaps: A dataset for image captioning with reading comprehension
Sidorov, O., R. Hu, M. Rohrbach, and A. Singh (2020) · 2020
Earlier work this paper cites.
Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., P. Sharma, N. Ding, and R. Soricut (2021) · 2021
Earlier work this paper cites.
FinQA: A dataset of numerical reasoning over financial data
Chen, Z., W. Chen, C. Smiley, S. Shah, I. Borova, D. Langdon, R. Moussa, M. Beane, T.-H. Huang, B. Routledge, and W. Y. Wang (2021, November) · 2021
Earlier work this paper cites.
Redcaps: Web-curated image-text data created by the people, for the people
Desai, K., G. Kaul, Z. Aysola, and J. Johnson (2021) · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt (2021) · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention
Jaegle, A., F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira (2021, 18–24 Jul) · 2021
Earlier work this paper cites.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Lu, P., R. Gong, S. Jiang, L. Qiu, S. Huang, X. Liang, and S.-C. Zhu (2021) · 2021
Earlier work this paper cites.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
Lu, P., L. Qiu, J. Chen, T. Xia, Y. Zhao, W. Zhang, Z. Yu, X. Liang, and S.-C. Zhu (2021) · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Mathew, M., D. Karatzas, and C. V. Jawahar (2021) · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021) · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki (2021) · 2021
Earlier work this paper cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
Srinivasan, K., K. Raman, J. Chen, M. Bendersky, and M. Najork (2021) · 2021
Earlier work this paper cites.
Visualmrc: Machine reading comprehension on document images
Tanaka, R., K. Nishida, and S. Yoshida (2021) · 2021
Earlier work this paper cites.
Screen2words: Automatic mobile ui summarization with multimodal learning
Wang, B., G. Li, X. Zhou, Z. Chen, T. Grossman, and Y. Li (2021) · 2021
Earlier work this paper cites.
TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance
Zhu, F., W. Lei, Y. Huang, C. Wang, S. Zhang, J. Lv, F. Feng, and T.-S. Chua (2021, August) · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. a. Bińkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan (2022) · 2022
Cited alongside, same era.
PromptSource: An integrated development environment and repository for natural language prompts
Bach, S., V. Sanh, Z. X. Yong, A. Webson, C. Raffel, N. V. Nayak, A. Sharma, T. Kim, M. S. Bari, T. Fevry, Z. Alyafeai, M. Dey, A. Santilli, Z. Sun, S. Ben-david, C. Xu, G. Chhablani, H. Wang, J. Fries, M. Al-shaibani, S. Sharma, U. Thakker, K. Almubarak, X. Tang, D. Radev, M. T.-j. Jiang, and A. Rush (2022, May) · 2022
Cited alongside, same era.
Ocr-idl: Ocr annotations for industry document library dataset
Biten, A. F., R. Tito, L. Gomez, E. Valveny, and D. Karatzas (2022) · 2022
Cited alongside, same era.
MapQA: A dataset for question answering on choropleth maps
Chang, S., D. Palzer, J. Li, E. Fosler-Lussier, and N. Xiao (2022) · 2022
Cited alongside, same era.
Pali: Scaling language-image learning in 100+ languages
Chen, X. and X. Wang (2022) · 2022
Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action
Lu, J., C. Clark, S. Lee, Z. Zhang, S. Khosla, R. Marten, D. Hoiem, and A. Kembhavi (2023) · 2023
Later among the works it cites.
Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning
Lu, P., L. Qiu, K.-W. Chang, Y. N. Wu, S.-C. Zhu, T. Rajpurohit, P. Clark, and A. Kalyan (2023) · 2023
Later among the works it cites.
MAPL: Parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting
Mañas, O., P. Rodriguez Lopez, S. Ahmadi, A. Nematzadeh, Y. Goyal, and A. Agrawal (2023, May) · 2023
Later among the works it cites.
The refinedweb dataset for falcon LLM: Outperforming curated corpora with web data only
Penedo, G., Q. Malartic, D. Hesslow, R. Cojocaru, H. Alobeidli, A. Cappelli, B. Pannier, E. Almazrouei, and J. Launay (2023) · 2023
Later among the works it cites.
ep-alm: Efficient perceptual augmentation of language models
Shukor, M., C. Dancette, and M. Cord (2023, oct) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
HiTab: A hierarchical table dataset for question answering and natural language generation
Cheng, Z., H. Dong, Z. Wang, R. Jia, J. Guo, Y. Gao, S. Han, J.-G. Lou, and D. Zhang (2022, May) · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) · 2022
Cited alongside, same era.
The bigscience roots corpus: A 1.6tb composite multilingual dataset
Laurençon, H., L. Saulnier, T. Wang, C. Akiki, A. Villanova del Moral, T. Le Scao, L. Von Werra, C. Mou, E. González Ponferrada, H. Nguyen, J. Frohberg, M. Šaško, Q. Lhoest, A. McMillan-Major, G. Dupont, S. Biderman, A. Rogers, L. Ben allal, F. De Toni, G. Pistilli, O. Nguyen, S. Nikpoor, M. Masoud, P. Colombo, J. de la Rosa, P. Villegas, T. Thrush, S. Longpre, S. Nagel, L. Weber, M. Muñoz, J. Zhu, D. Van Strien, Z. Alyafeai, K. Almubarak, M. C. Vu, I. Gonzalez-Dios, A. Soroa, K. Lo, M. Dey, P. Ortiz Suarez, A. Gokaslan, S. Bose, D. Adelani, L. Phan, H. Tran, I. Yu, S. Pai, J. Chim, V. Lepercq, S. Ilic, M. Mitchell, S. A. Luccioni, and Y. Jernite (2022) · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., D. Li, C. Xiong, and S. Hoi (2022) · 2022
Cited alongside, same era.
Clevr-math: A dataset for compositional language, visual, and mathematical reasoning
Lindström, A. D. (2022) · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan (2022) · 2022
Cited alongside, same era.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Masry, A., D. Long, J. Q. Tan, S. Joty, and E. Hoque (2022, May) · 2022
Cited alongside, same era.
Generative multimodal models are in-context learners
Sun, Q., Y. Cui, X. Zhang, F. Zhang, Q. Yu, Z. Luo, Y. Wang, Y. Rao, J. Liu, T. Huang, and X. Wang (2023) · 2023
Later among the works it cites.
Eva-clip: Improved training techniques for clip at scale
Sun, Q., Y. Fang, L. Wu, X. Wang, and Y. Cao (2023) · 2023
Later among the works it cites.
Aligning large multimodal models with factually augmented rlhf
Sun, Z., S. Shen, S. Cao, H. Liu, C. Li, Y. Shen, C. Gan, L.-Y. Gui, Y.-X. Wang, Y. Yang, K. Keutzer, and T. Darrell (2023) · 2023
Later among the works it cites.
VisText: A Benchmark for Semantically Rich Chart Captioning
Tang, B. J., A. Boggust, and A. Satyanarayan (2023) · 2023
Later among the works it cites.
Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants
Teknium (2023) · 2023
Later among the works it cites.
Identifying and eliminating csam in generative ml training data and models
Thiel, D. (2023) · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample (2023) · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom (2023) · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Zhai, X., B. Mustafa, A. Kolesnikov, and L. Beyer (2023) · 2023
Later among the works it cites.
Pmc-vqa: Visual instruction tuning for medical visual question answering
Zhang, X., C. Wu, Z. Zhao, W. Lin, Y. Zhang, Y. Wang, and W. Xie (2023) · 2023
Later among the works it cites.
RobuT: A systematic study of table QA robustness against human-annotated adversarial perturbations
Zhao, Y., C. Zhao, L. Nan, Z. Qi, W. Zhang, X. Tang, B. Mi, and D. Radev (2023, July) · 2023
Later among the works it cites.
LIMA: Less is more for alignment
Zhou, C., P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. YU, S. Zhang, G. Ghosh, M. Lewis, L. Zettlemoyer, and O. Levy (2023) · 2023
Later among the works it cites.
Multimodal c4: An open, billion-scale corpus of images interleaved with text
Zhu, W., J. Hessel, A. Awadalla, S. Y. Gadre, J. Dodge, A. Fang, Y. Yu, L. Schmidt, W. Y. Wang, and Y. Choi (2023) · 2023
Later among the works it cites.
Automatikz: Text-guided synthesis of scientific vector graphics with tikz
Belouadi, J., A. Lauscher, and S. Eger (2024) · 2024
Closest in time.
Chart-based reasoning: Transferring capabilities from llms to vlms
Carbune, V., H. Mansoor, F. Liu, R. Aralikatte, G. Baechler, J. Chen, and A. Sharma (2024) · 2024
Closest in time.
Mobilevlm v2: Faster and stronger baseline for vision language model
Chu, X., L. Qiao, X. Zhang, S. Xu, F. Wei, Y. Yang, X. Sun, Y. Hu, X. Lin, B. Zhang, and C. Shen (2024) · 2024
Closest in time.
Vision transformers need registers
Darcet, T., M. Oquab, J. Mairal, and P. Bojanowski (2024) · 2024
Closest in time.
Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
Gao, P., R. Zhang, C. Liu, L. Qiu, S. Huang, W. Lin, S. Zhao, S. Geng, Z. Lin, P. Jin, K. Zhang, W. Shao, C. Xu, C. He, J. He, H. Shao, P. Lu, H. Li, and Y. Qiao (2024) · 2024
Closest in time.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Hu, A., H. Xu, J. Ye, M. Yan, L. Zhang, B. Zhang, C. Li, J. Zhang, Q. Jin, F. Huang, and J. Zhou (2024) · 2024
Closest in time.
NEFTune: Noisy embeddings improve instruction finetuning
Jain, N., P. yeh Chiang, Y. Wen, J. Kirchenbauer, H.-M. Chu, G. Somepalli, B. R. Bartoldson, B. Kailkhura, A. Schwarzschild, A. Saha, M. Goldblum, J. Geiping, and T. Goldstein (2024) · 2024
Closest in time.
Prismatic vlms: Investigating the design space of visually-conditioned language models
Karamcheti, S., S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh (2024) · 2024
Closest in time.
Geomverse: A systematic evaluation of large models for geometric reasoning
Kazemi, M., H. Alvari, A. Anand, J. Wu, X. Chen, and R. Soricut (2024) · 2024
Closest in time.
Unlocking the conversion of web screenshots into html code with the websight dataset
Laurençon, H., L. Tronchon, and V. Sanh (2024) · 2024
Closest in time.
Moai: Mixture of all intelligence for large language and vision models
Lee, B.-K., B. Park, C. W. Kim, and Y. M. Ro (2024) · 2024
Closest in time.
Mini-gemini: Mining the potential of multi-modality vision language models
Li, Y., Y. Zhang, C. Wang, Z. Zhong, Y. Chen, R. Chu, S. Liu, and J. Jia (2024) · 2024
Closest in time.
Monkey: Image resolution and text label are important things for large multi-modal models
Li, Z., B. Yang, Q. Liu, Z. Ma, S. Zhang, J. Yang, Y. Sun, Y. Liu, and X. Bai (2024) · 2024
Closest in time.
Moe-llava: Mixture of experts for large vision-language models
Lin, B., Z. Tang, Y. Ye, J. Cui, B. Zhu, P. Jin, J. Huang, J. Zhang, M. Ning, and L. Yuan (2024) · 2024
Closest in time.
Vila: On pre-training for visual language models
Lin, J., H. Yin, W. Ping, Y. Lu, P. Molchanov, A. Tao, H. Mao, J. Kautz, M. Shoeybi, and S. Han (2024) · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge
Liu, H., C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee (2024, January) · 2024
Closest in time.
Infimm-hd: A leap forward in high-resolution multimodal understanding
Liu, H., Q. You, X. Han, Y. Wang, B. Zhai, Y. Liu, Y. Tao, H. Huang, R. He, and H. Yang (2024) · 2024
Closest in time.
Dora: Weight-decomposed low-rank adaptation
Liu, S.-Y., C.-Y. Wang, H. Yin, P. Molchanov, Y.-C. F. Wang, K.-T. Cheng, and M.-H. Chen (2024) · 2024
Closest in time.
Deepseek-vl: Towards real-world vision-language understanding
Lu, H., W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, Y. Sun, C. Deng, H. Xu, Z. Xie, and C. Ruan (2024) · 2024
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao (2024) · 2024
Closest in time.
Mm1: Methods, analysis & insights from multimodal llm pre-training
McKinzie, B., Z. Gan, J.-P. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, F. Weers, A. Belyi, H. Zhang, K. Singh, D. Kang, A. Jain, H. Hè, M. Schwarzer, T. Gunter, X. Kong, A. Zhang, J. Wang, C. Wang, N. Du, T. Lei, S. Wiseman, G. Yin, M. Lee, Z. Wang, R. Pang, P. Grasch, A. Toshev, and Y. Yang (2024) · 2024
Closest in time.
Orca-math: Unlocking the potential of slms in grade school math
Mitra, A., H. Khanpour, C. Rosset, and A. Awadallah (2024) · 2024
Closest in time.
Gpt-4 technical report
OpenAI (2024) · 2024
Closest in time.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, E., Y. Dong, and N. Abu-Ghazaleh (2024) · 2024
Closest in time.
Improved baselines for data-efficient perceptual augmentation of llms
Vallaeys, T., M. Shukor, M. Cord, and J. Verbeek (2024) · 2024
Closest in time.
Cogvlm: Visual expert for pretrained language models
Wang, W., Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, X. Song, J. Xu, B. Xu, J. Li, Y. Dong, M. Ding, and J. Tang (2024) · 2024
Closest in time.
Palm2-vadapter: Progressively aligned language model makes a strong vision-language adapter
Xiao, J., Z. Xu, A. Yuille, S. Yan, and B. Wang (2024) · 2024
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., W. Jiang, H. Shi, J. YU, Z. Liu, Y. Zhang, J. Kwok, Z. Li, A. Weller, and W. Liu (2024) · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen (2024) · 2024
Closest in time.
MAmmoTH: Building math generalist models through hybrid instruction tuning
Yue, X., X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen (2024) · 2024
Closest in time.
Tinyllava: A framework of small-scale large multimodal models
Zhou, B., Y. Hu, X. Weng, J. Jia, J. Luo, X. Liu, J. Wu, and L. Huang (2024) · 2024
Closest in time.