Fetching the paper…
Reading the bibliography…
We introduce NVLM 1.0, a family of frontier-class multimodal large language models (LLMs) that achieve state-of-the-art results on vision-language tasks, rivaling the leading proprietary models (e.g., GPT-4o) and open-access models (e.g., Llama 3-V 405B and InternVL 2).
The iam-database: an english sentence database for offline handwriting recognition
Marti, U.-V. and Bunke, H · 2002
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T · 2011
Earlier work this paper cites.
Icfhr 2014 competition on handwritten digit string recognition in challenging datasets (hdsrc 2014)
Diem, M., Fiel, S., Kleber, F., Sablatnig, R., Saavedra, J. M., Contreras, D., Barrios, J. M., and Oliveira, L. S · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Pasupat, P. and Liang, P · 2015
Earlier work this paper cites.
Solving geometry problems: Combining text and diagram interpretation
Seo, M., Hajishirzi, H., Farhadi, A., Etzioni, O., and Malcolm, C · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Coco-text: Dataset and benchmark for text detection and recognition in natural images
Veit, A., Matera, T., Neumann, L., Matas, J., and Belongie, S · 2016
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Visual7W: Grounded question answering in images
Zhu, Y., Groth, O., Bernstein, M., and Fei-Fei, L · 2016
Earlier work this paper cites.
Making the v in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Figureqa: An annotated figure dataset for visual reasoning
Kahou, S. E., Michalski, V., Atkinson, A., Kádár, Á., Trischler, A., and Bengio, Y · 2017
Earlier work this paper cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Kembhavi, A., Seo, M., Schwenk, D., Choi, J., Farhadi, A., and Hajishirzi, H · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al · 2017
Earlier work this paper cites.
Icdar2017 competition on reading chinese text in the wild (rctw-17)
Shi, B., Yao, C., Liao, M., Yang, M., Xu, P., Cui, L., Belongie, S., Lu, S., and Bai, X · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
DVQA: Understanding data visualizations via question answering
Kafle, K., Price, B., Cohen, S., and Kanan, C · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Tallyqa: Answering complex counting questions
Acharya, M., Kafle, K., and Kanan, C · 2019
Earlier work this paper cites.
Scene text visual question answering
Biten, A. F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D · 2019
Earlier work this paper cites.
Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art
Chng, C. K., Liu, Y., Sun, Y., Ng, C. C., Luo, C., Ni, Z., Fang, C., Zhang, S., Han, J., Ding, E., et al · 2019
Earlier work this paper cites.
Icdar2019 competition on scanned receipt ocr and information extraction
Huang, Z., Chen, K., He, J., Bai, X., Karatzas, D., Lu, S., and Jawahar, C · 2019
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Jaume, G., Ekenel, H. K., and Thiran, J.-P · 2019
Earlier work this paper cites.
OK-VQA: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
OCR-VQA: Visual question answering by reading text in images
Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Towards VQA models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt
Sun, Y., Ni, Z., Chng, C.-K., Liu, Y., Luo, C., Ng, C. C., Han, J., Ding, E., Liu, J., Karatzas, D., et al · 2019
Earlier work this paper cites.
Icdar 2019 robust reading challenge on reading chinese text on signboard
Zhang, R., Zhou, Y., Jiang, Q., Song, Q., Li, N., Zhou, K., Wang, L., Wang, D., Liao, M., Yang, M., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Plotqa: Reasoning over scientific plots
Methani, N., Ganguly, P., Khapra, M. M., and Kumar, P · 2020
Earlier work this paper cites.
Chart-to-text: Generating natural language descriptions for charts by adapting the transformer model
Obeid, J. and Hoque, E · 2020
Earlier work this paper cites.
Textcaps: a dataset for image captioning with reading comprehension
Sidorov, O., Hu, R., Rohrbach, M., and Singh, A · 2020
Earlier work this paper cites.
On the general value of evidence, and bilingual scene-text visual question answering
Wang, X., Liu, Y., Shen, C., Ng, C. C., Luo, C., Jin, L., Chan, C. S., Hengel, A. v. d., and Wang, L · 2020
Earlier work this paper cites.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., Piao, S., and Wei, F · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Hitab: A hierarchical table dataset for question answering and natural language generation
Cheng, Z., Dong, H., Wang, Z., Jia, R., Guo, J., Gao, Y., Han, S., Lou, J.-G., and Zhang, D · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
OpenCLIP, 2021
Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J · 2021
Cited alongside, same era.
An open source implementation of CLIP, 2021
mlfoundations · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Cited alongside, same era.
Llava-pretrain
Liu, H · 2023
Later among the works it cites.
Llama-Team · 2023
Later among the works it cites.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D · 2023
Later among the works it cites.
Orca: Progressive learning from complex explanation traces of gpt-4
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A · 2023
Later among the works it cites.
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
Singh, A., Pang, G., Toh, M., Huang, J., Galuba, W., and Hassner, T · 2021
Cited alongside, same era.
Visualmrc: Machine reading comprehension on document images
Tanaka, R., Nishida, K., and Yoshida, S · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, J · 2021
Cited alongside, same era.
Flamingo: A visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Coyo-700m: Image-text pair dataset
Byeon, M., Park, B., Kim, H., Lee, S., Baek, W., and Kim, S · 2022
Cited alongside, same era.
An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Cao, J. and Xiao, J · 2022
Cited alongside, same era.
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Tanaka, R., Nishida, K., Nishida, K., Hasegawa, T., Saito, I., and Saito, K · 2023
Later among the works it cites.
GPTeacher-General-Instruct, 2023
Teknium · 2023
Later among the works it cites.
ShareGPT-Vicuna, 2023
The-Vicuna-Team · 2023
Later among the works it cites.
Document understanding dataset and evaluation (dude)
Van Landeghem, J., Tito, R., Borchmann, Ł., Pietruszka, M., Joziak, P., Powalski, R., Jurkiewicz, D., Coustaty, M., Anckaert, B., Valveny, E., et al · 2023
Later among the works it cites.
Cogvlm: Visual expert for pretrained language models
Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., et al · 2023
Later among the works it cites.
Re-ViLM: Retrieval-augmented visual language model for zero and few-shot image captioning
Yang, Z., Ping, W., Liu, Z., Korthikanti, V., Nie, W., Huang, D.-A., Fan, L., Yu, Z., Lan, S., Li, B., et al · 2023
Later among the works it cites.
Ye, J., Hu, A., Xu, H., Ye, Q., Yan, M., Xu, G., Li, C., Tian, J., Qian, Q., Zhang, J., et al · 2023
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W · 2023
Later among the works it cites.
Mammoth: Building math generalist models through hybrid instruction tuning
Yue, X., Qu, X., Zhang, G., Fu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Later among the works it cites.
LLaVAR: Enhanced visual instruction tuning for text-rich image understanding
Zhang, Y., Zhang, R., Gu, J., Zhou, Y., Lipka, N., Yang, D., and Sun, T · 2023
Later among the works it cites.
RobuT: A systematic study of table QA robustness against human-annotated adversarial perturbations
Zhao, Y., Zhao, C., Nan, L., Qi, Z., Zhang, W., Tang, X., Mi, B., and Radev, D · 2023
Later among the works it cites.
Adept Fuyu-Heavy: A new multimodal model, 2024
Adept · 2024
Closest in time.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2024
Closest in time.
Dong, X., Zhang, P., Zang, Y., Cao, Y., Wang, B., Ouyang, L., Zhang, S., Duan, H., Zhang, W., Li, Y., et al · 2024
Closest in time.
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models, 2024
Duan, H., Yang, J., Qiao, Y., Fang, X., Chen, L., Liu, Y., Dong, X., Zang, Y., Zhang, P., Wang, J., Lin, D., and Chen, K · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini-Team · 2024
Closest in time.
ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots, 2024
Hsiao, Y.-C., Zubach, F., Wang, M., and Chen, J · 2024
Closest in time.
Sharegpt-4o, 2024
Laboratory, S. A · 2024
Closest in time.
Vila: On pre-training for visual language models
Lin, J., Yin, H., Ping, W., Molchanov, P., Shoeybi, M., and Han, S · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2024
Closest in time.
Docmatix - a huge dataset for document visual question answering, 2024
Marafioti, A. and Laurencon, H · 2024
Closest in time.
Orca-math: Unlocking the potential of slms in grade school math
Mitra, A., Khanpour, H., Rosset, C., and Awadallah, A · 2024
Closest in time.
Nous-Hermes-2-Yi-34B, 2024
Nous · 2024
Closest in time.
Qwen2 technical report, 2024
Qwen-Team · 2024
Closest in time.
Reka Core: Our Frontier Class Multimodal Language Model, 2024
Reka · 2024
Closest in time.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Tong, S., Brown, E., Wu, P., Woo, S., Middepogu, M., Akula, S. C., Yang, J., Yang, S., Iyer, A., Pan, X., et al · 2024
Closest in time.
Magicoder: Empowering code generation with oss-instruct
Wei, Y., Wang, Z., Liu, J., Ding, Y., and Zhang, L · 2024
Closest in time.
WizardLM: Empowering large pre-trained language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., Lin, Q., and Jiang, D · 2024
Closest in time.
Minicpm-v: A gpt-4v level mllm on your phone
Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., Cai, T., Li, H., Zhao, W., He, Z., et al · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., et al · 2024
Closest in time.
Advancing llm reasoning generalists with preference trees
Yuan, L., Cui, G., Wang, H., Ding, N., Wang, X., Deng, J., Shan, B., Chen, H., Xie, R., Lin, Y., et al · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W · 2024
Closest in time.
Mavis: Mathematical visual instruction tuning
Zhang, R., Wei, X., Jiang, D., Zhang, Y., Guo, Z., Tong, C., Liu, J., Zhou, A., Wei, B., Zhang, S., et al · 2024
Closest in time.