Fetching the paper…
Reading the bibliography…
We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, particularly in visual document understanding.
The iam-database: an english sentence database for offline handwriting recognition, 2002
U-V Marti and Horst Bunke · 2002
Earlier work this paper cites.
Implementation and benchmarking of perceptual image hash functions, 2010
Christoph Zauner · 2010
Earlier work this paper cites.
Scene text recognition using higher order language priors, 2012
Anand Mishra, Karteek Alahari, and CV Jawahar · 2012
Earlier work this paper cites.
Evaluation of deep convolutional nets for document image classification and retrieval, 2015
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis · 2015
Earlier work this paper cites.
A diagram is worth a dozen images, 2016
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Figureqa: An annotated figure dataset for visual reasoning, 2017
Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Ákos Kádár, Adam Trischler, and Yoshua Bengio · 2017
Earlier work this paper cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension, 2017
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people, 2018
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering, 2018
Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan · 2018
Earlier work this paper cites.
Scene text visual question answering, 2019
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images, 2019
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty · 2019
Earlier work this paper cites.
Towards vqa models that can read, 2019
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Earlier work this paper cites.
Plotqa: Reasoning over scientific plots, 2020
Nitesh Methani, Pritha Ganguly, Mitesh M Khapra, and Pratyush Kumar · 2020
Earlier work this paper cites.
Chart-to-text: Generating natural language descriptions for charts by adapting the transformer model, 2020
Jason Obeid and Enamul Hoque · 2020
Earlier work this paper cites.
Textcaps: a dataset for image captioning with reading comprehension, 2020
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh · 2020
Earlier work this paper cites.
Image-based table recognition: data, model, and evaluation, 2020
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes · 2020
Earlier work this paper cites.
Docformer: End-to-end transformer for document understanding, 2021
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R. Manmatha · 2021
Earlier work this paper cites.
Due: End-to-end document understanding benchmark, 2021
Łukasz Borchmann, Michał Pietruszka, Tomasz Stanislawek, Dawid Jurkiewicz, Michał Turski, Karolina Szyndler, and Filip Graliński · 2021
Earlier work this paper cites.
Finqa: A dataset of numerical reasoning over financial data, 2021
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, et al · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar · 2021
Earlier work this paper cites.
Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text, 2021
Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang, Wojciech Galuba, and Tal Hassner · 2021
Earlier work this paper cites.
Visualmrc: Machine reading comprehension on document images, 2021
Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida · 2021
Earlier work this paper cites.
Screen2words: Automatic mobile ui summarization with multimodal learning, 2021
Bryan Wang, Gang Li, Xin Zhou, Zhourong Chen, Tovi Grossman, and Yang Li · 2021
Earlier work this paper cites.
Tap: Text-aware pre-training for text-vqa and text-caption, 2021
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florencio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo · 2021
Earlier work this paper cites.
Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context
Xinyi Zheng, Doug Burdick, Lucian Popa, Peter Zhong, and Nancy Xin Ru Wang · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning, 2022
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan · 2022
Cited alongside, same era.
Hitab: A hierarchical table dataset for question answering and natural language generation, 2022
Zhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia, Jiaqi Guo, Yan Gao, Shi Han, Jian-Guang Lou, and Dongmei Zhang · 2022
Cited alongside, same era.
Turl: Table understanding through representation learning, 2022
Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu · 2022
Cited alongside, same era.
Screenqa: Large-scale question-answer pairs over mobile app screenshots, 2022
Yu-Chung Hsiao, Fedir Zubach, Gilles Baechler, Victor Carbune, Jason Lin, Maria Wang, Srinivas Sunkara, Yun Zhu, and Jindong Chen · 2022
Cited alongside, same era.
Nvlm: Open frontier-class multimodal llms, 2024
Wenliang Dai, Nayeon Lee, Boxin Wang, Zhuolin Yang, Zihan Liu, Jon Barker, Tuomas Rintamaki, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping · 2024
Later among the works it cites.
Hao Feng, Qi Liu, Hao Liu, Jingqun Tang, Wengang Zhou, Houqiang Li, and Can Huang · 2024
Later among the works it cites.
Granite 3.0 language models, 2024
IBM Granite Team · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, and … · 2024
Later among the works it cites.
Multimodal task vectors enable many-shot multimodal in-context learning, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei · 2022
Cited alongside, same era.
Rfc 9309 robots exclusion protocol, 2022
M Koster, G Illyes, H Zeller, and L Sassman · 2022
Cited alongside, same era.
Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2022
Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and Ashwin Kalyan · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque · 2022
Cited alongside, same era.
Infographicvqa, 2022
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar · 2022
Cited alongside, same era.
Tableformer: Table structure understanding with transformers
Ahmed Nassar, Nikolaos Livathinos, Maksym Lysak, and Peter Staar · 2022
Cited alongside, same era.
Doclaynet: A large human-annotated dataset for document-layout segmentation
Birgit Pfitzmann, Christoph Auer, Michele Dolfi, Ahmed S. Nassar, and Peter Staar · 2022
Cited alongside, same era.
Pubtables-1m: Towards comprehensive table extraction from unstructured documents
Brandon Smock, Rohith Pesala, and Robin Abraham · 2022
Cited alongside, same era.
Brandon Huang, Chancharik Mitra, Assaf Arbelle, Leonid Karlinsky, Trevor Darrell, and Roei Herzig · 2024
Later among the works it cites.
Smolvlm, 2024
Hugging-Face · 2024
Later among the works it cites.
Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web
Raghav Kapoor, Yash Parag Butala, Melisa Russak, Jing Yu Koh, Kiran Kamble, Waseem AlShikh, and Ruslan Salakhutdinov · 2024
Later among the works it cites.
Unlocking the conversion of web screenshots into html code with the websight dataset, 2024
Hugo Laurençon, Léo Tronchon, and Victor Sanh · 2024
Later among the works it cites.
Wenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang, Jun Huang, and Lianwen Jin · 2024
Later among the works it cites.
Sparse attention vectors: Generative multimodal model features are discriminative vision-language classifiers, 2024
Chancharik Mitra, Brandon Huang, Tianning Chai, Zhiqiu Lin, Assaf Arbelle, Rogerio Feris, Leonid Karlinsky, Trevor Darrell, Deva Ramanan, and Roei Herzig · 2024
Later among the works it cites.
Kvp10k: A comprehensive dataset for key-value pair extraction in business documents
Oshri Naparstek, Ophir Azulai, Inbar Shapira, Elad Amrani, Yevgeny Yaroker, Yevgeny Burshtein, Roi Pony, Nadav Rubinstein, Foad Abo Dahood, Orit Prince, et al · 2024
Later among the works it cites.
Juan Rodriguez, Xiangru Jian, Siba Smarak Panigrahi, Tianyu Zhang, Aarash Feizi, Abhay Puri, Akshay Kalkunte, François Savard, Ahmed Masry, Shravan Nayak, Rabiul Awal, Mahsa Massoud, Amirhossein Abaskohi, Zichao Li, Suyuchen Wang, Pierre-André Noël, Mats Leon Richter, Saverio Vadacchino, Shubbam Agarwal, Sanket Biswas, Sara Shanian, Ying Zhang, Noah Bolger, Kurt MacDonald, Simon Fauvel, Sathwik Tejaswi, Srinivas Sunkara, Joao Monteiro, Krishnamurthy DJ Dvijotham, Torsten Scholak, Nicolas Chapados, Sepideh Kharagani, Sean Hughes, M. Özsu, Siva Reddy, Marco Pedersoli, Yoshua Bengio, Christopher Pal, Issam Laradji, Spandanna Gella, Perouz Taslakian, David Vazquez, and Sai Rajeswar · 2024
Later among the works it cites.
NumeroLogic: Number encoding for enhanced LLMs’ numerical reasoning
Eli Schwartz, Leshem Choshen, Joseph Shtok, Sivan Doveh, Leonid Karlinsky, and Assaf Arbelle · 2024
Later among the works it cites.
Livexiv–a multi-modal live benchmark based on arxiv papers content, 2024
Nimrod Shabtay, Felipe Maia Polo, Sivan Doveh, Wei Lin, M Jehanzeb Mirza, Leshem Chosen, Mikhail Yurochkin, Yuekai Sun, Assaf Arbelle, Leonid Karlinsky, et al · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Ziteng Wang, Rob Fergus, Yann LeCun, and Saining Xie · 2024
Later among the works it cites.
The evolution of multimodal model architectures, 2024
Shakti N. Wadekar, Abhishek Chaurasia, Aman Chadha, and Eugenio Culurciello · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al · 2024
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al · 2024
Later among the works it cites.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models, 2024
Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, and Timothy Hospedales · 2024
Later among the works it cites.
Claude 3.5 sonnet, 2024
Anthropic · 2025
Closest in time.
Docmatix: A new approach to document understanding, 2025
HuggingFace · 2025
Closest in time.
Docling: An efficient open-source toolkit for ai-driven document conversion, 2025
Nikolaos Livathinos, Christoph Auer, Maksym Lysak, Ahmed Nassar, Michele Dolfi, Panos Vagenas, Cesar Berrospi Ramis, Matteo Omenetti, Kasper Dinkla, Yusik Kim, Shubham Gupta, Rafael Teixeira de Lima, Valery Weber, Lucas Morin, Ingmar Meijer, Viktor Kuropiatnyk, and Peter W. J. Staar · 2025
Closest in time.
Gpt-4v(ision) technical work and authors, 2023b
OpenAI · 2025
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2025
Closest in time.
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen · 2053
Closest in time.