Fetching the paper…
Reading the bibliography…
Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wider range of tasks.
Unifying vision-and-language tasks via text generation
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021 · 1942
Earlier work this paper cites.
LayoutLM: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Earlier work this paper cites.
Docformer: End-to-end transformer for document understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R. Manmatha. 2021 · 2021
Earlier work this paper cites.
Docvqa: A dataset for VQA on document images
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021 · 2021
Earlier work this paper cites.
LayoutLMv2: Multi-modal pre-training for visually-rich document understanding
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou. 2021 · 2021
Earlier work this paper cites.
Layoutlmv3: Pre-training for document AI with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022 · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq R. Joty, and Enamul Hoque. 2022 · 2022
Earlier work this paper cites.
Hierarchical multimodal transformers for multi-page docvqa
Rubèn Tito, Dimosthenis Karatzas, and Ernest Valveny. 2022 · 2022
Earlier work this paper cites.
Document understanding dataset and evaluation (DUDE)
Jordy Van Landeghem, Rafal Powalski, Rubèn Tito, Dawid Jurkiewicz, Matthew B. Blaschko, Lukasz Borchmann, Mickaël Coustaty, Sien Moens, Michal Pietruszka, Bertrand Anckaert, Tomasz Stanislawek, Pawel Józiak, and Ernest Valveny. 2023 · 2023
Earlier work this paper cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 · 2023
Earlier work this paper cites.
Qwen2-VL: Enhancing vision-language model’s perception of the world at any resolution
Alibaba. 2024 · 2024
Earlier work this paper cites.
Claude 3.5 sonnet
Anthropic. 2024 · 2024
Cited alongside, same era.
GRAM: global reasoning for multi-page VQA
Tsachi Blau, Sharon Fogel, Roi Ronen, Alona Golts, Shahar Tsiper, Elad Ben-Avraham, Aviad Aberdam, Roy Ganz, and Ron Litman. 2024 · 2024
Cited alongside, same era.
Yew Ken Chia, Liying Cheng, Hou Pong Chan, Chaoqun Liu, Maojia Song, Sharifah Mahani Aljunied, Soujanya Poria, and Lidong Bing. 2024 · 2024
Cited alongside, same era.
M3docrag: Multi-modal retrieval is what you need for multi-page multi-document understanding
Jaemin Cho, Debanjan Mahata, Ozan Irsoy, Yujie He, and Mohit Bansal. 2024 · 2024
Cited alongside, same era.
MMVQA: A comprehensive dataset for investigating multipage multimodal information retrieval in pdf-based visual question answering
Yihao Ding, Kaixuan Ren, Jiabin Huang, Siwen Luo, and Soyeon Caren Han. 2024a · 2024
Multi-page document visual question answering using self-attention scoring mechanism
Lei Kang, Rubèn Tito, Ernest Valveny, and Dimosthenis Karatzas. 2024 · 2024
Closest in time.
Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models
Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. 2024 · 2024
Closest in time.
TextMonkey: An ocr-free large multimodal model for understanding document
Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai. 2024 · 2024
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2024 · 2024
Closest in time.
Mmlongbench-doc: Benchmarking long-context document understanding with visualizations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-page document VQA with recurrent memory transformer
Qi Dong, Lei Kang, and Dimosthenis Karatzas. 2024a · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
GeminiTeam. 2024 · 2024
Cited alongside, same era.
mPLUG-DocOwl2: High-resolution compressing for ocr-free multi-page document understanding
Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. 2024 · 2024
Cited alongside, same era.
Leopard: A vision language model for text-rich multi-image tasks
Mengzhao Jia, Wenhao Yu, Kaixin Ma, Tianqing Fang, Zhihan Zhang, Siru Ouyang, Hongming Zhang, Meng Jiang, and Dong Yu. 2024 · 2024
Cited alongside, same era.
MANTIS: interleaved multi-image instruction tuning
Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, and Wenhu Chen. 2024 · 2024
Cited alongside, same era.
PDF-MVQA: A dataset for multimodal information retrieval in pdf-based visual question answering
Yihao Ding, Kaixuan Ren, Jiabin Huang, Siwen Luo, and Soyeon Caren Han. 2024b
Cited in the paper.
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, et al. 2024b
Cited in the paper.
Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, Pan Zhang, Liangming Pan, Yu-Gang Jiang, Jiaqi Wang, Yixin Cao, and Aixin Sun. 2024 · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta. 2024 · 2024
Closest in time.
Hello gpt-4o
OpenAI. 2024 · 2024
Closest in time.
Webquest: A benchmark for multimodal QA on web page sequences
Maria Wang, Srinivas Sunkara, Gilles Baechler, Jason Lin, Yun Zhu, Fedir Zubach, Lei Shu, and Jindong Chen. 2024 · 2024
Closest in time.
Visrag: Vision-based retrieval-augmented generation on multi-modality documents
Shi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui, Junhao Ran, Yukun Yan, Zhenghao Liu, Shuo Wang, Xu Han, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Closest in time.
CREAM: coarse-to-fine retrieval and multi-modal efficient tuning for document VQA
Jinxu Zhang, Yongqi Yu, and Yu Zhang. 2024b · 2024
Closest in time.