Fetching the paper…
Reading the bibliography…
Text-rich document understanding (TDU) requires comprehensive analysis of documents containing substantial textual content and complex layouts.
StrucTexT: Structured Text Understanding with Multi-Modal Transformers
Yulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin, Chengquan Zhang, Yan Liu, Kun Yao, Junyu Han, Jingtuo Liu, and Errui Ding · 1920
Earlier work this paper cites.
A Prototype Document Image Analysis System for Technical Journals
G. Nagy, S. Seth, and M. Viswanathan · 1992
Earlier work this paper cites.
Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval
Adam W. Harley, Alex Ufkes, and Konstantinos G. Derpanis · 2015
Earlier work this paper cites.
Compositional Semantic Parsing on Semi-Structured Tables
Panupong Pasupat and Percy Liang · 2015
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and C. V. Jawahar · 2019
Earlier work this paper cites.
FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran · 2019
Earlier work this paper cites.
CORD: A Consolidated Receipt Dataset for Post-OCR Parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee · 2019
Earlier work this paper cites.
PubLayNet: Largest Dataset Ever for Document Layout Analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno-Yepes · 2019
Earlier work this paper cites.
Language Models are Few-shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
TabFact: A Large-scale Dataset for Table-based Fact Verification
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang · 2020
Earlier work this paper cites.
DocBank: A Benchmark Dataset for Document Layout Analysis
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou · 2020
Earlier work this paper cites.
LayoutLM: Pre-training of Text and Layout for Document Image Understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou · 2020
Earlier work this paper cites.
Image-Based Table Recognition: Data, Model, and Evaluation
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno-Yepes · 2020
Earlier work this paper cites.
DocFormer: End-to-End Transformer for Document Understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R. Manmatha · 2021
Earlier work this paper cites.
LAMBERT: Layout-Aware Language Modeling for Information Extraction
Łukasz Garncarek, Rafał Powalski, Tomasz Stanisławek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Graliński · 2021
Earlier work this paper cites.
DocVQA: A Dataset for VQA on Document Images
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar · 2021
Earlier work this paper cites.
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer
Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michał Pietruszka, and Gabriela Pałka · 2021
Earlier work this paper cites.
Kleister: Key Information Extraction Datasets Involving Long Documents with Complex Layouts
Tomasz Stanislawek, Filip Gralinski, Anna Wróblewska, Dawid Lipinski, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemyslaw Biecek · 2021
Earlier work this paper cites.
VisualMRC: Machine Reading Comprehension on Document Images
Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida · 2021
Earlier work this paper cites.
LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou · 2021
Earlier work this paper cites.
BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents
Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park · 2022
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei · 2022
Cited alongside, same era.
OCR-Free Document Understanding Transformer
Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park · 2022
Cited alongside, same era.
Relational Representation Learning in Visually-Rich Documents
Xin Li, Yan Zheng, Yiqing Hu, Haoyu Cao, Yunfei Wu, Deqiang Jiang, Yinsong Liu, and Bo Ren · 2022
Cited alongside, same era.
InfographicVQA
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and C.V. Jawahar · 2022
Cited alongside, same era.
ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding
DocFormerv2: Local Features for Document Understanding
Srikar Appalaraju, Peng Tang, Qi Dong, Nishant Sankaran, Yichu Zhou, and R. Manmatha · 2024
Closest in time.
LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
Masato Fujitake · 2024
Closest in time.
CogAgent: A Visual Language Model for GUI Agents
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, and Jie Tang · 2024
Closest in time.
Building and better understanding vision-language models: insights and future directions
Hugo Laurençon, Andrés Marafioti, Victor Sanh, and Leo Tronchon · 2024
Closest in time.
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Yuhui Cao, Weichong Yin, Yongfeng Chen, Yin Zhang, Shikun Feng, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2022
Cited alongside, same era.
DocLayNet: A Large Human-Annotated Dataset for Document-Layout Segmentation
Birgit Pfitzmann, Christoph Auer, Michele Dolfi, Ahmed S. Nassar, and Peter W. J. Staar · 2022
Cited alongside, same era.
LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding
Jiapeng Wang, Lianwen Jin, and Kai Ding · 2022
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Cited alongside, same era.
Disentangling Writer and Character Styles for Handwriting Generation
Gang Dai, Yifan Zhang, Qingfeng Wang, Qing Du, Zhuliang Yu, Zhuoman Liu, and Shuangping Huang · 2023
Cited alongside, same era.
ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information Extraction
Jiabang He, Lei Wang, Yi Hu, Ning Liu, Hui Liu, Xing Xu, and Heng Tao Shen · 2023
Cited alongside, same era.
Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models
Geewook Kim, Hodong Lee, Daehee Kim, Haeji Jung, Sanghee Park, Yoonsik Kim, Sangdoo Yun, Taeho Kil, Bado Lee, and Seunghyun Park · 2023
Cited alongside, same era.
Meta AI Llama Team · 2024
Closest in time.
Jinghui Lu, Haiyang Yu, Yanjie Wang, Yongjie Ye, Jingqun Tang, Ziwei Yang, Binghong Wu, Qi Liu, Hao Feng, Han Wang, Hao Liu, and Can Huang · 2024
Closest in time.
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
Chuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng, Zhi Yu, and Cong Yao · 2024
Closest in time.
Hello GPT-4o, 2024
OpenAI · 2024
Closest in time.
LMDX: Language Model-based Document Information Extraction and Localization
Vincent Perot, Kai Kang, Florian Luisier, Guolong Su, Xiaoyu Sun, Ramya Sree Boppana, Zilong Wang, Zifeng Wang, Jiaqi Mu, Hao Zhang, Chen-Yu Lee, and Nan Hua · 2024
Closest in time.
DeepForm: Understand Structured Documents at Scale
Stacey Svetlichnaya · 2024
Closest in time.
InstructDoc: A Dataset for Zero-shot Generalization of Visual Document Understanding with Instructions
Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito, and Jun Suzuki · 2024
Closest in time.
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, Wei Shi, Yuliang Liu, Hao Liu, Yuan Xie, Xiang Bai, and Can Huang · 2024
Closest in time.
QwQ: Reflect Deeply on the Boundaries of the Unknown, 2024
Qwen Team · 2024
Closest in time.
OmniParser: A Unified Framework for Text Spotting Key Information Extraction and Table Recognition
Jianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu, Wenqing Cheng, Fei Huang, Xiang Bai, Cong Yao, and Zhibo Yang · 2024
Closest in time.
DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding
Dongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, and Xiaomo Liu · 2024
Closest in time.
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
Jiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie, and Lianwen Jin · 2024
Closest in time.
One-dm: One-shot diffusion mimicker for handwritten text generation
Gang Dai, Yifan Zhang, Quhui Ke, Qiangya Guo, and Shuangping Huang · 2025
Closest in time.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI · 2025
Closest in time.
LLaVA-OneVision: Easy Visual Task Transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li · 2025
Closest in time.
LiLTv2: Language-substitutable layout-image transformer for visual information extraction
Jiapeng Wang, Zening Lin, Dayi Huang, Longfei Xiong, and Lianwen Jin · 2025
Closest in time.