Fetching the paper…
Reading the bibliography…
Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems.
Tablebank: Table benchmark for image-based table detection and recognition
Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, Ming Zhou, and Zhoujun Li · 1925
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein et al · 1966
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Docbank: A benchmark dataset for document layout analysis
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou · 2006
Earlier work this paper cites.
Adapting the tesseract open source ocr engine for multilingual ocr
Ray Smith, Daria Antonova, and Dar-Shyang Lee · 2009
Earlier work this paper cites.
Tabtransformer: Tabular data modeling using contextual embeddings. arxiv 2020
Xin Huang, Ashish Khetan, Milan Cvitkovic, and Zohar Karnin · 2012
Earlier work this paper cites.
Icdar 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, Faisal Shafait, Seiichi Uchida, and Ernest Valveny · 2015
Earlier work this paper cites.
Image-to-markup generation with coarse-to-fine attention
Yuntian Deng, Anssi Kanervisto, Jeffrey Ling, and Alexander M Rush · 2017
Earlier work this paper cites.
Multi-scale attention with dense encoder for handwritten mathematical expression recognition
Jianshu Zhang, Jun Du, and Lirong Dai · 2018
Earlier work this paper cites.
Publaynet: largest dataset ever for document layout analysis
Zhong Xu, Jianbin Tang, and Antonio Jimeno Yepes · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Improving attention-based handwritten mathematical expression recognition with scale augmentation and drop attention
Zhe Li, Lianwen Jin, Songxuan Lai, and Yecheng Zhu · 2020
Earlier work this paper cites.
Abcnet: Real-time scene text spotting with adaptive bezier-curve network
Yuliang Liu, Hao Chen, Chunhua Shen, Tong He, Lianwen Jin, and Liangwei Wang · 2020
Earlier work this paper cites.
Subhojeet Pramanik, Shashank Mujumdar, and Hima Patel · 2020
Earlier work this paper cites.
Image-based table recognition: data, model, and evaluation
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes · 2020
Cited alongside, same era.
Tablex: a benchmark dataset for structure and content information extraction from scientific tables
Harsh Desai, Pratik Kayal, and Mayank Singh · 2021
Cited alongside, same era.
Unidoc: Unified pretraining framework for document understanding
Jiuxiang Gu, Jason Kuen, Vlad I Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, and Tong Sun · 2021
Cited alongside, same era.
Spatial dependency parsing for semi-structured document information extraction
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo · 2021
Cited alongside, same era.
Pgnet: Real-time arbitrarily-shaped text spotting with point gathering network
Pengfei Wang, Chengquan Zhang, Fei Qi, Shanshan Liu, Xiaoqiang Zhang, Pengyuan Lyu, Junyu Han, Jingtuo Liu, Errui Ding, and Guangming Shi · 2021
Cited alongside, same era.
Hello gpt 4o, 2024
Open AI · 2024
Closest in time.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2024
Closest in time.
Graphkd: Exploring knowledge distillation towards document object detection with structured graph creation
Ayan Banerjee, Sanket Biswas, Josep Lladós, and Umapada Pal · 2024
Closest in time.
pix2tex - latex ocr
Lukas Blecher · 2024
Closest in time.
Nougat: Neural optical understanding for academic documents
Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Doclaynet: A large human-annotated dataset for document-layout segmentation
Birgit Pfitzmann, Christoph Auer, Michele Dolfi, Ahmed S Nassar, and Peter Staar · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Swindocsegmenter: An end-to-end unified domain adaptive transformer for document instance segmentation
Ayan Banerjee, Sanket Biswas, Josep Lladós, and Umapada Pal · 2023
Cited alongside, same era.
M6doc: A large-scale multi-format, multi-type, multi-layout, multi-language, multi-annotation category dataset for modern document layout analysis
Hiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang, Qiyuan Zhu, Zecheng Xie, Jing Li, Kai Ding, and Lianwen Jin · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang · 2023
Cited alongside, same era.
Improving table structure recognition with visual-alignment sequential coordinate modeling
Yongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li, Zecheng Xie, Shenggao Zhu, Liangcai Gao, and Wei Peng · 2023
Cited alongside, same era.
Rapidtable
RapidAI · 2023
Cited alongside, same era.
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai · 2024
Closest in time.
mplug-docowl2: High-resolution compressing for ocr-free multi-page document understanding
Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou · 2024
Closest in time.
Kosmos-2.5: A multimodal literate model, 2024
Tengchao Lv, Yupan Huang, Jingye Chen, Yuzhong Zhao, Yilin Jia, Lei Cui, Shuming Ma, Yaoyao Chang, Shaohan Huang, Wenhui Wang, Li Dong, Weiyao Luo, Shaoxiang Wu, Guoxin Wang, Cha Zhang, and Furu Wei · 2024
Closest in time.
Marker, 2024
Vik Paruchuri · 2024
Closest in time.
General ocr theory: Towards ocr-2.0 via a unified end-to-end model
Haoran Wei, Chenglong Liu, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, et al · 2024
Closest in time.
Qintong Zhang, Victor Shea-Jay Huang, Bin Wang, Junyuan Zhang, Zhengren Wang, Hao Liang, Shawn Wang, Matthieu Lin, Wentao Zhang, and Conghui He · 2024
Closest in time.
Doclayout-yolo: Enhancing document layout analysis through diverse synthetic data and global-to-local adaptive perception, 2024
Zhiyuan Zhao, Hengrui Kang, Bin Wang, and Conghui He · 2024
Closest in time.
Structeqtable-deploy: A high-efficiency open-source toolkit for table-to-latex transformation
Hongbin Zhou, Xiangchao Yan, and Bo Zhang · 2024
Closest in time.
Vary: Scaling up the vision vocabulary for large vision-language model
Haoran Wei, Lingyu Kong, Jinyue Chen, Liang Zhao, Zheng Ge, Jinrong Yang, Jianjian Sun, Chunrui Han, and Xiangyu Zhang · 2025
Closest in time.