Fetching the paper…
Reading the bibliography…
We introduce MonkeyOCR, a document parsing model that advances the state of the art by leveraging a Structure-Recognition-Relation (SRR) triplet paradigm.
Icfhr 2014 competition on recognition of on-line handwritten mathematical expressions (crohme 2014)
Harold Mouchere, Christian Viard-Gaudin, Richard Zanibbi, and Utpal Garain · 2014
Earlier work this paper cites.
Icfhr2016 crohme: Competition on recognition of online handwritten mathematical expressions
Harold Mouchère, Christian Viard-Gaudin, Richard Zanibbi, and Utpal Garain · 2016
Earlier work this paper cites.
Icdar 2019 crohme+ tfd: Competition on recognition of handwritten mathematical expressions and typeset formula detection
Mahshad Mahdavi, Richard Zanibbi, Harold Mouchere, Christian Viard-Gaudin, and Utpal Garain · 2019
Earlier work this paper cites.
Image-based table recognition: data, model, and evaluation
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes · 2020
Earlier work this paper cites.
Pingan-vcgroup’s solution for icdar 2021 competition on scientific table image recognition to latex
Yelin He, Xianbiao Qi, Jiaquan Ye, Peng Gao, Yihao Chen, Bingcong Li, Xin Tang, and Rong Xiao · 2021
Earlier work this paper cites.
Cdla: A chinese document layout analysis (cdla) dataset
Hang Li · 2021
Earlier work this paper cites.
Tabular transformers for modeling multivariate time series
Inkit Padhi, Yair Schiff, Igor Melnyk, Mattia Rigotti, Youssef Mroueh, Pierre Dognin, Jerret Ross, Ravi Nair, and Erik Altman · 2021
Earlier work this paper cites.
Layoutreader: Pre-training of text and layout for reading order detection
Zilong Wang, Yiheng Xu, Lei Cui, Jingbo Shang, and Furu Wei · 2021
Earlier work this paper cites.
Jiaquan Ye, Xianbiao Qi, Yelin He, Yihao Chen, Dengyi Gu, Peng Gao, and Rong Xiao · 2021
Earlier work this paper cites.
Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context
Xinyi Zheng, Douglas Burdick, Lucian Popa, Xu Zhong, and Nancy Xin Ru Wang · 2021
Earlier work this paper cites.
Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding
Zhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan, Weiqiang Wang, Ming Gu, and Liqing Zhang · 2022
Earlier work this paper cites.
Layoutlmv3: Pre-training for document ai with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei · 2022
Earlier work this paper cites.
Ocr-free document understanding transformer
Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park · 2022
Earlier work this paper cites.
Pp-ocrv3: More attempts for the improvement of ultra lightweight ocr system
Chenxia Li, Weiwei Liu, Ruoyu Guo, Xiaoting Yin, Kaitao Jiang, Yongkun Du, Yuning Du, Lingfeng Zhu, Baohua Lai, Xiaoguang Hu, et al · 2022
Earlier work this paper cites.
Spts: single-point text spotting
Dezhi Peng, Xinyu Wang, Yuliang Liu, Jiaxin Zhang, Mingxin Huang, Songxuan Lai, Jing Li, Shenggao Zhu, Dahua Lin, Chunhua Shen, et al · 2022
Earlier work this paper cites.
Doclaynet: A large human-annotated dataset for document-layout segmentation
Birgit Pfitzmann, Christoph Auer, Michele Dolfi, Ahmed S Nassar, and Peter Staar · 2022
Earlier work this paper cites.
Syntax-aware network for handwritten mathematical expression recognition
Ye Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu, Zhilong Ji, Zhongqin Wu, and Xiang Bai · 2022
Earlier work this paper cites.
Nougat: Neural optical understanding for academic documents
Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic · 2023
Earlier work this paper cites.
M6doc: a large-scale multi-format, multi-type, multi-layout, multi-language, multi-annotation category dataset for modern document layout analysis
Hiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang, Qiyuan Zhu, Zecheng Xie, Jing Li, Kai Ding, and Lianwen Jin · 2023
Cited alongside, same era.
Vision grid transformer for document layout analysis
Cheng Da, Chuwei Luo, Qi Zheng, and Cong Yao · 2023
Cited alongside, same era.
Ultralytics yolov8
Glenn Jocher, Ayush Chaurasia, and Jing Qiu · 2023
Cited alongside, same era.
Spts v2: single-point scene text spotting
Yuliang Liu, Jiaxin Zhang, Dezhi Peng, Mingxin Huang, Xinyu Wang, Jingqun Tang, Can Huang, Dahua Lin, Chunhua Shen, Xiang Bai, et al · 2023
Cited alongside, same era.
pix2tex - latex ocr
Lukas Blecher · 2024
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al · 2024
Open parse: Visually-driven document chunking for llm applications
Sergey Filimonov · 2025
Closest in time.
Gemini 2.5
Google DeepMind · 2025
Closest in time.
The unreasonable ineffectiveness of the deeper layers
Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A. Roberts · 2025
Closest in time.
Docling: An efficient open-source toolkit for ai-driven document conversion
Nikolaos Livathinos, Christoph Auer, Maksym Lysak, Ahmed Nassar, Michele Dolfi, Panos Vagenas, Cesar Berrospi Ramis, Matteo Omenetti, Kasper Dinkla, Yusik Kim, et al · 2025
Closest in time.
Nanonets-ocr-s: A model for transforming documents into structured markdown with intelligent content recognition and semantic tagging
Souvik Mandal, Ashish Talewar, Paras Ahuja, and Prathamesh Juvatkar · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Easyocr: Ready-to-use ocr with 80+ supported languages
Jaided AI · 2024
Cited alongside, same era.
Monkey: Image resolution and text label are important things for large multi-modal models
Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai · 2024
Cited alongside, same era.
Textmonkey: An ocr-free large multimodal model for understanding document
Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai · 2024
Cited alongside, same era.
Hello gpt-4o
OpenAI · 2024
Cited alongside, same era.
Omniparser: A unified framework for text spotting key information extraction and table recognition
Jianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu, Wenqing Cheng, Fei Huang, Xiang Bai, Cong Yao, and Zhibo Yang · 2024
Cited alongside, same era.
General ocr theory: Towards ocr-2.0 via a unified end-to-end model
Haoran Wei, Chenglong Liu, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, et al · 2024
Cited alongside, same era.
Mathpix · 2025
Closest in time.
Shortgpt: Layers in large language models are more redundant than you expect
Xin Men, Mingyu Xu, Qingyu Zhang, Qianhao Yuan, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen · 2025
Closest in time.
Mistral ocr: Free online ai ocr tool to extract text
Mistral OCR · 2025
Closest in time.
Smoldocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
Ahmed Nassar, Andres Marafioti, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos, Christoph Auer, Lucas Morin, Rafael Teixeira de Lima, Yusik Kim, A Said Gurbuz, et al · 2025
Closest in time.
Texteller: An end-to-end formula recognition model
OleehyO · 2025
Closest in time.
Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations
Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, et al · 2025
Closest in time.
Surya: A lightweight document ocr and analysis toolkit
Vikas Paruchuri and Datalab Team · 2025
Closest in time.
olmocr: Unlocking trillions of tokens in pdfs with vision language models
Jake Poznanski, Jon Borchardt, Jason Dunkelberger, Regan Huff, Daniel Lin, Aman Rangapur, Christopher Wilhelm, Kyle Lo, and Luca Soldaini · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu · 2025
Closest in time.
Unstructured: Open-source etl for complex document transformation
Unstructured-IO · 2025
Closest in time.
Deepseek-ocr: Contexts optical compression
Haoran Wei, Yaofeng Sun, and Yukun Li · 2025
Closest in time.
Why lvlms are more prone to hallucinations in longer responses: The role of context
Ge Zheng, Jiaye Qian, Jiajin Tang, and Sibei Yang · 2025
Closest in time.