Fetching the paper…
Reading the bibliography…
Document parsing is essential for analyzing complex document structures and extracting fine-grained information, supporting numerous downstream applications.
Binary codes capable of correcting deletions, insertions and reversals
V.I. Levenshtein. 1966 · 1966
Earlier work this paper cites.
Ambiguity and constraint in mathematical expression recognition
Erik G Miller and Paul A Viola. 1998 · 1998
Earlier work this paper cites.
An overview of the tesseract ocr engine
R. Smith. 2007 · 2007
Earlier work this paper cites.
Layoutlmv2: Multi-modal pre-training for visually-rich document understanding
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou. 2022 · 2012
Earlier work this paper cites.
Deep textspotter: An end-to-end trainable scene text localization and recognition framework
Michal Bušta, Lukàš Neumann, and Jirí Matas. 2017 · 2017
Earlier work this paper cites.
Fully convolutional neural networks for page segmentation of historical document images
Christoph Wick and Frank Puppe. 2018 · 2018
Earlier work this paper cites.
Pattern generation strategies for improving recognition of handwritten mathematical expressions
Anh Duc Le, Bipin Indurkhya, and Masaki Nakagawa. 2019 · 2019
Earlier work this paper cites.
Rethinking table recognition using graph neural networks
Shah Rukh Qasim, Hassan Mahmood, and Faisal Shafait. 2019 · 2019
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020 · 2020
Earlier work this paper cites.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei. 2021 · 2021
Earlier work this paper cites.
Scene text telescope: Text-focused scene image super-resolution
Jingye Chen, Bin Li, and Xiangyang Xue. 2021 · 2021
Earlier work this paper cites.
Davit: Dual attention vision transformers
Mingyu Ding, Bin Xiao, Noel Codella, Ping Luo, Jingdong Wang, and Lu Yuan. 2022 · 2022
Cited alongside, same era.
Layoutlmv3: Pre-training for document ai with unified text and image masking (2022)
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022 · 2022
Cited alongside, same era.
When counting meets hmer: Counting-aware network for handwritten mathematical expression recognition
Bohan Li, Ye Yuan, Dingkang Liang, Xiao Liu, Zhilong Ji, Jinfeng Bai, Wenyu Liu, and Xiang Bai. 2022 · 2022
Cited alongside, same era.
Doclaynet: A large human-annotated dataset for document-layout segmentation
Birgit Pfitzmann, Christoph Auer, Michele Dolfi, Ahmed S Nassar, and Peter W J Staar. 2022 · 2022
Cited alongside, same era.
Florence-2: Advancing a unified representation for a variety of vision tasks (2023)
Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan. 2023 · 2023
Later among the works it cites.
pix2tex - latex ocr
Lukas Blecher. 2022 · 2024
Closest in time.
Openparse
Sergey Filimonov. 2024 · 2024
Closest in time.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024 · 2024
Closest in time.
Llamaparse
Jerry Liu. 2024 · 2024
Closest in time.
Visually guided generative text-layout pre-training for document intelligence
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hang Zhang, Yeyun Gong, Yelong Shen, Jiancheng Lv, Nan Duan, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 · 2023
Cited alongside, same era.
Nougat: Neural optical understanding for academic documents
Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic. 2023 · 2023
Cited alongside, same era.
Improving table structure recognition with visual-alignment sequential coordinate modeling
Yongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li, Zecheng Xie, Shenggao Zhu, Liangcai Gao, and Wei Peng. 2023 · 2023
Cited alongside, same era.
Structgpt: A general framework for large language model to reason over structured data
Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Wayne Xin Zhao, and Ji-Rong Wen. 2023 · 2023
Cited alongside, same era.
Rocketqav2: A joint training method for dense passage retrieval and passage re-ranking
Ruiyang Ren, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ji-Rong Wen. 2023 · 2023
Cited alongside, same era.
Unimernet: A universal network for real-world mathematical expression recognition
Bin Wang, Zhuangcheng Gu, Guang Liang, Chao Xu, Bo Zhang, Botian Shi, and Conghui He. 2024a
Cited in the paper.
Cdm: A reliable metric for fair and accurate formula recognition evaluation
Bin Wang, Fan Wu, Linke Ouyang, Zhuangcheng Gu, Rui Zhang, Renqiu Xia, Bo Zhang, and Conghui He. 2024b
Cited in the paper.
Zhiming Mao, Haoli Bai, Lu Hou, Jiansheng Wei, Xin Jiang, Qun Liu, and Kam-Fai Wong. 2024 · 2024
Closest in time.
Adocrnet: A deep learning ocr for arabic documents recognition
Lamia Mosbah, Ikram Moalla, Tarek M. Hamdani, Bilel Neji, Taha Beyrouthy, and Adel M. Alimi. 2024 · 2024
Closest in time.
General ocr theory: Towards ocr-2.0 via a unified end-to-end model
Haoran Wei, Chenglong Liu, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, Chunrui Han, and Xiangyu Zhang. 2024 · 2024
Closest in time.
Renqiu Xia, Song Mao, Xiangchao Yan, Hongbin Zhou, Bo Zhang, Haoyang Peng, Jiahao Pi, Daocheng Fu, Wenjie Wu, Hancheng Ye, et al. 2024 · 2024
Closest in time.
Deepdoc
Zhichang Yu. 2024 · 2024
Closest in time.