Fetching the paper…
Reading the bibliography…
This technical report introduces PaddleOCR 3.0, an Apache-licensed open-source toolkit for OCR and document parsing.
Recursive xy cut using bounding boxes of connected components
J. Ha, R. M. Haralick, and I. T. Phillips · 1995
Earlier work this paper cites.
A survey of methods and strategies in character segmentation
R. Casey and E. Lecolinet · 1996
Earlier work this paper cites.
Optical Character Recognition
S. Mori, H. Nishida, and H. Yamada · 1999
Earlier work this paper cites.
I. J. Goodfellow, Y. Bulatov, J. Ibarz, S. Arnoud, and V. Shet · 2014
Earlier work this paper cites.
B. Shi, X. Bai, and C. Yao · 2015
Earlier work this paper cites.
Paddlepaddle: An open-source deep learning platform from industrial practice
Y. Ma, D. Yu, T. Wu, and H. Wang · 2019
Earlier work this paper cites.
Pp-ocr: A practical ultra lightweight ocr system
Y. Du, C. Li, R. Guo, X. Yin, W. Liu, J. Zhou, Y. Bai, Z. Yu, Y. Yang, Q. Dang, et al · 2020
Earlier work this paper cites.
Gtc: Guided training of ctc towards efficient and accurate scene text recognition
W. Hu, X. Cai, J. Hou, S. Yi, and Z. Lin · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
Pp-lcnet: A lightweight cpu convolutional neural network, 2021
C. Cui, T. Gao, S. Wei, Y. Du, R. Guo, S. Dong, B. Lu, Y. Zhou, X. Lv, Q. Liu, X. Hu, D. Yu, and Y. Ma · 2021
Earlier work this paper cites.
Pp-ocrv2: Bag of tricks for ultra lightweight ocr system
Y. Du, C. Li, R. Guo, C. Cui, W. Liu, J. Zhou, B. Lu, Y. Yang, Q. Liu, X. Hu, et al · 2021
Earlier work this paper cites.
Svtr: Scene text recognition with a single visual model
Y. Du, Z. Chen, C. Jia, X. Yin, T. Zheng, C. Li, Y. Du, and Y.-G. Jiang · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Nougat: Neural optical understanding for academic documents, 2023
L. Blecher, G. Cucurull, T. Scialom, and R. Stojnic · 2023
Cited alongside, same era.
Uvdoc: neural grid-based document unwarping
F. Verhoeven, T. Magne, and O. Sorkine-Hornung · 2023
Cited alongside, same era.
Z. Chen, W. Wang, Y. Cao, Y. Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al · 2024
Cited alongside, same era.
Mineru: An open-source solution for precise document content extraction
B. Wang, C. Xu, X. Zhao, L. Ouyang, F. Wu, Z. Zhao, R. Xu, K. Liu, Y. Qu, F. Shang, et al · 2024
Cited alongside, same era.
General ocr theory: Towards ocr-2.0 via a unified end-to-end model
H. Wei, C. Liu, J. Chen, J. Wang, L. Kong, Y. Xu, Z. Ge, L. Zhao, J. Sun, Y. Peng, et al · 2024
Cited alongside, same era.
Pp-formulanet: Bridging accuracy and efficiency in advanced formula recognition
H. Liu, C. Cui, Y. Du, Y. Liu, and G. Pan · 2025
Closest in time.
ONNX Runtime
Microsoft Corporation · 2025
Closest in time.
Smoldocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
A. Nassar, A. Marafioti, M. Omenetti, M. Lysak, N. Livathinos, C. Auer, L. Morin, R. T. de Lima, Y. Kim, A. S. Gurbuz, et al · 2025
Closest in time.
Pp-docbee: Improving multimodal document understanding through a bag of tricks
F. Ni, K. Huang, Y. Lu, W. Lv, G. Wang, Z. Chen, and Y. Liu · 2025
Closest in time.
TensorRT
NVIDIA Corporation · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2024
Cited alongside, same era.
Rolmocr: A faster, lighter open source ocr model, 2025
R. AI · 2025
Cited alongside, same era.
Ernie 4.5 technical report, 2025
Baidu-ERNIE-Team · 2025
Cited alongside, same era.
Pix2text
breezedeus · 2025
Cited alongside, same era.
open-parse
Filimoa · 2025
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Cited alongside, same era.
OpenVINO Toolkit
Intel Corporation · 2025
Cited alongside, same era.
NVIDIA Corporation · 2025
Closest in time.
Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations
L. Ouyang, Y. Qu, H. Zhou, J. Zhu, R. Zhang, Q. Lin, B. Wang, Z. Zhao, M. Jiang, X. Zhao, et al · 2025
Closest in time.
Ai studio
PaddlePaddle Team · 2025
Closest in time.
olmocr: Unlocking trillions of tokens in pdfs with vision language models
J. Poznanski, J. Borchardt, J. Dunkelberger, R. Huff, D. Lin, A. Rangapur, C. Wilhelm, K. Lo, and L. Soldaini · 2025
Closest in time.
Pp-doclayout: A unified document layout detection model to accelerate large-scale data construction
T. Sun, C. Cui, Y. Du, and Y. Liu · 2025
Closest in time.
unstructured
Unstructured-IO · 2025
Closest in time.
A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al · 2025
Closest in time.