Fetching the paper…
Reading the bibliography…
We present DeepSeek-OCR as an initial investigation into the feasibility of compressing long contexts via optical 2D mapping.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Towards vqa models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Earlier work this paper cites.
Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark
J. Gu, X. Meng, G. Lu, L. Hou, N. Minzhe, X. Liang, L. Yao, R. Huang, W. Zhang, X. Jiang, et al · 2022
Earlier work this paper cites.
Opt-iml: Scaling language model instruction meta learning through the lens of generalization
S. Iyer, X. V. Lin, R. Pasunuru, T. Mihaylov, D. Simig, P. Yu, K. Shuster, T. Wang, Q. Liu, P. S. Koura, et al · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque · 2022
Earlier work this paper cites.
Nougat: Neural optical understanding for academic documents
L. Blecher, G. Cucurull, T. Scialom, and R. Stojnic · 2023
Cited alongside, same era.
Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al · 2023
Cited alongside, same era.
HAI-LLM: Efficient and lightweight training tool for large models, 2023
High-flyer · 2023
Cited alongside, same era.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Paddleocr 3.0 technical report
C. Cui, T. Sun, M. Lin, T. Gao, Y. Zhang, J. Liu, X. Wang, Z. Zhang, C. Zhou, H. Liu, et al · 2025
Closest in time.
Dolphin: Document image parsing via heterogeneous anchor prompting
H. Feng, S. Wei, X. Fei, W. Shi, Y. Han, L. Liao, J. Lu, B. Wu, Q. Liu, C. Lin, et al · 2025
Closest in time.
Monkeyocr: Document parsing with a structure-recognition-relation triplet paradigm
Z. Li, Y. Liu, Q. Liu, Z. Ma, Z. Zhang, S. Zhang, Z. Guo, J. Zhang, X. Wang, and X. Bai · 2025
Closest in time.
Smoldocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
A. Nassar, A. Marafioti, M. Omenetti, M. Lysak, N. Livathinos, C. Auer, L. Morin, R. T. de Lima, Y. Kim, A. S. Gurbuz, et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang · 2023
Cited alongside, same era.
Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding
Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y. Ma, C. Wu, B. Wang, et al · 2024
Cited alongside, same era.
URL https://github.com/chatdoc-com/OCRFlux
Ocrflux, 2025 · 2025
Cited alongside, same era.
Gemini 2.5-pro, 2025
G. AI · 2025
Cited alongside, same era.
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin · 2025
Cited alongside, same era.
URL https://github.com/datalab-to/marker
Marker
Cited in the paper.
URL https://mathpix.com/
Mathpix
Cited in the paper.
Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations
L. Ouyang, Y. Qu, H. Zhou, J. Zhu, R. Zhang, Q. Lin, B. Wang, Z. Zhao, M. Jiang, X. Zhao, et al · 2025
Closest in time.
olmocr: Unlocking trillions of tokens in pdfs with vision language models
J. Poznanski, A. Rangapur, J. Borchardt, J. Dunkelberger, R. Huff, D. Lin, C. Wilhelm, K. Lo, and L. Soldaini · 2025
Closest in time.
dots.ocr, 2025
Rednote · 2025
Closest in time.
Pp-doclayout: A unified document layout detection model to accelerate large-scale data construction
T. Sun, C. Cui, Y. Du, and Y. Liu · 2025
Closest in time.
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models
J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, H. Tian, Y. Duan, W. Su, J. Shao, et al · 2025
Closest in time.