Fetching the paper…
Reading the bibliography…
Document Structured Extraction (DSE) aims to extract structured content from raw documents.
An overview of the tesseract ocr engine
Ray Smith. 2007 · 2007
Earlier work this paper cites.
Icdar 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al. 2015 · 2015
Earlier work this paper cites.
Cermine: automatic extraction of structured metadata from scientific literature
Dominika Tkaczyk, Paweł Szostek, Mateusz Fedoryszak, Piotr Jan Dendek, and Łukasz Bolikowski. 2015 · 2015
Earlier work this paper cites.
Image-to-markup generation with coarse-to-fine attention
Yuntian Deng, Anssi Kanervisto, Jeffrey Ling, and Alexander M Rush. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Publaynet: largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. 2019 · 2019
Earlier work this paper cites.
Scirex: A challenge dataset for document-level information extraction
Sarthak Jain, Madeleine van Zuylen, Hannaneh Hajishirzi, and Iz Beltagy. 2020 · 2020
Earlier work this paper cites.
Docbank: A benchmark dataset for document layout analysis
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou. 2020 · 2020
Earlier work this paper cites.
Cord-19: The covid-19 open research dataset
Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Doug Burdick, Darrin Eide, Kathryn Funk, Yannis Katsis, Rodney Kinney, et al. 2020 · 2020
Earlier work this paper cites.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Earlier work this paper cites.
Image-based table recognition: data, model, and evaluation
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes. 2020 · 2020
Earlier work this paper cites.
Docformer: End-to-end transformer for document understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R Manmatha. 2021 · 2021
Earlier work this paper cites.
Tablex: a benchmark dataset for structure and content information extraction from scientific tables
Harsh Desai, Pratik Kayal, and Mayank Singh. 2021 · 2021
Earlier work this paper cites.
Layoutreader: Pre-training of text and layout for reading order detection
Zilong Wang, Yiheng Xu, Lei Cui, Jingbo Shang, and Furu Wei. 2021 · 2021
Earlier work this paper cites.
Multimodal tree decoder for table of contents extraction in document images
Pengfei Hu, Zhenrong Zhang, Jianshu Zhang, Jun Du, and Jiajia Wu. 2022 · 2022
Earlier work this paper cites.
Layoutlmv3: Pre-training for document ai with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Pp-structurev2: A stronger document analysis system
Chenxia Li, Ruoyu Guo, Jun Zhou, Mengtao An, Yuning Du, Lingfeng Zhu, Yi Liu, Xiaoguang Hu, and Dianhai Yu. 2022 · 2022
Cited alongside, same era.
Vila: Improving structured content extraction from scientific pdfs using visual layout groups
Zejiang Shen, Kyle Lo, Lucy Lu Wang, Bailey Kuehl, Daniel S Weld, and Doug Downey. 2022 · 2022
Cited alongside, same era.
Pubtables-1m: Towards comprehensive table extraction from unstructured documents
Brandon Smock, Rohith Pesala, and Robin Abraham. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Christoph Auer, Maksym Lysak, Ahmed Nassar, Michele Dolfi, Nikolaos Livathinos, Panos Vagenas, Cesar Berrospi Ramis, Matteo Omenetti, Fabian Lindlbauer, Kasper Dinkla, Lokesh Mishra, Yusik Kim, Shubham Gupta, Rafael Teixeira de Lima, Valery Weber, Lucas Morin, Ingmar Meijer, Viktor Kuropiatnyk, and Peter W. J. Staar. 2024 · 2024
Closest in time.
Pix2text (p2t)
Breezedeus. 2022 · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. 2024 · 2024
Closest in time.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, et al. 2024 · 2024
Closest in time.
Monkey: Image resolution and text label are important things for large multi-modal models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Nougat: Neural optical understanding for academic documents
Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic. 2023 · 2023
Cited alongside, same era.
Hao Feng, Qi Liu, Hao Liu, Wengang Zhou, Houqiang Li, and Can Huang. 2023 · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang. 2023 · 2023
Cited alongside, same era.
Large multilingual models pivot zero-shot multimodal learning across languages
Jinyi Hu, Yuan Yao, Chongyi Wang, Shan Wang, Yinxu Pan, Qianyu Chen, Tianyu Yu, Hanghao Wu, Yue Zhao, Haoye Zhang, Xu Han, Yankai Lin, Jiao Xue, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Cited alongside, same era.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023 · 2023
Cited alongside, same era.
Papermage: A unified toolkit for processing, representing, and manipulating visually-rich scientific documents
Kyle Lo, Zejiang Shen, Benjamin Newman, Joseph Z Chang, Russell Authur, Erin Bransom, Stefan Candra, Yoganand Chandrasekhar, Regan Huff, Bailey Kuehl, et al. 2023 · 2023
Cited alongside, same era.
Kosmos-2.5: A multimodal literate model
Tengchao Lv, Yupan Huang, Jingye Chen, Lei Cui, Shuming Ma, Yaoyao Chang, Shaohan Huang, Wenhui Wang, Li Dong, Weiyao Luo, et al. 2023 · 2023
Cited alongside, same era.
Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. 2024 · 2024
Closest in time.
Deepseek-vl: towards real-world vision-language understanding
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Yaofeng Sun, et al. 2024 · 2024
Closest in time.
Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations
Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, et al. 2024 · 2024
Closest in time.
Pymupdf4llm
PyMuPDF. 2024 · 2024
Closest in time.
pdfplumber
Jeremy Singer-Vine and The pdfplumber contributors. 2024 · 2024
Closest in time.
Renqiu Xia, Song Mao, Xiangchao Yan, Hongbin Zhou, Bo Zhang, Haoyang Peng, Jiahao Pi, Daocheng Fu, Wenjie Wu, Hancheng Ye, et al. 2024 · 2024
Closest in time.
Tabpedia: Towards comprehensive visual table understanding with concept synergy
Weichao Zhao, Hao Feng, Qi Liu, Jingqun Tang, Shu Wei, Binghong Wu, Lei Liao, Yongjie Ye, Hao Liu, Houqiang Li, et al. 2024 · 2024
Closest in time.
Mineru: A one-stop, open-source, high-quality data extraction tool
MinerU Contributors. 2024 · 2025
Closest in time.
Markitdown
microsoft. 2024 · 2025
Closest in time.
Marker: Convert pdf to markdown quickly with high accuracy
Vik Paruchuri and Samuel Lampa. 2023 · 2025
Closest in time.