Fetching the paper…
Reading the bibliography…
Document parsing (DP) transforms unstructured or semi-structured documents into structured, machine-readable representations, enabling downstream applications such as knowledge base construction and retrieval-augmented generation (RAG).
Tablebank: Table benchmark for image-based table detection and recognition. In Proceedings of the Twelfth Language Resources and Evaluation Conference . 1918–1925
Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, Ming Zhou, and Zhoujun Li. 2020a · 1925
Earlier work this paper cites.
Chartocr: Data extraction from charts images via a deep hybrid framework. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 1917–1925
Junyu Luo, Zekun Li, Jinpeng Wang, and Chin-Yew Lin. 2021 · 1925
Earlier work this paper cites.
Syntax-directed recognition of hand-printed two-dimensional mathematics. In Symposium on interactive systems for experimental applied mathematics: Proceedings of the Association for Computing Machinery Inc. Symposium . 436–459
Robert H Anderson. 1967 · 1967
Earlier work this paper cites.
User’s reference manual for the UW english/technical document image database III
Ihsin Tsaiyun Phillips. 1996 · 1996
Earlier work this paper cites.
Performance evaluation of document layout analysis algorithms on the UW data set. In Document Recognition IV , Vol. 3027. SPIE, 149–160
Jisheng Liang, Ihsin T Phillips, and Robert M Haralick. 1997 · 1997
Earlier work this paper cites.
Document structure analysis algorithms: a literature survey
Song Mao, Azriel Rosenfeld, and Tapas Kanungo. 2003 · 2003
Earlier work this paper cites.
ICDAR 2003 robust reading competitions: entries, results, and future directions
Simon M Lucas, Alex Panaretos, Luis Sosa, Anthony Tang, Shirley Wong, Robert Young, Kazuki Ashida, Hiroki Nagai, Masayuki Okamoto, Hiroaki Yamamoto, et al · 2005
Earlier work this paper cites.
A ground-truthed mathematical character and symbol image database. In Eighth International Conference on Document Analysis and Recognition (ICDAR’05) . IEEE, 675–679
Masakazu Suzuki, Seiichi Uchida, and Akihiro Nomura. 2005 · 2005
Earlier work this paper cites.
A realistic dataset for performance evaluation of document layout analysis. In 2009 10th International Conference on Document Analysis and Recognition . IEEE, 296–300
Apostolos Antonacopoulos, David Bridson, Christos Papadopoulos, and Stefan Pletschacher. 2009 · 2009
Earlier work this paper cites.
An open approach towards the benchmarking of table structure recognition systems. In Proceedings of the 9th IAPR International Workshop on Document Analysis Systems . 113–120
Asif Shahab, Faisal Shafait, Thomas Kieninger, and Andreas Dengel. 2010 · 2010
Earlier work this paper cites.
Automatic segmentation of subfigure image panels for multimodal biomedical document retrieval. In Document Recognition and Retrieval XVIII , Vol. 7874. SPIE, 294–304
Beibei Cheng, Sameer Antani, R Joe Stanley, and George R Thoma. 2011 · 2011
Earlier work this paper cites.
Revision: Automated classification, analysis and redesign of chart images. In Proceedings of the 24th annual ACM symposium on User interface software and technology . 393–402
Manolis Savva, Nicholas Kong, Arti Chhajta, Li Fei-Fei, Maneesh Agrawala, and Jeffrey Heer. 2011 · 2011
Earlier work this paper cites.
Dataset, ground-truth and performance metrics for table detection evaluation. In 2012 10th IAPR International Workshop on Document Analysis Systems . IEEE, 445–449
Jing Fang, Xin Tao, Zhi Tang, Ruiheng Qiu, and Ying Liu. 2012 · 2012
Earlier work this paper cites.
View: Visual information extraction widget for improving chart images accessibility. In 2012 19th IEEE international conference on image processing . IEEE, 2865–2868
Jinglun Gao, Yin Zhou, and Kenneth E Barner. 2012 · 2012
Earlier work this paper cites.
Performance evaluation of mathematical formula identification. In 2012 10th IAPR International Workshop on Document Analysis Systems . IEEE, 287–291
Xiaoyan Lin, Liangcai Gao, Zhi Tang, Xiaofan Lin, and Xuan Hu. 2012 · 2012
Earlier work this paper cites.
Scene text recognition using higher order language priors. In BMVC-British machine vision conference . BMVA
Anand Mishra, Karteek Alahari, and CV Jawahar. 2012 · 2012
Earlier work this paper cites.
CVC-UAB’s Participation in the Flowchart Recognition Task of CLEF-IP 2012.. In CLEF (Online Working Notes/Labs/Workshop)
Marçal Rusinol, Lluís-Pere de las Heras, Joan Mas, Oriol Ramos Terrades, Dimosthenis Karatzas, Anjan Dutta, Gemma Sánchez, and Josep Lladós. 2012 · 2012
Earlier work this paper cites.
End-to-end text recognition with convolutional neural networks. In Proceedings of the 21st international conference on pattern recognition (ICPR2012) . IEEE, 3304–3308
Tao Wang, David J Wu, Adam Coates, and Andrew Y Ng. 2012 · 2012
Earlier work this paper cites.
Detecting texts of arbitrary orientations in natural images. In 2012 IEEE conference on computer vision and pattern recognition . IEEE, 1083–1090
Cong Yao, Xiang Bai, Wenyu Liu, Yi Ma, and Zhuowen Tu. 2012 · 2012
Earlier work this paper cites.
Image retrieval from scientific publications: Text and image content processing to separate multipanel figures
Emilia Apostolova, Daekeun You, Zhiyun Xue, Sameer Antani, Dina Demner-Fushman, and George R Thoma. 2013 · 2013
Earlier work this paper cites.
Fusion of statistical and structural information for flowchart recognition. In 2013 12th International Conference on Document Analysis and Recognition . IEEE, 1210–1214
Céres Carton, Aurélie Lemaitre, and Bertrand Coüasnon. 2013 · 2013
Earlier work this paper cites.
ICDAR 2013 table competition. In 2013 12th international conference on document analysis and recognition . IEEE, 1449–1453
Max Göbel, Tamir Hassan, Ermelinda Oro, and Giorgio Orsi. 2013 · 2013
Earlier work this paper cites.
ICDAR 2013 robust reading competition. In 2013 12th international conference on document analysis and recognition . IEEE, 1484–1493
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras. 2013 · 2013
Earlier work this paper cites.
A framework for biomedical figure segmentation towards image-based document retrieval
Luis D Lopez, Jingyi Yu, Cecilia Arighi, Catalina O Tudor, Manabu Torii, Hongzhan Huang, K Vijay-Shanker, and Cathy Wu. 2013 · 2013
Earlier work this paper cites.
Automatic extraction of figures from scientific publications in high-energy physics
Piotr Adam Praczyk and Javier Nogueras-Iso. 2013 · 2013
Earlier work this paper cites.
Synthetic data and artificial neural networks for natural scene text recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
ICFHR 2014 competition on recognition of on-line handwritten mathematical expressions (CROHME 2014). In 2014 14th International Conference on Frontiers in Handwriting Recognition . IEEE, 791–796
Harold Mouchere, Christian Viard-Gaudin, Richard Zanibbi, and Utpal Garain. 2014 · 2014
Earlier work this paper cites.
A robust arbitrary text detection system for natural scene images
Anhar Risnumawan, Palaiahankote Shivakumara, Chee Seng Chan, and Chew Lim Tan. 2014 · 2014
Earlier work this paper cites.
Automated data extraction from scholarly line graphs. In Proc. Int. Workshop Graph. Recognit
Sagnik Ray Choudhury, Shuting Wang, Prasenjit Mitra, and C Lee Giles. 2015 · 2015
Earlier work this paper cites.
Densebox: Unifying landmark localization with end to end object detection
Lichao Huang, Yi Yang, Yafeng Deng, and Yinan Yu. 2015 · 2015
Earlier work this paper cites.
ICDAR 2015 competition on robust reading. In 2015 13th international conference on document analysis and recognition (ICDAR) . IEEE, 1156–1160
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al · 2015
Earlier work this paper cites.
Junction-based table detection in camera-captured document images
Wonkyo Seo, Hyung Il Koo, and Nam Ik Cho. 2015 · 2015
Earlier work this paper cites.
Scalable algorithms for scholarly figure mining and semantics. In Proceedings of the International Workshop on Semantic Big Data . 1–6
Sagnik Ray Choudhury, Shuting Wang, and C Lee Giles. 2016 · 2016
Earlier work this paper cites.
R-fcn: Object detection via region-based fully convolutional networks
Jifeng Dai, Yi Li, Kaiming He, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Synthetic data for text localisation in natural images. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2315–2324
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman. 2016 · 2016
Earlier work this paper cites.
A table detection method for pdf documents based on convolutional neural networks. In 2016 12th IAPR Workshop on Document Analysis Systems (DAS) . IEEE, 287–292
Leipeng Hao, Liangcai Gao, Xiaohan Yi, and Zhi Tang. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Reading text in the wild with convolutional neural networks
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2016 · 2016
Earlier work this paper cites.
BCE-Arabic-v1 dataset: Towards interpreting Arabic document images for people with visual impairments. In Proceedings of the 9th ACM International Conference on PErvasive Technologies Related to Assistive Environments . 1–8
Rana SM Saad, Randa I Elanwar, NS Abdel Kader, Samia Mashali, and Margrit Betke. 2016 · 2016
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao. 2016a · 2016
Earlier work this paper cites.
Figureseer: Parsing result-figures in research papers. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part VII 14 . Springer, 664–680
Noah Siegel, Zachary Horvitz, Roie Levin, Santosh Divvala, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
Deepchart: Combining deep convolutional networks and deep belief networks in chart classification
Binbin Tang, Xiao Liu, Jie Lei, Mingli Song, Dapeng Tao, Shuifa Sun, and Fangmin Dong. 2016 · 2016
Earlier work this paper cites.
Detecting text in natural image with connectionist text proposal network. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part VIII 14 . Springer, 56–72
Zhi Tian, Weilin Huang, Tong He, Pan He, and Yu Qiao. 2016 · 2016
Earlier work this paper cites.
Coco-text: Dataset and benchmark for text detection and recognition in natural images
Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge Belongie. 2016 · 2016
Earlier work this paper cites.
A machine learning approach for semantic structuring of scientific charts in scholarly documents. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 31. 4644–4649
Rabah Al-Zaidy and C Giles. 2017 · 2017
Earlier work this paper cites.
Deep textspotter: An end-to-end trainable scene text localization and recognition framework. In Proceedings of the IEEE international conference on computer vision . 2204–2212
Michal Busta, Lukas Neumann, and Jiri Matas. 2017 · 2017
Earlier work this paper cites.
Total-text: A comprehensive dataset for scene text detection and recognition. In 2017 14th IAPR international conference on document analysis and recognition (ICDAR) , Vol. 1. IEEE, 935–942
Chee Kheng Ch’ng and Chee Seng Chan. 2017 · 2017
Earlier work this paper cites.
Scatteract: Automated extraction of data from scatter plots. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2017, Skopje, Macedonia, September 18–22, 2017, Proceedings, Part I 10 . Springer, 135–150
Mathieu Cliche, David Rosenberg, Dhruv Madeka, and Connie Yee. 2017 · 2017
Earlier work this paper cites.
generation with coarse-to-fine attention. In International Conference on Machine Learning . PMLR, 980–989
Yuntian Deng, Anssi Kanervisto, Jeffrey Ling, and Alexander M Rush. 2017 · 2017
Earlier work this paper cites.
Intercoder reliability and validity of WebPlotDigitizer in extracting graphed data
Daniel Drevon, Sophie R Fursa, and Allura L Malcolm. 2017 · 2017
Earlier work this paper cites.
ICDAR2017 competition on page object detection. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR) , Vol. 1. IEEE, 1417–1422
Liangcai Gao, Xiaohan Yi, Zhuoren Jiang, Leipeng Hao, and Zhi Tang. 2017a · 2017
Earlier work this paper cites.
A deep learning-based formula detection method for PDF documents. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR) , Vol. 1. IEEE, 553–558
Liangcai Gao, Xiaohan Yi, Yuan Liao, Zhuoren Jiang, Zuoyu Yan, and Zhi Tang. 2017b · 2017
Earlier work this paper cites.
Table detection using deep learning. In 2017 14th IAPR international conference on document analysis and recognition (ICDAR) , Vol. 1. IEEE, 771–776
Azka Gilani, Shah Rukh Qasim, Imran Malik, and Faisal Shafait. 2017 · 2017
Earlier work this paper cites.
The MUSCIMA++ dataset for handwritten optical music recognition. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR) , Vol. 1. IEEE, 39–46
Jan Hajič and Pavel Pecina. 2017 · 2017
Earlier work this paper cites.
Chartsense: Interactive data extraction from chart images. In Proceedings of the 2017 chi conference on human factors in computing systems . 6706–6717
Daekyoung Jung, Wonjae Kim, Hyunjoo Song, Jeong-in Hwang, Bongshin Lee, Bohyoung Kim, and Jinwook Seo. 2017 · 2017
Earlier work this paper cites.
Towards end-to-end text spotting with convolutional recurrent neural networks. In Proceedings of the IEEE international conference on computer vision . 5238–5246
Hui Li, Peng Wang, and Chunhua Shen. 2017 · 2017
Earlier work this paper cites.
Textboxes: A fast text detector with a single deep neural network. In Proceedings of the AAAI conference on artificial intelligence , Vol. 31
Minghui Liao, Baoguang Shi, Xiang Bai, Xinggang Wang, and Wenyu Liu. 2017 · 2017
Earlier work this paper cites.
Fast CNN-based document layout analysis. In 2017 IEEE International Conference on Computer Vision Workshops (ICCVW) . IEEE, 1173–1180
Dario Augusto Borges Oliveira and Matheus Palhares Viana. 2017 · 2017
Earlier work this paper cites.
Reverse-engineering visualizations: Recovering visual encodings from chart images. In Computer graphics forum , Vol. 36. Wiley Online Library, 353–363
Jorge Poco and Jeffrey Heer. 2017 · 2017
Earlier work this paper cites.
A study on optical character recognition techniques
Narendra Sahu and Manoj Sonkusare. 2017 · 2017
Earlier work this paper cites.
Deepdesrt: Deep learning for detection and structure recognition of tables in document images. In 2017 14th IAPR international conference on document analysis and recognition (ICDAR) , Vol. 1. IEEE, 1162–1167
Sebastian Schreiber, Stefan Agne, Ivo Wolf, Andreas Dengel, and Sheraz Ahmed. 2017 · 2017
Earlier work this paper cites.
Icdar2017 competition on layout analysis for challenging medieval manuscripts. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR) , Vol. 1. IEEE, 1361–1370
Fotini Simistira, Manuel Bouillon, Mathias Seuret, Marcel Würsch, Michele Alberti, Rolf Ingold, and Marcus Liwicki. 2017 · 2017
Earlier work this paper cites.
Learning to extract semantic structure from documents using multimodal fully convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5315–5324
Xiao Yang, Ersin Yumer, Paul Asente, Mike Kraley, Daniel Kifer, and C Lee Giles. 2017 · 2017
Earlier work this paper cites.
CNN based page object detection in document images. In 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR) , Vol. 1. IEEE, 230–235
Xiaohan Yi, Liangcai Gao, Yuan Liao, Xiaode Zhang, Runtao Liu, and Zhuoren Jiang. 2017 · 2017
Earlier work this paper cites.
Deeptext: A new approach for text proposal generation and text detection in natural images. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 1208–1212
Zhuoyao Zhong, Lianwen Jin, and Shuangping Huang. 2017 · 2017
Earlier work this paper cites.
East: an efficient and accurate scene text detector. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition . 5551–5560
Xinyu Zhou, Cong Yao, He Wen, Yuzhi Wang, Shuchang Zhou, Weiran He, and Jiajun Liang. 2017 · 2017
Earlier work this paper cites.
Evaluation of convolutional neural network architectures for chart image classification. In 2018 International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–8
Paulo Chagas, Rafael Akiyama, Aruanda Meiguins, Carlos Santos, Filipe Saraiva, Bianchi Meiguins, and Jefferson Morais. 2018 · 2018
Earlier work this paper cites.
Aon: Towards arbitrarily-oriented text recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5571–5579
Zhanzhan Cheng, Yangliu Xu, Fan Bai, Yi Niu, Shiliang Pu, and Shuigeng Zhou. 2018 · 2018
Earlier work this paper cites.
Chart decoder: Generating textual and numeric information from chart images automatically
Wenjing Dai, Meng Wang, Zhibin Niu, and Jiawan Zhang. 2018 · 2018
Earlier work this paper cites.
Pixellink: Detecting scene text via instance segmentation. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32
Dan Deng, Haifeng Liu, Xuelong Li, and Deng Cai. 2018 · 2018
Earlier work this paper cites.
R 2 cnn: Rotational region cnn for arbitrarily-oriented scene text detection. In 2018 24th International conference on pattern recognition (ICPR) . IEEE, 3610–3615
Yingying Jiang, Xiangyu Zhu, Xiaobing Wang, Shuli Yang, Wei Li, Hua Wang, Pei Fu, and Zhenbo Luo. 2018 · 2018
Earlier work this paper cites.
Chargrid: Towards understanding 2d documents
Anoop Raveendra Katti, Christian Reisswig, Cordula Guder, Sebastian Brarda, Steffen Bickel, Johannes Höhne, and Jean Baptiste Faddoul. 2018 · 2018
Earlier work this paper cites.
Page object detection from pdf document images by deep structured prediction and supervised clustering. In 2018 24th International Conference on Pattern Recognition (ICPR) . IEEE, 3627–3632
Xiao-Hui Li, Fei Yin, and Cheng-Lin Liu. 2018 · 2018
Earlier work this paper cites.
Textboxes++: A single-shot oriented scene text detector
Minghui Liao, Baoguang Shi, and Xiang Bai. 2018 · 2018
Earlier work this paper cites.
Fots: Fast oriented text spotting with a unified network. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5676–5685
Xuebo Liu, Ding Liang, Shi Yan, Dagui Chen, Yu Qiao, and Junjie Yan. 2018 · 2018
Earlier work this paper cites.
Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes. In Proceedings of the European conference on computer vision (ECCV) . 67–83
Pengyuan Lyu, Minghui Liao, Cong Yao, Wenhao Wu, and Xiang Bai. 2018 · 2018
Earlier work this paper cites.
Arbitrary-oriented scene text detection via rotation proposals
Jianqi Ma, Weiyuan Shao, Hao Ye, Li Wang, Hong Wang, Yingbin Zheng, and Xiangyang Xue. 2018 · 2018
Earlier work this paper cites.
Multi-task handwritten document layout analysis
Lorenzo Quirós. 2018 · 2018
Earlier work this paper cites.
Decnt: Deep deformable cnn for table detection
Shoaib Ahmed Siddiqui, Muhammad Imran Malik, Stefan Agne, Andreas Dengel, and Sheraz Ahmed. 2018 · 2018
Earlier work this paper cites.
Extracting scientific figures with distantly supervised neural networks. In Proceedings of the 18th ACM/IEEE on joint conference on digital libraries . 223–232
Noah Siegel, Nicholas Lourie, Russell Power, and Waleed Ammar. 2018 · 2018
Earlier work this paper cites.
Deepscores-a dataset for segmentation, detection and classification of tiny objects. In 2018 24th International Conference on Pattern Recognition (ICPR) . IEEE, 3704–3709
Lukas Tuggener, Ismail Elezi, Jurgen Schmidhuber, Marcello Pelillo, and Thilo Stadelmann. 2018 · 2018
Earlier work this paper cites.
Fully convolutional neural networks for page segmentation of historical document images. In 2018 13th IAPR International Workshop on Document Analysis Systems (DAS) . IEEE, 287–292
Christoph Wick and Frank Puppe. 2018 · 2018
Earlier work this paper cites.
Qiangpeng Yang, Mengli Cheng, Wenmeng Zhou, Yan Chen, Minghui Qiu, Wei Lin, and Wei Chu. 2018 · 2018
Earlier work this paper cites.
ESIR: End-to-end Scene Text Recognition via Iterative Rectification
Fangneng Zhan and Shijian Lu. 2018 · 2018
Earlier work this paper cites.
Multi-scale attention with dense encoder for handwritten mathematical expression recognition. In 2018 24th international conference on pattern recognition (ICPR) . IEEE, 2245–2250
Jianshu Zhang, Jun Du, and Lirong Dai. 2018 · 2018
Earlier work this paper cites.
Character region awareness for text detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 9365–9374
Youngmin Baek, Bado Lee, Dongyoon Han, Sangdoo Yun, and Hwalsuk Lee. 2019 · 2019
Earlier work this paper cites.
Document layout analysis: a comprehensive survey
Galal M Binmakhashen and Sabri A Mahmoud. 2019 · 2019
Earlier work this paper cites.
Complicated table structure recognition
Zewen Chi, Heyan Huang, Heng-Da Xu, Houjin Yu, Wanxuan Yin, and Xian-Ling Mao. 2019 · 2019
Earlier work this paper cites.
ICDAR 2019 competition on harvesting raw tables from infographics (chart-infographics). In 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 1594–1599
Kenny Davila, Bhargava Urala Kota, Srirangaraj Setlur, Venu Govindaraju, Christopher Tensmeyer, Sumit Shekhar, and Ritwick Chaudhry. 2019 · 2019
Earlier work this paper cites.
Challenges in end-to-end neural scientific table recognition. In 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 894–901
Yuntian Deng, David Rosenberg, and Gideon Mann. 2019 · 2019
Earlier work this paper cites.
Bertgrid: Contextualized embedding for 2d document representation and understanding
Timo I Denk and Christian Reisswig. 2019 · 2019
Earlier work this paper cites.
Textdragon: An end-to-end framework for arbitrary shaped text spotting. In Proceedings of the IEEE/CVF international conference on computer vision . 9076–9085
Wei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang, and Cheng-Lin Liu. 2019 · 2019
Earlier work this paper cites.
ICDAR 2019 competition on table detection and recognition (cTDaR). In 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 1510–1515
Liangcai Gao, Yilun Huang, Hervé Déjean, Jean-Luc Meunier, Qinqin Yan, Yu Fang, Florian Kleber, and Eva Lang. 2019 · 2019
Earlier work this paper cites.
A two-stage method for text line detection in historical documents
Tobias Grüning, Gundram Leifert, Tobias Strauß, Johannes Michael, and Roger Labahn. 2019 · 2019
Earlier work this paper cites.
A YOLO-based table detection method. In 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 813–818
Yilun Huang, Qinqin Yan, Yibo Li, Yifan Chen, Xiong Wang, Liangcai Gao, and Zhi Tang. 2019b · 2019
Earlier work this paper cites.
Icdar2019 competition on scanned receipt ocr and information extraction. In 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 1516–1520
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar. 2019a · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents. In 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW) , Vol. 2. IEEE, 1–6
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Earlier work this paper cites.
Docfigure: A dataset for scientific document figure classification. In 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW) , Vol. 1. IEEE, 74–79
KV Jobin, Ajoy Mondal, and CV Jawahar. 2019 · 2019
Earlier work this paper cites.
Table structure extraction with bi-directional gated recurrent unit networks. In 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 1366–1371
Saqib Ali Khan, Syed Muhammad Daniyal Khalid, Muhammad Ali Shahzad, and Faisal Shafait. 2019 · 2019
Cited alongside, same era.
DECO: A dataset of annotated spreadsheets for layout and table recognition. In 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 1280–1285
Elvis Koci, Maik Thiele, Josephine Rehak, Oscar Romero, and Wolfgang Lehner. 2019 · 2019
Cited alongside, same era.
Pattern generation strategies for improving recognition of handwritten mathematical expressions
Anh Duc Le, Bipin Indurkhya, and Masaki Nakagawa. 2019 · 2019
Cited alongside, same era.
Show, attend and read: A simple and strong baseline for irregular text recognition. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 8610–8617
Hui Li, Peng Wang, Chunhua Shen, and Guyu Zhang. 2019 · 2019
Cited alongside, same era.
Segmentation for document layout analysis: not dead yet
Logan Markewich, Hao Zhang, Yubin Xing, Navid Lambert-Shirzad, Zhexin Jiang, Roy Ka-Wei Lee, Zhi Li, and Seok-Bum Ko. 2022 · 2022
Later among the works it cites.
Continual learning for table detection in document images
Mohammad Minouei, Khurram Azeem Hashmi, Mohammad Reza Soheili, Muhammad Zeshan Afzal, and Didier Stricker. 2022 · 2022
Later among the works it cites.
Tableformer: Table structure understanding with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4614–4623
Ahmed Nassar, Nikolaos Livathinos, Maksym Lysak, and Peter Staar. 2022 · 2022
Later among the works it cites.
TableSegNet: a fully convolutional network for table detection and segmentation in document images
Duc-Dung Nguyen. 2022 · 2022
Later among the works it cites.
Pagenet: Towards end-to-end weakly supervised page-level handwritten chinese text recognition
Dezhi Peng, Lianwen Jin, Yuliang Liu, Canjie Luo, and Songxuan Lai. 2022a · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scene text recognition from two-dimensional perspective. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 8714–8721
Minghui Liao, Jian Zhang, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lyu, Cong Yao, and Xiang Bai. 2019 · 2019
Cited alongside, same era.
Graph convolution for multimodal information extraction from visually rich documents
Xiaojing Liu, Feiyu Gao, Qiong Zhang, and Huasha Zhao. 2019a · 2019
Cited alongside, same era.
Yuliang Liu, Tong He, Hao Chen, Xinyu Wang, Canjie Luo, Shuaitao Zhang, Chunhua Shen, and Lianwen Jin. 2019b · 2019
Cited alongside, same era.
Curved scene text detection via transverse and longitudinal sequence connection
Yuliang Liu, Lianwen Jin, Shuaitao Zhang, Canjie Luo, and Sheng Zhang. 2019c · 2019
Cited alongside, same era.
Moran: A multi-object rectified attention network for scene text recognition
Canjie Luo, Lianwen Jin, and Zenghui Sun. 2019 · 2019
Cited alongside, same era.
ICDAR 2019 CROHME+ TFD: Competition on recognition of handwritten mathematical expressions and typeset formula detection. In 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 1533–1538
Mahshad Mahdavi, Richard Zanibbi, Harold Mouchere, Christian Viard-Gaudin, and Utpal Garain. 2019 · 2019
Cited alongside, same era.
Detecting mathematical expressions in scientific document images using a u-net trained on a diverse dataset
Wataru Ohyama, Masakazu Suzuki, and Seiichi Uchida. 2019 · 2019
Cited alongside, same era.
Tablenet: Deep learning model for end-to-end table detection and tabular data extraction from scanned document images. In 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 128–133
Shubham Singh Paliwal, D Vishwanath, Rohit Rahul, Monika Sharma, and Lovekesh Vig. 2019 · 2019
Cited alongside, same era.
Doclaynet: A large humanannotated dataset for document-layout analysis (2022)
B Pfitzmann, C Auer, M Dolfi, AS Nassar, and PWJ Staar. [n. d.] · 2022
Later among the works it cites.
Visual understanding of complex table structures from document images. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 2299–2308
Sachin Raja, Ajoy Mondal, and CV Jawahar. 2022 · 2022
Later among the works it cites.
Glass: Global to local attention for scene-text spotting. In European Conference on Computer Vision . Springer, 249–266
Roi Ronen, Shahar Tsiper, Oron Anschel, Inbal Lavi, Amir Markovitz, and R Manmatha. 2022 · 2022
Later among the works it cites.
FormulaNet: A benchmark dataset for mathematical formula detection
Felix M Schmitt-Koopmann, Elaine M Huang, Hans-Peter Hutter, Thilo Stadelmann, and Alireza Darvishy. 2022 · 2022
Later among the works it cites.
PubTables-1M: Towards comprehensive table extraction from unstructured documents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4634–4642
Brandon Smock, Rohith Pesala, and Robin Abraham. 2022 · 2022
Later among the works it cites.
Vision-language pre-training for boosting scene text detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15681–15691
Sibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang, Wenqing Cheng, Xiang Bai, and Cong Yao. 2022 · 2022
Later among the works it cites.
FR-DETR: End-to-end flowchart recognition with precision and robustness
Lianshan Sun, Hanchao Du, and Tao Hou. 2022 · 2022
Later among the works it cites.
Few could be better than all: Feature sampling and grouping for scene text detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 4563–4572
Jingqun Tang, Wenqing Zhang, Hongye Liu, MingKun Yang, Bo Jiang, Guanglong Hu, and Xiang Bai. 2022 · 2022
Later among the works it cites.
A new ChEMBL dataset for the similarity-based target fishing engine FastTargetPred: Annotation of an exhaustive list of linear tetrapeptides
Shivalika Tanwar, Patrick Auberger, Germain Gillet, Mario DiPaola, Katya Tsaioun, and Bruno O Villoutreix. 2022 · 2022
Later among the works it cites.
Toward understanding wordart: Corner-guided transformer for scene text recognition. In European conference on computer vision . Springer, 303–321
Xudong Xie, Ling Fu, Zhifei Zhang, Zhaowen Wang, and Xiang Bai. 2022 · 2022
Later among the works it cites.
Syntax-aware network for handwritten mathematical expression recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 4553–4562
Ye Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu, Zhilong Ji, Zhongqin Wu, and Xiang Bai. 2022 · 2022
Later among the works it cites.
Split, embed and merge: An accurate table structure recognizer
Zhenrong Zhang, Jianshu Zhang, Jun Du, and Fengren Wang. 2022b · 2022
Later among the works it cites.
Comer: Modeling coverage for transformer-based handwritten mathematical expression recognition. In European conference on computer vision . Springer, 392–408
Wenqi Zhao and Liangcai Gao. 2022 · 2022
Later among the works it cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 · 2023
Later among the works it cites.
Nougat: Neural optical understanding for academic documents
Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic. 2023 · 2023
Later among the works it cites.
Llava-interactive: An all-in-one demo for image chat, segmentation, generation and editing
Wei-Ge Chen, Irina Spiridonova, Jianwei Yang, Jianfeng Gao, and Chunyuan Li. 2023 · 2023
Later among the works it cites.
M6doc: A large-scale multi-format, multi-type, multi-layout, multi-language, multi-annotation category dataset for modern document layout analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15138–15147
Hiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang, Qiyuan Zhu, Zecheng Xie, Jing Li, Kai Ding, and Lianwen Jin. 2023 · 2023
Later among the works it cites.
Vision grid transformer for document layout analysis. In Proceedings of the IEEE/CVF international conference on computer vision . 19462–19472
Cheng Da, Chuwei Luo, Qi Zheng, and Cong Yao. 2023 · 2023
Later among the works it cites.
A survey and approach to chart classification. In International Conference on Document Analysis and Recognition . Springer, 67–82
Anurag Dhote, Mohammed Javed, and David S Doermann. 2023 · 2023
Later among the works it cites.
Context perception parallel decoder for scene text recognition
Yongkun Du, Zhineng Chen, Caiyan Jia, Xiaoting Yin, Chenxia Li, Yuning Du, and Yu-Gang Jiang. 2023 · 2023
Later among the works it cites.
Hao Feng, Qi Liu, Hao Liu, Wengang Zhou, Houqiang Li, and Can Huang. 2023 · 2023
Later among the works it cites.
Chartllama: A multimodal llm for chart understanding and generation
Yucheng Han, Chi Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, and Hanwang Zhang. 2023 · 2023
Later among the works it cites.
Lineex: data extraction from scientific line charts. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 6213–6221
Muhammad Yusuf Hassan, Mayank Singh, et al · 2023
Later among the works it cites.
Improving table structure recognition with visual-alignment sequential coordinate modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11134–11143
Yongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li, Zecheng Xie, Shenggao Zhu, Liangcai Gao, and Wei Peng. 2023 · 2023
Later among the works it cites.
Tables to LaTeX: structure and content extraction from scientific tables
Pratik Kayal, Mrinal Anand, Harsh Desai, and Mayank Singh. 2023 · 2023
Later among the works it cites.
Recent trends in mathematical expressions recognition: An LDA-based analysis
Vinay Kukreja et al · 2023
Later among the works it cites.
Trocr: Transformer-based optical character recognition with pre-trained models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 13094–13102
Minghao Li, Tengchao Lv, Jingye Chen, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, and Furu Wei. 2023 · 2023
Later among the works it cites.
Llava-plus: Learning to use tools for creating multimodal agents
Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, et al · 2023
Later among the works it cites.
Spts v2: single-point scene text spotting
Yuliang Liu, Jiaxin Zhang, Dezhi Peng, Mingxin Huang, Xinyu Wang, Jingqun Tang, Can Huang, Dahua Lin, Chunhua Shen, Xiang Bai, et al · 2023
Later among the works it cites.
Geolayoutlm: Geometric pre-training for visual information extraction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 7092–7101
Chuwei Luo, Changxu Cheng, Qi Zheng, and Cong Yao. 2023 · 2023
Later among the works it cites.
Rethinking image-based table recognition using weakly supervised methods
Nam Tuan Ly, Atsuhiro Takasu, Phuc Nguyen, and Hideaki Takeda. 2023 · 2023
Later among the works it cites.
Robust table detection and structure recognition from heterogeneous document images
Chixiang Ma, Weihong Lin, Lei Sun, and Qiang Huo. 2023 · 2023
Later among the works it cites.
ChartEye: A Deep Learning Framework for Chart Information Extraction. In 2023 International Conference on Digital Image Computing: Techniques and Applications (DICTA) . IEEE, 554–561
Osama Mustafa, Muhammad Khizer Ali, Momina Moetesum, and Imran Siddiqi. 2023 · 2023
Later among the works it cites.
Formerge: Recover spanning cells in complex table structure using transformer network. In International Conference on Document Analysis and Recognition . Springer, 522–534
Nam Quan Nguyen, Anh Duy Le, Anh Khoa Lu, Xuan Toan Mai, and Tuan Anh Tran. 2023 · 2023
Later among the works it cites.
Structure Diagram Recognition in Financial Announcements. In International Conference on Document Analysis and Recognition . Springer, 20–44
Meixuan Qiao, Jun Wang, Junfu Xiang, Qiyu Hou, and Ruixuan Li. 2023 · 2023
Later among the works it cites.
Text-DIAE: a self-supervised degradation invariant autoencoder for text recognition and document enhancement. In proceedings of the AAAI conference on artificial intelligence , Vol. 37. 2330–2338
Mohamed Ali Souibgui, Sanket Biswas, Andres Mafla, Ali Furkan Biten, Alicia Fornés, Yousri Kessentini, Josep Lladós, Lluis Gomez, and Dimosthenis Karatzas. 2023 · 2023
Later among the works it cites.
Contextual transformer sequence-based recognition network for medical examination reports
Honglin Wan, Zongfeng Zhong, Tianping Li, Huaxiang Zhang, and Jiande Sun. 2023 · 2023
Later among the works it cites.
DocLLM: A layout-aware generative language model for multimodal document understanding
Dongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, and Xiaomo Liu. 2023c · 2023
Later among the works it cites.
Robust table structure recognition with dynamic queries enhanced detection transformer
Jiawei Wang, Weihong Lin, Chixiang Ma, Mingze Li, Zheng Sun, Lei Sun, and Qiang Huo. 2023b · 2023
Later among the works it cites.
Vary: Scaling up the vision vocabulary for large vision-language models
Haoran Wei, Lingyu Kong, Jinyue Chen, Liang Zhao, Zheng Ge, Jinrong Yang, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. 2023 · 2023
Later among the works it cites.
Structchart: Perception, structuring, reasoning for visual chart understanding
Renqiu Xia, Bo Zhang, Haoyang Peng, Hancheng Ye, Xiangchao Yan, Peng Ye, Botian Shi, Yu Qiao, and Junchi Yan. 2023 · 2023
Later among the works it cites.
Table detection for visually rich document images
Bin Xiao, Murat Simsek, Burak Kantarci, and Ala Abu Alkheir. 2023 · 2023
Later among the works it cites.
Chartdetr: A multi-shape detection network for visual chart recognition
Wenyuan Xue, Dapeng Chen, Baosheng Yu, Yifei Chen, Sai Zhou, and Wei Peng. 2023 · 2023
Later among the works it cites.
A large-scale dataset for end-to-end table recognition in the wild
Fan Yang, Lei Hu, Xinwu Liu, Shuangping Huang, and Zhenghui Gu. 2023 · 2023
Later among the works it cites.
Docxchain: A powerful open-source toolchain for document parsing and beyond
Cong Yao. 2023 · 2023
Later among the works it cites.
Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Guohai Xu, Chenliang Li, Junfeng Tian, Qi Qian, Ji Zhang, et al · 2023
Later among the works it cites.
Turning a clip model into a scene text detector. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6978–6988
Wenwen Yu, Yuliang Liu, Wei Hua, Deqiang Jiang, Bo Ren, and Xiang Bai. 2023 · 2023
Later among the works it cites.
Arbitrary shape text detection via boundary transformer
Shi-Xue Zhang, Chun Yang, Xiaobin Zhu, and Xu-Cheng Yin. 2023b · 2023
Later among the works it cites.
Table Structure Recognition Method Based on Lightweight Network and Channel Attention
Tao Zhang, Yi Sui, Shunyao Wu, Fengjing Shao, and Rencheng Sun. 2023a · 2023
Later among the works it cites.
Abdelrahman Abdallah, Daniel Eberharter, Zoe Pfister, and Adam Jatowt. 2024 · 2024
Closest in time.
Optimum Deep Learning Method for Document Layout Analysis in Low Resource Languages. In Proceedings of the 2024 ACM Southeast Conference . 199–204
Md Mutasim Billah Abu Noman Akanda, Maruf Ahmed, AKM Shahariar Azad Rabby, and Fuad Rahman. 2024 · 2024
Closest in time.
SemiDocSeg: harnessing semi-supervised learning for document layout analysis
Ayan Banerjee, Sanket Biswas, Josep Lladós, and Umapada Pal. 2024 · 2024
Closest in time.
pix2tex - LaTeX OCR
Lukas Blecher. 2022 · 2024
Closest in time.
OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
Jinyue Chen, Lingyu Kong, Haoran Wei, Chenglong Liu, Zheng Ge, Liang Zhao, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. 2024a · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al · 2024
Closest in time.
Swin-chart: An efficient approach for chart classification
Anurag Dhote, Mohammed Javed, and David S Doermann. 2024 · 2024
Closest in time.
Improved manta ray foraging optimizer-based SVM for feature selection problems: a medical case study
Adel Got, Djaafar Zouache, Abdelouahab Moussaoui, Laith Abualigah, and Ahmed Alsayat. 2024 · 2024
Closest in time.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, et al · 2024
Closest in time.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, et al · 2024
Closest in time.
mplug-docowl2: High-resolution compressing for ocr-free multi-page document understanding
Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye, Ming Yan, Ji Zhang, Qin Jin, Fei Huang, and Jingren Zhou. 2024c · 2024
Closest in time.
Mathematical formula detection in document images: A new dataset and a new approach
Kai Hu, Zhuoyao Zhong, Lei Sun, and Qiang Huo. 2024d · 2024
Closest in time.
Readoc: A unified benchmark for realistic document structured extraction
Zichao Li, Aizier Abulaiti, Yaojie Lu, Xuanang Chen, Jia Zheng, Hongyu Lin, Xianpei Han, and Le Sun. 2024 · 2024
Closest in time.
Revolutionizing retrieval-augmented generation with enhanced PDF structure recognition
Demiao Lin. 2024 · 2024
Closest in time.
Focus Anywhere for Fine-grained Multi-page Document Understanding
Chenglong Liu, Haoran Wei, Jinyue Chen, Lingyu Kong, Zheng Ge, Zining Zhu, Liang Zhao, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. 2024b · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge (January 2024)
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. [n. d.] · 2024
Closest in time.
Textmonkey: An ocr-free large multimodal model for understanding document
Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai. 2024c · 2024
Closest in time.
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15630–15640
Chuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng, Zhi Yu, and Cong Yao. 2024 · 2024
Closest in time.
ADOCRNet: A Deep Learning OCR for Arabic Documents Recognition
Lamia Mosbah, Ikram Moalla, Tarek M Hamdani, Bilel Neji, Taha Beyrouthy, and Adel M Alimi. 2024 · 2024
Closest in time.
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu, Rui Zhang, Qunshu Lin, Bin Wang, Zhiyuan Zhao, Man Jiang, Xiaomeng Zhao, et al · 2024
Closest in time.
Machine learning and non-machine learning methods in mathematical recognition systems: Two decades’ systematic literature review
Sakshi and Vinay Kukreja. 2024 · 2024
Closest in time.
C2F-CHART: A Curriculum Learning Approach to Chart Classification
Nour Shaheen, Tamer Elsharnouby, and Marwan Torki. 2024 · 2024
Closest in time.
LOCR: Location-Guided Transformer for Optical Character Recognition
Yu Sun, Dongzhan Zhou, Chen Lin, Conghui He, Wanli Ouyang, and Han-Sen Zhong. 2024 · 2024
Closest in time.
Chart classification: a survey and benchmarking of different state-of-the-art methods
Jennil Thiyam, Sanasam Ranbir Singh, and Prabin Kumar Bora. 2024 · 2024
Closest in time.
OmniParser: A Unified Framework for Text Spotting Key Information Extraction and Table Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15641–15653
Jianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu, Wenqing Cheng, Fei Huang, Xiang Bai, Cong Yao, and Zhibo Yang. 2024 · 2024
Closest in time.
UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition
Bin Wang, Zhuangcheng Gu, Chao Xu, Bo Zhang, Botian Shi, and Conghui He. 2024b · 2024
Closest in time.
Cdm: A reliable metric for fair and accurate formula recognition evaluation
Bin Wang, Fan Wu, Linke Ouyang, Zhuangcheng Gu, Rui Zhang, Renqiu Xia, Bo Zhang, and Conghui He. 2024d · 2024
Closest in time.
Jiawei Wang, Kai Hu, Zhuoyao Zhong, Lei Sun, and Qiang Huo. 2024c · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al · 2024
Closest in time.
General ocr theory: Towards ocr-2.0 via a unified end-to-end model
Haoran Wei, Chenglong Liu, Jinyue Chen, Jia Wang, Lingyu Kong, Yanming Xu, Zheng Ge, Liang Zhao, Jianjian Sun, Yuang Peng, et al · 2024
Closest in time.
End-to-end video text spotting with transformer
Weijia Wu, Yuanqiang Cai, Chunhua Shen, Debing Zhang, Ying Fu, Hong Zhou, and Ping Luo. 2024 · 2024
Closest in time.
Renqiu Xia, Song Mao, Xiangchao Yan, Hongbin Zhou, Bo Zhang, Haoyang Peng, Jiahao Pi, Daocheng Fu, Wenjie Wu, Hancheng Ye, et al · 2024
Closest in time.
Chartx & chartvlm: A versatile benchmark and foundation model for complicated chart reasoning
Renqiu Xia, Bo Zhang, Hancheng Ye, Xiangchao Yan, Qi Liu, Hongbin Zhou, Zijun Chen, Min Dou, Botian Shi, Junchi Yan, et al · 2024
Closest in time.
WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
Xudong Xie, Liang Yin, Hao Yan, Yang Liu, Jing Ding, Minghui Liao, Yuliang Liu, Wei Chen, and Xiang Bai. 2024 · 2024
Closest in time.
Empowering 1000 tokens/second on-device llm prefilling with mllm-npu
Daliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu, Gang Huang, Mengwei Xu, and Xuanzhe Liu. 2024 · 2024
Closest in time.
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
Shi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui, Junhao Ran, Yukun Yan, Zhenghao Liu, Shuo Wang, Xu Han, Zhiyuan Liu, et al · 2024
Closest in time.
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
Junyuan Zhang, Qintong Zhang, Bin Wang, Linke Ouyang, Zichen Wen, Ying Li, Ka-Ho Chow, Conghui He, and Wentao Zhang. 2024b · 2024
Closest in time.
SEMv2: Table separation line detection based on instance segmentation
Zhenrong Zhang, Pengfei Hu, Jiefeng Ma, Jun Du, Jianshu Zhang, Baocai Yin, Bing Yin, and Cong Liu. 2024a · 2024
Closest in time.
Retrieval-augmented generation for ai-generated content: A survey
Penghao Zhao, Hailin Zhang, Qinhan Yu, Zhengren Wang, Yunteng Geng, Fangcheng Fu, Ling Yang, Wentao Zhang, and Bin Cui. 2024b · 2024
Closest in time.
Zhiyuan Zhao, Hengrui Kang, Bin Wang, and Conghui He. 2024a · 2024
Closest in time.
ICAL: Implicit Character-Aided Learning for Enhanced Handwritten Mathematical Expression Recognition. In International Conference on Document Analysis and Recognition . Springer, 21–37
Jianhua Zhu, Liangcai Gao, and Wenqi Zhao. 2024 · 2024
Closest in time.
DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems
Anni Zou, Wenhao Yu, Hongming Zhang, Kaixin Ma, Deng Cai, Zhuosheng Zhang, Hai Zhao, and Dong Yu. 2024 · 2024
Closest in time.
Vary: Scaling up the Vision Vocabulary for Large Vision-Language Model. In European Conference on Computer Vision . Springer, 408–424
Haoran Wei, Lingyu Kong, Jinyue Chen, Liang Zhao, Zheng Ge, Jinrong Yang, Jianjian Sun, Chunrui Han, and Xiangyu Zhang. 2025 · 2025
Closest in time.
Tensormask: A foundation for dense object segmentation. In Proceedings of the IEEE/CVF international conference on computer vision . 2061–2069
Xinlei Chen, Ross Girshick, Kaiming He, and Piotr Dollár. 2019 · 2069
Closest in time.