Fetching the paper…
Reading the bibliography…
Recently, visually-situated text parsing (VsTP) has experienced notable advancements, driven by the increasing demand for automated document understanding and the emergence of Generative Large Language Models (LLMs) capable of processing document-based questions.
Icdar 2013 robust reading competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras · 2013
Earlier work this paper cites.
Icdar 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al · 2015
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao · 2016
Earlier work this paper cites.
Image-to-markup generation with coarse-to-fine attention
Yuntian Deng, Anssi Kanervisto, Jeffrey Ling, and Alexander M Rush · 2017
Earlier work this paper cites.
Icdar2017 robust reading challenge on coco-text
Raul Gomez, Baoguang Shi, Lluis Gomez, Lukas Numann, Andreas Veit, Jiri Matas, Serge Belongie, and Dimosthenis Karatzas · 2017
Earlier work this paper cites.
Towards end-to-end text spotting with convolutional recurrent neural networks
Hui Li, Peng Wang, and Chunhua Shen · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Earlier work this paper cites.
Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification-rrc-mlt
Nibal Nayef, Fei Yin, Imen Bizid, Hyunsoo Choi, Yuan Feng, Dimosthenis Karatzas, Zhenbo Luo, Umapada Pal, Christophe Rigaud, Joseph Chazalon, et al · 2017
Earlier work this paper cites.
East: an efficient and accurate scene text detector
Xinyu Zhou, Cong Yao, He Wen, Yuzhi Wang, Shuchang Zhou, Weiran He, and Jiajun Liang · 2017
Earlier work this paper cites.
An end-to-end textspotter with explicit alignment and attention
Tong He, Zhi Tian, Weilin Huang, Chunhua Shen, Yu Qiao, and Changming Sun · 2018
Earlier work this paper cites.
Fots: Fast oriented text spotting with a unified network
Xuebo Liu, Ding Liang, Shi Yan, Dagui Chen, Yu Qiao, and Junjie Yan · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes
Pengyuan Lyu, Minghui Liao, Cong Yao, Wenhao Wu, and Xiang Bai · 2018
Earlier work this paper cites.
Corpus conversion service: A machine learning platform to ingest documents at scale
Peter WJ Staar, Michele Dolfi, Christoph Auer, and Costas Bekas · 2018
Earlier work this paper cites.
Textdragon: An end-to-end framework for arbitrary shaped text spotting
Wei Feng, Wenhao He, Fei Yin, Xu-Yao Zhang, and Cheng-Lin Liu · 2019
Earlier work this paper cites.
Icdar2019 competition on scanned receipt ocr and information extraction
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar · 2019
Earlier work this paper cites.
Curved scene text detection via transverse and longitudinal sequence connection
Yuliang Liu, Lianwen Jin, Shuaitao Zhang, Canjie Luo, and Sheng Zhang · 2019
Earlier work this paper cites.
Cord: A consolidated receipt dataset for post-ocr parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee · 2019
Earlier work this paper cites.
Towards unconstrained end-to-end text spotting
Siyang Qin, Alessandro Bissacco, Michalis Raptis, Yasuhisa Fujii, and Ying Xiao · 2019
Earlier work this paper cites.
Textnet: Irregular text reading from images with an end-to-end trainable network
Yipeng Sun, Chengquan Zhang, Zuming Huang, Jiaming Liu, Junyu Han, and Errui Ding · 2019
Earlier work this paper cites.
Convolutional character networks
Linjie Xing, Zhi Tian, Weilin Huang, and Matthew R Scott · 2019
Earlier work this paper cites.
Publaynet: largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes · 2019
Earlier work this paper cites.
Character region attention for text spotting
Youngmin Baek, Seung Shin, Jeonghun Baek, Sungrae Park, Junyeop Lee, Daehyun Nam, and Hwalsuk Lee · 2020
Earlier work this paper cites.
Total-text: toward orientation robustness in scene text detection
Chee-Kheng Ch’ng, Chee Seng Chan, and Cheng-Lin Liu · 2020
Earlier work this paper cites.
Mask textspotter v3: Segmentation proposal network for robust scene text spotting
Minghui Liao, Guan Pang, Jing Huang, Tal Hassner, and Xiang Bai · 2020
Earlier work this paper cites.
Text perceptron: Towards end-to-end arbitrary-shaped text spotting
Liang Qiao, Sanli Tang, Zhanzhan Cheng, Yunlu Xu, Yi Niu, Shiliang Pu, and Fei Wu · 2020
Earlier work this paper cites.
All you need is boundary: Toward arbitrary-shaped text spotting
Hao Wang, Pu Lu, Hui Zhang, Mingkun Yang, Xiang Bai, Yongchao Xu, Mengchao He, Yongpan Wang, and Wenyu Liu · 2020
Earlier work this paper cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Cited alongside, same era.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou · 2020
Cited alongside, same era.
Trie: end-to-end text reading and information extraction for document understanding
Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu · 2020
Cited alongside, same era.
Image-based table recognition: data, model, and evaluation
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes · 2020
Cited alongside, same era.
Docformer: End-to-end transformer for document understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R Manmatha · 2021
Cited alongside, same era.
Pix2seq: A language modeling framework for object detection
Real-time scene text detection with differentiable binarization and adaptive scale fusion
Minghui Liao, Zhisheng Zou, Zhaoyi Wan, Cong Yao, and Xiang Bai · 2022
Later among the works it cites.
Tsrformer: Table structure recognition with transformers
Weihong Lin, Zheng Sun, Chixiang Ma, Mingze Li, Jiawei Wang, Lei Sun, and Qiang Huo · 2022
Later among the works it cites.
Towards end-to-end unified scene text detection and layout analysis
Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, and Michalis Raptis · 2022
Later among the works it cites.
Tableformer: Table structure understanding with transformers
Ahmed Nassar, Nikolaos Livathinos, Maksym Lysak, and Peter Staar · 2022
Later among the works it cites.
Ernie-layout: Layout knowledge enhanced pre-training for visually-rich document understanding
Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Yuhui Cao, Weichong Yin, Yongfeng Chen, Yin Zhang, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ting Chen, Saurabh Saxena, Lala Li, David J Fleet, and Geoffrey Hinton · 2021
Cited alongside, same era.
Unidoc: Unified pretraining framework for document understanding
Jiuxiang Gu, Jason Kuen, Vlad I Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, and Tong Sun · 2021
Cited alongside, same era.
Spatial dependency parsing for semi-structured document information extraction
Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo · 2021
Cited alongside, same era.
Open images v5 text annotation and yet another mask text spotter
Ilya Krylov, Sergei Nosov, and Vladislav Sovrasov · 2021
Cited alongside, same era.
Parsing table structures in the wild
Rujiao Long, Wen Wang, Nan Xue, Feiyu Gao, Zhibo Yang, Yongpan Wang, and Gui-Song Xia · 2021
Cited alongside, same era.
Mango: A mask attention guided one-stage scene text spotter
Liang Qiao, Ying Chen, Zhanzhan Cheng, Yunlu Xu, Yi Niu, Shiliang Pu, and Fei Wu · 2021
Cited alongside, same era.
Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang, Wojciech Galuba, and Tal Hassner · 2021
Cited alongside, same era.
Glass: Global to local attention for scene-text spotting
Roi Ronen, Shahar Tsiper, Oron Anschel, Inbal Lavi, Amir Markovitz, and R Manmatha · 2022
Later among the works it cites.
Vision-language pre-training for boosting scene text detectors
Sibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang, Wenqing Cheng, Xiang Bai, and Cong Yao · 2022
Later among the works it cites.
Tpsnet: Reverse thinking of thin plate splines for arbitrary shape scene text representation
Wei Wang, Yu Zhou, Jiahao Lv, Dayan Wu, Guoqing Zhao, Ning Jiang, and Weipinng Wang · 2022
Later among the works it cites.
Structextv2: Masked visual-textual prediction for document image pre-training
Yuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang, Zengyuan Guo, Xiameng Qin, Kun Yao, Junyu Han, Errui Ding, and Jingdong Wang · 2022
Later among the works it cites.
Text spotting transformers
Xiang Zhang, Yongwen Su, Subarna Tripathi, and Zhuowen Tu · 2022
Later among the works it cites.
Multi-granularity prediction with learnable fusion for scene text recognition
Cheng Da, Peng Wang, and Cong Yao · 2023
Later among the works it cites.
Docparser: End-to-end ocr-free information extraction from visually rich documents
Mohamed Dhouib, Ghassen Bettaieb, and Aymen Shabou · 2023
Later among the works it cites.
Hao Feng, Zijian Wang, Jingqun Tang, Jinghui Lu, Wengang Zhou, Houqiang Li, and Can Huang · 2023
Later among the works it cites.
Improving table structure recognition with visual-alignment sequential coordinate modeling
Yongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li, Zecheng Xie, Shenggao Zhu, Liangcai Gao, and Wei Peng · 2023
Later among the works it cites.
Towards unified scene text spotting based on sequence generation
Taeho Kil, Seonghyeon Kim, Sukmin Seo, Yoonsik Kim, and Daehee Kim · 2023
Later among the works it cites.
Visual information extraction in the wild: practical dataset and end-to-end solution
Jianfeng Kuang, Wei Hua, Dingkang Liang, Mingkun Yang, Deqiang Jiang, Bo Ren, and Xiang Bai · 2023
Later among the works it cites.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Later among the works it cites.
Spts v2: single-point scene text spotting
Yuliang Liu, Jiaxin Zhang, Dezhi Peng, Mingxin Huang, Xinyu Wang, Jingqun Tang, Can Huang, Dahua Lin, Chunhua Shen, Xiang Bai, et al · 2023
Later among the works it cites.
Geolayoutlm: Geometric pre-training for visual information extraction
Chuwei Luo, Changxu Cheng, Qi Zheng, and Cong Yao · 2023
Later among the works it cites.
An end-to-end local attention based model for table recognition
Nam Tuan Ly and Atsuhiro Takasu · 2023
Later among the works it cites.
Gridformer: Towards accurate table structure recognition via grid prediction
Pengyuan Lyu, Weihong Ma, Hongyi Wang, Yuechen Yu, Chengquan Zhang, Kun Yao, Yang Xue, and Jingdong Wang · 2023
Later among the works it cites.
GPT-4V(ision) System Card
OpenAI · 2023
Later among the works it cites.
Unifying vision, text, and layout for universal document processing
Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, and Mohit Bansal · 2023
Later among the works it cites.
Ppn: Parallel pointer-based network for key information extraction with complex layouts
Kaiwen Wei, Jie Yao, Jingyuan Zhang, Yangyang Kang, Fubang Zhao, Yating Zhang, Changlong Sun, Xin Jin, and Xin Zhang · 2023
Later among the works it cites.
Modeling entities as semantic points for visual information extraction in the wild
Zhibo Yang, Rujiao Long, Pengfei Wang, Sibo Song, Humen Zhong, Wenqing Cheng, Xiang Bai, and Cong Yao · 2023
Later among the works it cites.
Deepsolo: Let transformer decoder with explicit points solo for text spotting
Maoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu, Tongliang Liu, Bo Du, and Dacheng Tao · 2023
Later among the works it cites.
Reading order matters: Information extraction from visually-rich documents by token path prediction
Chong Zhang, Ya Guo, Yi Tu, Huan Chen, Jinyang Tang, Huijia Zhu, Qi Zhang, and Tao Gui · 2023
Later among the works it cites.