Fetching the paper…
Reading the bibliography…
This paper presents a comprehensive evaluation of the Optical Character Recognition (OCR) capabilities of the recently released GPT-4V(ision), a Large Multimodal Model (LMM).
The IAM-database: an english sentence database for offline handwriting recognition
U-V Marti and Horst Bunke · 2002
Earlier work this paper cites.
Handwritten Chinese text recognition by integrating multiple contexts
Qiu-Feng Wang, Fei Yin, and Cheng-Lin Liu · 2011
Earlier work this paper cites.
ICDAR 2013 Chinese handwriting recognition competition
Fei Yin, Qiu-Feng Wang, Xu-Yao Zhang, and Cheng-Lin Liu · 2013
Earlier work this paper cites.
A robust arbitrary text detection system for natural scene images
Anhar Risnumawan, Palaiahnakote Shivakumara, Chee Seng Chan, and Chew Lim Tan · 2014
Earlier work this paper cites.
ICFHR 2014 competition on recognition of on-line handwritten mathematical expressions (CROHME 2014)
Harold Mouchere, Christian Viard-Gaudin, Richard Zanibbi, and Utpal Garain · 2014
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao · 2016
Earlier work this paper cites.
Joint line segmentation and transcription for end-to-end handwritten paragraph recognition
Théodore Bluche · 2016
Earlier work this paper cites.
Learning spatial-semantic context with fully convolutional recurrent network for online handwritten Chinese text recognition
Zecheng Xie, Zenghui Sun, Lianwen Jin, Hao Ni, and Terry Lyons · 2017
Earlier work this paper cites.
Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition
Jianshu Zhang, Jun Du, Shiliang Zhang, Dan Liu, Yulong Hu, Jinshui Hu, Si Wei, and Lirong Dai · 2017
Earlier work this paper cites.
Total-Text: A comprehensive dataset for scene text detection and recognition
Chee-Kheng Chng and Chee Seng Chan · 2017
Earlier work this paper cites.
Aster: An attentional scene text recognizer with flexible rectification
Baoguang Shi, Mingkun Yang, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai · 2018
Earlier work this paper cites.
Mask TextSpotter: An end-to-end trainable neural network for spotting text with arbitrary shapes
Pengyuan Lyu, Minghui Liao, Cong Yao, Wenhao Wu, and Xiang Bai · 2018
Earlier work this paper cites.
Multi-scale attention with dense encoder for handwritten mathematical expression recognition
Jianshu Zhang, Jun Du, and Lirong Dai · 2018
Earlier work this paper cites.
Offline continuous handwriting recognition using sequence to sequence neural networks
Jorge Sueiras, Victoria Ruiz, Angel Sanchez, and Jose F Velez · 2018
Earlier work this paper cites.
Moran: A multi-object rectified attention network for scene text recognition
Canjie Luo, Lianwen Jin, and Zenghui Sun · 2019
Earlier work this paper cites.
Mask TextSpotter: An end-to-end trainable neural network for spotting text with arbitrary shapes
Minghui Liao, Pengyuan Lyu, Minghang He, Cong Yao, Wenhao Wu, and Xiang Bai · 2019
Earlier work this paper cites.
A fast and accurate fully convolutional network for end-to-end handwritten Chinese text segmentation and recognition
Dezhi Peng, Lianwen Jin, Yaqiang Wu, Zhepeng Wang, and Mingxiang Cai · 2019
Earlier work this paper cites.
Curved scene text detection via transverse and longitudinal sequence connection
Yuliang Liu, Lianwen Jin, Shuaitao Zhang, Canjie Luo, and Sheng Zhang · 2019
Earlier work this paper cites.
ICDAR 2019 robust reading challenge on reading Chinese text on signboard
Rui Zhang, Yongsheng Zhou, Qianyi Jiang, Qi Song, Nan Li, Kai Zhou, Lei Wang, Dong Wang, Minghui Liao, Mingkun Yang, et al · 2019
Earlier work this paper cites.
ICDAR2019 robust reading challenge on multi-lingual scene text detection and recognition—RRC-MLT-2019
Nibal Nayef, Yash Patel, Michal Busta, Pinaki Nath Chowdhury, Dimosthenis Karatzas, Wafa Khlif, Jiri Matas, Umapada Pal, Jean-Christophe Burie, Cheng-lin Liu, et al · 2019
Earlier work this paper cites.
Complicated table structure recognition
Zewen Chi, Heyan Huang, Heng-Da Xu, Houjin Yu, Wanxuan Yin, and Xian-Ling Mao · 2019
Earlier work this paper cites.
FUNSD: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran · 2019
Earlier work this paper cites.
Decoupled attention network for text recognition
Tianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo, Xiaoxue Chen, Yaqiang Wu, Qianying Wang, and Mingxiang Cai · 2020
Earlier work this paper cites.
Mask TextSpotter v3: Segmentation proposal network for robust scene text spotting
Minghui Liao, Guan Pang, Jing Huang, Tal Hassner, and Xiang Bai · 2020
Earlier work this paper cites.
OrigamiNet: weakly-supervised, segmentation-free, one-step, full page text recognition by learning to unfold
Mohamed Yousef and Tom E Bishop · 2020
Earlier work this paper cites.
Image-based table recognition: data, model, and evaluation
Xu Zhong, Elaheh ShafieiBavani, and Antonio Jimeno Yepes · 2020
Earlier work this paper cites.
LayoutLM: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou · 2020
Earlier work this paper cites.
Writer-aware CNN for parsimonious HMM-based offline handwritten Chinese text recognition
Zi-Rui Wang, Jun Du, and Jia-Ming Wang · 2020
Earlier work this paper cites.
High performance offline handwritten Chinese text recognition with a new data preprocessing and augmentation pipeline
Canyu Xie, Songxuan Lai, Qianying Liao, and Lianwen Jin · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition
Shancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao, and Yongdong Zhang · 2021
Cited alongside, same era.
ABCNet v2: Adaptive bezier-curve network for real-time end-to-end text spotting
Yuliang Liu, Chunhua Shen, Lianwen Jin, Tong He, Peng Chen, Chongyu Liu, and Hao Chen · 2021
Cited alongside, same era.
Jiaquan Ye, Xianbiao Qi, Yelin He, Yihao Chen, Dengyi Gu, Peng Gao, and Rong Xiao · 2021
Cited alongside, same era.
Show, read and reason: Table structure recognition with flexible context aggregator
Hao Liu, Xin Li, Bing Liu, Deqiang Jiang, Yinsong Liu, Bo Ren, and Rongrong Ji · 2021
Cited alongside, same era.
PICK: Processing key information extraction from documents using improved graph learning-convolutional networks
Wenwen Yu, Ning Lu, Xianbiao Qi, Ping Gong, and Rong Xiao · 2021
Improvising the CNN feature maps through integration of channel attention for handwritten text recognition
BN Shashank, S Nagesh Bhattu, and K Sri Phani Krishna · 2022
Later among the works it cites.
Stanford Alpaca: An instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
ERNIE Bot
Baidu · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards robust visual information extraction in real world: New dataset and novel solution
Jiapeng Wang, Chongyu Liu, Lianwen Jin, Guozhi Tang, Jiaxin Zhang, Shuaitao Zhang, Qianying Wang, Yaqiang Wu, and Mingxiang Cai · 2021
Cited alongside, same era.
Tag, copy or predict: A unified weakly-supervised learning framework for visual information extraction using sequences
Jiapeng Wang, Tianwei Wang, Guozhi Tang, Lianwen Jin, Weihong Ma, Kai Ding, and Yichao Huang · 2021
Cited alongside, same era.
SelfDoc: Self-supervised document representation learning
Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu · 2021
Cited alongside, same era.
DocFormer: End-to-end transformer for document understanding
Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R. Manmatha · 2021
Cited alongside, same era.
Parsing table structures in the wild
Rujiao Long, Wen Wang, Nan Xue, Feiyu Gao, Zhibo Yang, Yongpan Wang, and Gui-Song Xia · 2021
Cited alongside, same era.
GLM-130B: An open bilingual pre-trained model
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al · 2022
Cited alongside, same era.
SGBANet: semantic gan and balanced attention network for arbitrarily oriented scene text recognition
Dajian Zhong, Shujing Lyu, Palaiahnakote Shivakumara, Bing Yin, Jiajia Wu, Umapada Pal, and Yue Lu · 2022
Cited alongside, same era.
Baichuan · 2023
Closest in time.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Openflamingo, March 2023
Anas Awadalla, Irena Gao, Joshua Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Closest in time.
MiniGPT-4: Enhancing vision-language understanding with advanced large language models, 2023
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Closest in time.
mPLUG-Owl: Modularization empowers large language models with multimodality, 2023
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang · 2023
Closest in time.
The dawn of LMMs: Preliminary explorations with GPT-4V (ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang · 2023
Closest in time.
DeepSolo: Let transformer decoder with explicit points solo for text spotting
Maoyuan Ye, Jing Zhang, Shanshan Zhao, Juhua Liu, Tongliang Liu, Bo Du, and Dacheng Tao · 2023
Closest in time.
ESTextSpotter: Towards better scene text spotting with explicit synergy in transformer
Mingxin Huang, Jiaxin Zhang, Dezhi Peng, Hao Lu, Can Huang, Yuliang Liu, Xiang Bai, and Lianwen Jin · 2023
Closest in time.
SPTS v2: Single-point scene text spotting
Yuliang Liu, Jiaxin Zhang, Dezhi Peng, Mingxin Huang, Xinyu Wang, Jingqun Tang, Can Huang, Dahua Lin, Chunhua Shen, Xiang Bai, and Lianwen Jin · 2023
Closest in time.
Arbitrary shape text detection via boundary transformer
Shi-Xue Zhang, Chun Yang, Xiaobin Zhu, and Xu-Cheng Yin · 2023
Closest in time.
SegCTC: Offline handwritten Chinese text recognition via better fusion between explicit and implicit segmentation
Jiarong Huang, Dezhi Peng, Hongliang Li, Hao Ni, and Lianwen Jin · 2023
Closest in time.
DAN: a segmentation-free document attention network for handwritten document recognition
Denis Coquenet, Clément Chatelain, and Thierry Paquet · 2023
Closest in time.
Improving handwritten mathematical expression recognition via similar symbol distinguishing
Zhe Li, Xinyu Wang, Yuliang Liu, Lianwen Jin, Yichao Huang, and Kai Ding · 2023
Closest in time.
Improving table structure recognition with visual-alignment sequential coordinate modeling
Yongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li, Zecheng Xie, Shenggao Zhu, Liangcai Gao, and Wei Peng · 2023
Closest in time.
Divide rows and conquer cells: Towards structure recognition for large tables
Huawen Shen, Xiang Gao, Jin Wei, Liang Qiao, Yu Zhou, Qiang Li, and Zhanzhan Cheng · 2023
Closest in time.
StrucTexTv2: Masked visual-textual prediction for document image pre-training
Yuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang, Zengyuan Guo, Xiameng Qin, Kun Yao, Junyu Han, Errui Ding, and Jingdong Wang · 2023
Closest in time.
On the hidden mystery of OCR in large multimodal models
Yuliang Liu, Zhang Li, Hongliang Li, Wenwen Yu, Mingxin Huang, Dezhi Peng, Mingyu Liu, Mingrui Chen, Chunyuan Li, Lianwen Jin, et al · 2023
Closest in time.
Revisiting scene text recognition: A data perspective
Qing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu, and Lianwen Jin · 2023
Closest in time.
Looking and listening: Audio guided text recognition
Wenwen Yu, Mingyu Liu, Biao Yang, Enming Zhang, Deqiang Jiang, Xing Sun, Yuliang Liu, and Xiang Bai · 2023
Closest in time.
A comprehensive handwritten paragraph text recognition system: Lexiconnet
Lalita Kumari, Sukhdeep Singh, Vaibhav Varish Singh Rathore, and Anuj Sharma · 2023
Closest in time.
TrOCR: Transformer-based optical character recognition with pre-trained models
Minghao Li, Tengchao Lv, Jingye Chen, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, and Furu Wei · 2023
Closest in time.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning, 2023
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Closest in time.
InstructionGPT-4: A 200-instruction paradigm for fine-tuning MiniGPT-4
Lai Wei, Zihao Jiang, Weiran Huang, and Lichao Sun · 2023
Closest in time.