Fetching the paper…
Reading the bibliography…
The automatic reading of text-intensive images represents a significant advancement toward achieving Artificial General Intelligence (AGI).
TableBank: A Benchmark Dataset for Table Detection and Recognition
Li, M.; Cui, L.; Huang, S.; Wei, F.; Zhou, M.; and Li, Z. 2020 · 1903
Earlier work this paper cites.
Simple fast algorithms for the editing distance between trees and related problems
Zhang, K.; and Shasha, D. 1989 · 1989
Earlier work this paper cites.
On the resemblance and containment of documents
Broder, A. Z. 1997 · 1997
Earlier work this paper cites.
Building a test collection for complex document information processing
Lewis, D.; Agam, G.; Argamon, S.; Frieder, O.; Grossman, D.; and Heard, J. 2006 · 2006
Earlier work this paper cites.
DocVQA: A Dataset for VQA on Document Images
Mathew, M.; Karatzas, D.; and Jawahar, C. V. 2020 · 2007
Earlier work this paper cites.
An Overview of the Tesseract OCR Engine
Smith, R. 2007 · 2007
Earlier work this paper cites.
Synthetic data and artificial neural networks for natural scene text recognition
Jaderberg, e. a. 2014 · 2014
Earlier work this paper cites.
ICFHR 2014 competition on recognition of on-line handwritten mathematical expressions (CROHME 2014)
Mouchere, H.; Viard-Gaudin, C.; Zanibbi, R.; and Garain, U. 2014 · 2014
Earlier work this paper cites.
Compositional Semantic Parsing on Semi-Structured Tables
Pasupat, P.; and Liang, P. 2015 · 2015
Earlier work this paper cites.
Synthetic data for text localisation in natural images
Gupta, e. a. 2016 · 2016
Earlier work this paper cites.
Image-to-markup generation with coarse-to-fine attention
Deng, Y.; Kanervisto, A.; Ling, J.; and Rush, A. M. 2017 · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Earlier work this paper cites.
Icdar2019 competition on scanned receipt ocr and information extraction
Huang, Z.; Chen, K.; He, J.; Bai, X.; Karatzas, D.; Lu, S.; and Jawahar, C. 2019 · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Jaume, e. a. 2019 · 2019
Earlier work this paper cites.
CORD: A Consolidated Receipt Dataset for Post-OCR Parsing
Park, S.; Shin, S.; Lee, B.; Lee, J.; Surh, J.; Seo, M.; and Lee, H. 2019 · 2019
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N. M.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2019 · 2019
Earlier work this paper cites.
Towards VQA Models That Can Read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Earlier work this paper cites.
TabFact : A Large-scale Dataset for Table-based Fact Verification
Chen, W.; Wang, H.; Chen, J.; Zhang, Y.; Wang, H.; Li, S.; Zhou, X.; and Wang, W. Y. 2020 · 2020
Earlier work this paper cites.
TextCaps: A Dataset for Image Captioning with Reading Comprehension
Sidorov, O.; Hu, R.; Rohrbach, M.; and Singh, A. 2020 · 2020
Earlier work this paper cites.
DeepForm: Understand structured documents at scale
Svetlichnaya, S. 2020 · 2020
Earlier work this paper cites.
Layoutlm: Pre-training of text and layout for document image understanding
Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; and Zhou, M. 2020 · 2020
Earlier work this paper cites.
Docformer: End-to-end transformer for document understanding
Appalaraju, S.; Jasani, B.; Kota, B. U.; Xie, Y.; and Manmatha, R. 2021 · 2021
Earlier work this paper cites.
Document AI: Benchmarks, Models and Applications
Cui, L.; Xu, Y.; Lv, T.; and Wei, F. 2021 · 2021
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Earlier work this paper cites.
Seeing out of the box: End-to-end pre-training for vision-language representation learning
Huang, Z.; Zeng, Z.; Huang, Y.; Liu, B.; Fu, D.; and Fu, J. 2021 · 2021
Earlier work this paper cites.
Donut: Document understanding transformer without ocr
Kim, G.; Hong, T.; Yim, M.; Park, J.; Yim, J.; Hwang, W.; Yun, S.; Han, D.; and Park, S. 2021 · 2021
Cited alongside, same era.
Deduplicating training data makes language models better
Lee, K.; Ippolito, D.; Nystrom, A.; Zhang, C.; Eck, D.; Callison-Burch, C.; and Carlini, N. 2021 · 2021
Cited alongside, same era.
Kleister: Key Information Extraction Datasets Involving Long Documents with Complex Layouts
Stanislawek, T.; Gralinski, F.; Wróblewska, A.; Lipinski, D.; Kaliska, A.; Rosalska, P.; Topolski, B.; and Biecek, P. 2021 · 2021
Cited alongside, same era.
VisualMRC: Machine Reading Comprehension on Document Images
Tanaka, R.; Nishida, K.; and Yoshida, S. 2021 · 2021
Cited alongside, same era.
LayoutReader: Pre-training of Text and Layout for Reading Order Detection
Wang, Z.; Xu, Y.; Cui, L.; Shang, J.; and Wei, F. 2021 · 2021
Videopoet: A large language model for zero-shot video generation
Kondratyuk, D.; Yu, L.; Gu, X.; Lezama, J.; Huang, J.; Hornung, R.; Adam, H.; Akbari, H.; Alon, Y.; Birodkar, V.; et al. 2023 · 2023
Closest in time.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
Lee, K.; Joshi, M.; Turc, I. R.; Hu, H.; Liu, F.; Eisenschlos, J. M.; Khandelwal, U.; Shaw, P.; Chang, M.-W.; and Toutanova, K. 2023 · 2023
Closest in time.
Taskmatrix. ai: Completing tasks by connecting foundation models with millions of apis
Liang, Y.; Wu, C.; Song, T.; Wu, W.; Xia, Y.; Liu, Y.; Ou, Y.; Lu, S.; Ji, L.; Mao, S.; et al. 2023 · 2023
Closest in time.
Cheap and quick: Efficient vision-language instruction tuning for large language models
Luo, G.; Zhou, Y.; Ren, T.; Chen, S.; Sun, X.; and Ji, R. 2023 · 2023
Closest in time.
Kosmos-2: Grounding Multimodal Large Language Models to the World
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Probing inter-modality: Visual parsing with self-attention for vision-and-language pre-training
Xue, H.; Huang, Y.; Liu, B.; Peng, H.; Fu, J.; Li, H.; and Luo, J. 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Cited alongside, same era.
https://openai.com/blog/chatgpt
ChatGPT. 2022 · 2022
Cited alongside, same era.
Xdoc: Unified pre-training for cross-format document understanding
Chen, J.; Lv, T.; Cui, L.; Zhang, C.; and Wei, F. 2022 · 2022
Cited alongside, same era.
End-to-end document recognition and understanding with dessurt
Davis, B.; Morse, B.; Price, B.; Tensmeyer, C.; Wigington, C.; and Morariu, V. 2022 · 2022
Cited alongside, same era.
Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding
Gu, Z.; Meng, C.; Wang, K.; Lan, J.; Wang, W.; Gu, M.; and Zhang, L. 2022 · 2022
Cited alongside, same era.
Language Models are General-Purpose Interfaces
Hao, Y.; Song, H.; Dong, L.; Huang, S.; Chi, Z.; Wang, W.; Ma, S.; and Wei, F. 2022 · 2022
Cited alongside, same era.
Peng, Z.; Wang, W.; Dong, L.; Hao, Y.; Huang, S.; Ma, S.; and Wei, F. 2023 · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; and Zhuang, Y. 2023 · 2023
Closest in time.
Pandagpt: One model to instruction-follow them all
Su, Y.; Lan, T.; Li, H.; Xu, J.; Wang, Y.; and Cai, D. 2023 · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
Surís, D.; Menon, S.; and Vondrick, C. 2023 · 2023
Closest in time.
Codi-2: In-context, interleaved, and interactive any-to-any generation
Tang, Z.; Yang, Z.; Khademi, M.; Liu, Y.; Zhu, C.; and Bansal, M. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
Wang, W.; Chen, Z.; Chen, X.; Wu, J.; Zhu, X.; Zeng, G.; Luo, P.; Lu, T.; Zhou, J.; Qiao, Y.; et al. 2023 · 2023
Closest in time.
Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Wei, H.; Kong, L.; Chen, J.; Zhao, L.; Ge, Z.; Yang, J.; Sun, J.; Han, C.; and Zhang, X. 2023 · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Wu, C.; Yin, S.; Qi, W.; Wang, X.; Tang, Z.; and Duan, N. 2023 · 2023
Closest in time.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Yang, Z.; Li, L.; Wang, J.; Lin, K.; Azarnasab, E.; Ahmed, F.; Liu, Z.; Liu, C.; Zeng, M.; and Wang, L. 2023 · 2023
Closest in time.
UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Ye, J.; Hu, A.; Xu, H.; Ye, Q.; Yan, M.; Xu, G.; Li, C.; Tian, J.; Qian, Q.; Zhang, J.; Jin, Q.; He, L.; Lin, X.; and Huang, F. 2023b · 2023
Closest in time.
StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training
Yu, Y.; Li, Y.; Zhang, C.; Zhang, X.; Guo, Z.; Qin, X.; Yao, K.; Han, J.; Ding, E.; and Wang, J. 2023 · 2023
Closest in time.
Chatspot: Bootstrapping multimodal llms via precise referring instruction tuning
Zhao, L.; Yu, E.; Ge, Z.; Yang, J.; Wei, H.; Zhou, H.; Sun, J.; Peng, Y.; Dong, R.; Han, C.; et al. 2023 · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Closest in time.
Textdiffuser: Diffusion models as text painters
Chen, J.; Huang, Y.; Lv, T.; Cui, L.; Chen, Q.; and Wei, F. 2024 · 2024
Closest in time.
Cogagent: A visual language model for gui agents
Hong, W.; Wang, W.; Lv, Q.; Xu, J.; Yu, W.; Ji, J.; Wang, Y.; Wang, Z.; Dong, Y.; Ding, M.; et al. 2024 · 2024
Closest in time.
mplug-docowl 1.5: Unified structure learning for ocr-free document understanding
Hu, A.; Xu, H.; Ye, J.; Yan, M.; Zhang, L.; Zhang, B.; Li, C.; Zhang, J.; Jin, Q.; Huang, F.; et al. 2024 · 2024
Closest in time.
Small Language Model Meets with Reinforced Vision Vocabulary
Wei, H.; Kong, L.; Chen, J.; Zhao, L.; Ge, Z.; Yu, E.; Sun, J.; Han, C.; and Zhang, X. 2024 · 2024
Closest in time.
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Yao, Y.; Yu, T.; Zhang, A.; Wang, C.; Cui, J.; Zhu, H.; Cai, T.; Li, H.; Zhao, W.; He, Z.; Chen, Q.; Zhou, H.; Zou, Z.; Zhang, H.; Hu, S.; Zheng, Z.; Zhou, J.; Cai, J.; Han, X.; Zeng, G.; Li, D.; Liu, Z.; and Sun, M. 2024 · 2024
Closest in time.
TRINS: Towards Multimodal Language Models that Can Read
Zhang, R.; Zhang, Y.; Chen, J.; Zhou, Y.; Gu, J.; Chen, C.; and Sun, T. 2024 · 2024
Closest in time.