Fetching the paper…
Reading the bibliography…
This work presents DocPedia, a novel large multimodal model (LMM) for versatile OCR-free document understanding, capable of parsing images up to 2,560$\times$2,560 resolution.
Language models are few-shot learners
Brown T, Mann B, Ryder N, et al · 1901
Earlier work this paper cites.
StrucTexT: Structured text understanding with multi-modal transformers
Li Y, Qian Y, Yu Y, et al · 1920
Earlier work this paper cites.
Discrete cosine transform
Ahmed N, Natarajan T, Rao K R · 1974
Earlier work this paper cites.
Document image understanding
Srihari S N, Lam S W, Govindaraju V, et al · 1986
Earlier work this paper cites.
The jpeg still picture compression standard
Wallace G K · 1991
Earlier work this paper cites.
End-to-end scene text recognition
Wang K, Babenko B, Belongie S · 2011
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Shi B, Bai X, Yao C · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani A, Shazeer N, Parmar N, et al · 2017
Earlier work this paper cites.
FigureQA: An annotated figure dataset for visual reasoning
Kahou S E, Michalski V, Atkinson A, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov I, Hutter F · 2017
Earlier work this paper cites.
FOTS: Fast oriented text spotting with a unified network
Liu X, Liang D, Yan S, et al · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin J, Chang M W, Lee K, et al · 2018
Earlier work this paper cites.
Post-OCR parsing: building simple and robust parser via bio tagging
Hwang W, Kim S, Seo M, et al · 2019
Earlier work this paper cites.
A comprehensive survey of deep learning for image captioning
Hossain M Z, Sohel F, Shiratuddin M F, et al · 2019
Earlier work this paper cites.
OCR-VQA: Visual question answering by reading text in images
Mishra A, Shekhar S, Singh A K, et al · 2019
Earlier work this paper cites.
Towards VQA models that can read
Singh A, Natarajan V, Shah M, et al · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Jaume G, Ekenel H K, Thiran J P · 2019
Earlier work this paper cites.
ICDAR 2019 competition on scanned receipt OCR and information extraction
Huang Z, Chen K, He J, et al · 2019
Earlier work this paper cites.
Super-convergence: Very fast training of neural networks using large learning rates
Smith L N, Topin N · 2019
Earlier work this paper cites.
ICDAR 2019 competition on scene text visual question answering
Biten A F, Tito R, Mafla A, et al · 2019
Earlier work this paper cites.
LayoutLM: Pre-training of text and layout for document image understanding
Xu Y, Li M, Cui L, et al · 2020
Earlier work this paper cites.
Real-time scene text detection with differentiable binarization
Liao M, Wan Z, Yao C, et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy A, Beyer L, Kolesnikov A, et al · 2020
Cited alongside, same era.
LayoutxLM: Multimodal pre-training for multilingual visually-rich document understanding
Xu Y, Lv T, Cui L, et al · 2021
Cited alongside, same era.
DocFormer: End-to-end transformer for document understanding
Appalaraju S, Jasani B, Kota B U, et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford A, Kim J W, Hallacy C, et al · 2021
Cited alongside, same era.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu Z, Lin Y, Cao Y, et al · 2021
Cited alongside, same era.
mPLUG-DocOwl: Modularized multimodal large language model for document understanding
Ye J, Hu A, Xu H, et al · 2023
Closest in time.
Pix2Struct: Screenshot parsing as pretraining for visual language understanding
Lee K, Joshi M, Turc I R, et al · 2023
Closest in time.
On the hidden mystery of OCR in large multimodal models
Liu Y, Li Z, Li H, et al · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Touvron H, Lavril T, Izacard G, et al · 2023
Closest in time.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality
Chiang W L, Li Z, Lin Z, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mathew M, Karatzas D, Jawahar C · 2021
Cited alongside, same era.
OCR-free document understanding transformer
Kim G, Hong T, Yim M, et al · 2022
Cited alongside, same era.
LayoutLMv3: Pre-training for document ai with unified text and image masking
Huang Y, Lv T, Cui L, et al · 2022
Cited alongside, same era.
Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents
Hong T, Kim D, Ji M, et al · 2022
Cited alongside, same era.
Wukong-Reader: Multi-modal pre-training for fine-grained visual document understanding
Bai H, Liu Z, Meng X, et al · 2022
Cited alongside, same era.
Ernie-layout: Layout knowledge enhanced pre-training for visually-rich document understanding
Peng Q, Pan Y, Wang W, et al · 2022
Cited alongside, same era.
NomMer: Nominate synergistic context in vision transformer for visual recognition
Liu H, Jiang X, Li X, et al · 2022
Cited alongside, same era.
The devil is in the frequency: Geminated gestalt autoencoder for self-supervised visual pre-training
Liu H, Jiang X, Li X, et al · 2023
Closest in time.
Liu H, Li C, Wu Q, et al · 2023
Closest in time.
StrucTexTv2: Masked visual-textual prediction for document image pre-training
Yu Y, Li Y, Zhang C, et al · 2023
Closest in time.
KOSMOS-2: Grounding multimodal large language models to the world
Peng Z, Wang W, Dong L, et al · 2023
Closest in time.
MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Zhu D, Chen J, Shen X, et al · 2023
Closest in time.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li J, Li D, Savarese S, et al · 2023
Closest in time.
Improved baselines with visual instruction tuning
Liu H, Li C, Li Y, et al · 2023
Closest in time.
Towards improving document understanding: An exploration on text-grounding via MLLMs
Wang Y, Zhou W, Feng H, et al · 2023
Closest in time.
Visual information extraction in the wild: practical dataset and end-to-end solution
Kuang J, Hua W, Liang D, et al · 2023
Closest in time.
Qwen-VL: A frontier large vision-language model with versatile abilities
Bai J, Bai S, Yang S, et al · 2023
Closest in time.
Hierarchical multimodal transformers for multipage DocVQA
Tito R, Karatzas D, Valveny E · 2023
Closest in time.
Monkey: Image resolution and text label are important things for large multi-modal models
Li Z, Yang B, Liu Q, et al · 2024
Closest in time.
TextMonkey: An OCR-free large multimodal model for understanding document
Liu Y, Yang B, Liu Q, et al · 2024
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai W, Li J, Li D, et al · 2024
Closest in time.
Bliva: A simple multimodal LLM for better handling of text-rich visual questions
Hu W, Xu Y, Li Y, et al · 2024
Closest in time.
mPLUG-owl2: Revolutionizing multi-modal large language model with modality collaboration
Ye Q, Xu H, Ye J, et al · 2024
Closest in time.
Dong X, Zhang P, Zang Y, et al · 2024
Closest in time.