Fetching the paper…
Reading the bibliography…
We study the problem of completing various visual document understanding (VDU) tasks, e.g., question answering and information extraction, on real-world documents through human-written instructions.
The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
Kuznetsova, A.; Rom, H.; Alldrin, N.; Uijlings, J.; Krasin, I.; Pont-Tuset, J.; Kamali, S.; Popov, S.; Malloci, M.; Kolesnikov, A.; et al. 2020 · 1981
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval
Harley, A. W.; Ufkes, A.; and Derpanis, K. G. 2015 · 2015
Earlier work this paper cites.
A Diagram Is Worth A Dozen Images
Kembhavi, A.; Salvato, M.; Kolve, E.; Seo, M.; Hajishirzi, H.; and Farhadi, A. 2016 · 2016
Earlier work this paper cites.
Are You Smarter Than a Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension
Kembhavi, A.; Seo, M. J.; Schwenk, D.; Choi, J.; Farhadi, A.; and Hajishirzi, H. 2017 · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.; Shamma, D. A.; Bernstein, M. S.; and Fei-Fei, L. 2017 · 2017
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Scene Text Visual Question Answering
Biten, A. F.; Tito, R.; Mafla, A.; i Bigorda, L. G.; Rusiñol, M.; Jawahar, C. V.; Valveny, E.; and Karatzas, D. 2019 · 2019
Earlier work this paper cites.
ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
Huang, Z.; Chen, K.; He, J.; Bai, X.; Karatzas, D.; Lu, S.; and Jawahar, C. 2019 · 2019
Earlier work this paper cites.
FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents
Jaume, G.; Ekenel, H. K.; and Thiran, J.-P. 2019 · 2019
Earlier work this paper cites.
OCR-VQA: Visual Question Answering by Reading Text in Images
Mishra, A.; Shekhar, S.; Singh, A. K.; and Chakraborty, A. 2019 · 2019
Earlier work this paper cites.
CORD: A Consolidated Receipt Dataset for Post-OCR Parsing
Park, S.; Shin, S.; Lee, B.; Lee, J.; Surh, J.; Seo, M.; and Lee, H. 2019 · 2019
Earlier work this paper cites.
Towards VQA Models That Can Read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Earlier work this paper cites.
DocBank: A Benchmark Dataset for Document Layout Analysis
Li, M.; Xu, Y.; Cui, L.; Huang, S.; Wei, F.; Li, Z.; and Zhou, M. 2020 · 2020
Earlier work this paper cites.
LayoutLM: Pre-training of Text and Layout for Document Image Understanding
Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; and Zhou, M. 2020 · 2020
Earlier work this paper cites.
Docformer: End-to-End Transformer for Document Understanding
Appalaraju, S.; Jasani, B.; Kota, B. U.; Xie, Y.; and Manmatha, R. 2021 · 2021
Earlier work this paper cites.
DUE: End-to-End Document Understanding Benchmark
Borchmann, Ł.; Pietruszka, M.; Stanislawek, T.; Jurkiewicz, D.; Turski, M.; Szyndler, K.; and Graliński, F. 2021 · 2021
Earlier work this paper cites.
WebSRC: A Dataset for Web-Based Structural Reading Comprehension
Chen, X.; Zhao, Z.; Chen, L.; Ji, J.; Zhang, D.; Luo, A.; Xiong, Y.; and Yu, K. 2021 · 2021
Earlier work this paper cites.
SciCap: Generating Captions for Scientific Figures
Hsu, T.-Y.; Giles, C. L.; and Huang, T.-H. 2021 · 2021
Cited alongside, same era.
IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Lu, P.; Qiu, L.; Chen, J.; Xia, T.; Zhao, Y.; Zhang, W.; Yu, Z.; Liang, X.; and Zhu, S.-C. 2021 · 2021
Cited alongside, same era.
DocVQA: A Dataset for VQA on Document Images
Mathew, M.; Karatzas, D.; and Jawahar, C. V. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models from Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Spatial Dual-Modality Graph Reasoning for Key Information Extraction
Sun, H.; Kuang, Z.; Yue, X.; Lin, C.; and Zhang, W. 2021 · 2021
Cited alongside, same era.
Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Mishra, S.; Khashabi, D.; Baral, C.; and Hajishirzi, H. 2022 · 2022
Later among the works it cites.
DocLayNet: A Large Human-Annotated Dataset for Document-Layout Segmentation
Pfitzmann, B.; Auer, C.; Dolfi, M.; Nassar, A. S.; and Staar, P. 2022 · 2022
Later among the works it cites.
Recognition-free Question Answering on Handwritten Document Collections
Tüselmann, O.; Müller, F.; Wolf, F.; and Fink, G. A. 2022 · 2022
Later among the works it cites.
Towards Complex Document Understanding by Discrete Reasoning
Zhu, F.; Lei, W.; Feng, F.; Wang, C.; Zhang, H.; and Chua, T.-S. 2022 · 2022
Later among the works it cites.
DocFormerv2: Local Features for Document Understanding
Appalaraju, S.; Tang, P.; Dong, Q.; Sankaran, N.; Zhou, Y.; and Manmatha, R. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
VisualMRC: Machine Reading Comprehension on Document Images
Tanaka, R.; Nishida, K.; and Yoshida, S. 2021 · 2021
Cited alongside, same era.
Screen2words: Automatic mobile UI summarization with multimodal learning
Wang, B.; Li, G.; Zhou, X.; Chen, Z.; Grossman, T.; and Li, Y. 2021 · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2021 · 2021
Cited alongside, same era.
LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding
Xu, Y.; Xu, Y.; Lv, T.; Cui, L.; Wei, F.; Wang, G.; Lu, Y.; Florêncio, D. A. F.; Zhang, C.; Che, W.; Zhang, M.; and Zhou, L. 2021 · 2021
Cited alongside, same era.
Flamingo: a Visual Language Model for Few-Shot Learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Cited alongside, same era.
PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts
Bach, S.; Sanh, V.; Yong, Z. X.; Webson, A.; Raffel, C.; Nayak, N. V.; Sharma, A.; Kim, T.; Bari, M. S.; Fevry, T.; Alyafeai, Z.; Dey, M.; Santilli, A.; Sun, Z.; Ben-david, S.; Xu, C.; Chhablani, G.; Wang, H.; Fries, J.; Al-shaibani, M.; Sharma, S.; Thakker, U.; Almubarak, K.; Tang, X.; Radev, D.; Jiang, M. T.-j.; and Rush, A. 2022 · 2022
Cited alongside, same era.
Scaling Instruction-Finetuned Language Models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, E.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2022 · 2022
Cited alongside, same era.
Chen, X.; Djolonga, J.; Padlewski, P.; Mustafa, B.; Changpinyo, S.; Wu, J.; Ruiz, C. R.; Goodman, S.; Wang, X.; Tay, Y.; et al. 2023 · 2023
Later among the works it cites.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Later among the works it cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Later among the works it cites.
Document Understanding Dataset and Evaluation (DUDE)
Landeghem, J.; Tito, R.; Borchmann, Ł.; Pietruszka, M.; Józiak, P.; Powalski, R.; Jurkiewicz, D.; Coustaty, M.; Ackaert, B.; Valveny, E.; et al. 2023 · 2023
Later among the works it cites.
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
Lee, K.; Joshi, M.; Turc, I. R.; Hu, H.; Liu, F.; Eisenschlos, J. M.; Khandelwal, U.; Shaw, P.; Chang, M.-W.; and Toutanova, K. 2023 · 2023
Later among the works it cites.
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Later among the works it cites.
The FLAN collection: Designing Data and Methods for Effective Instruction Tuning
Longpre, S.; Hou, L.; Vu, T.; Webson, A.; Chung, H. W.; Tay, Y.; Zhou, D.; Le, Q. V.; Zoph, B.; Wei, J.; et al. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
DocILE Benchmark for Document Information Localization and Extraction
Šimsa, Š.; Šulc, M.; Uřičář, M.; Patel, Y.; Hamdi, A.; Kocián, M.; Skalickỳ, M.; Matas, J.; Doucet, A.; Coustaty, M.; and Karatzas, D. 2023 · 2023
Later among the works it cites.
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Tanaka, R.; Nishida, K.; Nishida, K.; Hasegawa, T.; Saito, I.; and Saito, K. 2023 · 2023
Later among the works it cites.
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Xu, Z.; Shen, Y.; and Huang, L. 2023 · 2023
Later among the works it cites.
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Zhang, Y.; Zhang, R.; Gu, J.; Zhou, Y.; Lipka, N.; Yang, D.; and Sun, T. 2023 · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.