Fetching the paper…
Reading the bibliography…
Layout-aware pre-trained models has achieved significant progress on document image question answering.
CUTIE: Learning to Understand Documents with Convolutional Universal Text Information Extractor
Zhao, X.; Niu, E.; Wu, Z.; and Wang, X. 2019 · 1903
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019b · 1907
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
Building a Test Collection for Complex Document Information Processing
Lewis, D.; Agam, G.; Argamon, S.; Frieder, O.; Grossman, D.; and Heard, J. 2006 · 2006
Earlier work this paper cites.
Big Bird: Transformers for Longer Sequences
Zaheer, M.; Guruganesh, G.; Dubey, A.; Ainslie, J.; Alberti, C.; Ontanon, S.; Pham, P.; Ravula, A.; Wang, Q.; Yang, L.; and Ahmed, A. 2021 · 2007
Earlier work this paper cites.
Compositional Semantic Parsing on Semi-Structured Tables
Pasupat, P.; and Liang, P. 2015 · 2015
Earlier work this paper cites.
Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Networks
Yang, X.; Yumer, E.; Asente, P.; Kraley, M.; Kifer, D.; and Giles, C. L. 2017 · 2017
Earlier work this paper cites.
Chargrid: Towards Understanding 2D Documents
Katti, A. R.; Reisswig, C.; Guder, C.; Brarda, S.; Bickel, S.; Höhne, J.; and Faddoul, J. B. 2018 · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2018 · 2018
Earlier work this paper cites.
Scene Text Visual Question Answering
Biten, A. F.; Tito, R.; Mafla, A.; Gomez, L.; Rusiñol, M.; Valveny, E.; Jawahar, C. V.; and Karatzas, D. 2019 · 2019
Earlier work this paper cites.
BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding
Denk, T. I.; and Reisswig, C. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Graph Convolution for Multimodal Information Extraction from Visually Rich Documents
Liu, X.; Gao, F.; Zhang, Q.; and Zhao, H. 2019a · 2019
Earlier work this paper cites.
GraphIE: A Graph-Based Framework for Information Extraction
Qian, Y.; Santus, E.; Jin, Z.; Guo, J.; and Barzilay, R. 2019 · 2019
Earlier work this paper cites.
Deterministic Routing between Layout Abstractions for Multi-Scale Classification of Visually Rich Documents
Sarkhel, R.; and Nandi, A. 2019 · 2019
Earlier work this paper cites.
UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
Bao, H.; Dong, L.; Wei, F.; Wang, W.; Yang, N.; Liu, X.; Wang, Y.; Gao, J.; Piao, S.; Zhou, M.; and Hon, H.-W. 2020 · 2020
Earlier work this paper cites.
Named Entity Recognition and Relation Extraction with Graph Neural Networks in Semi Structured Documents
Carbonell, M.; Riba, P.; Villegas, M.; Fornes, A.; and Llados, J. 2021 · 2020
Earlier work this paper cites.
Representation Learning for Information Extraction from Form-like Documents
Majumder, B. P.; Potti, N.; Tata, S.; Wendt, J. B.; Zhao, Q.; and Najork, M. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
DocStruct: A Multimodal Method to Extract Hierarchy Structure in Document for General Form Understanding
Wang, Z.; Zhan, M.; Liu, X.; and Liang, D. 2020 · 2020
Earlier work this paper cites.
Robust Layout-aware IE for Visually Rich Documents with Pre-trained Language Models
Wei, M.; He, Yi.; and Zhang, Q. 2020 · 2020
Earlier work this paper cites.
LayoutLM: Pre-training of Text and Layout for Document Image Understanding
Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; and Zhou, M. 2020 · 2020
Earlier work this paper cites.
PICK: Processing Key Information Extraction from Documents Using Improved Graph Learning-Convolutional Networks
Yu, W.; Lu, N.; Qi, X.; Gong, P.; and Xiao, R. 2020 · 2020
Cited alongside, same era.
TRIE: End-to-End Text Reading and Information Extraction for Document Understanding
Zhang, P.; Xu, Y.; Cheng, Z.; Pu, S.; Lu, J.; Qiao, L.; Niu, Y.; and Wu, F. 2020 · 2020
Cited alongside, same era.
DocFormer: End-to-End Transformer for Document Understanding
Appalaraju, S.; Jasani, B.; Kota, B. U.; Xie, Y.; and Manmatha, R. 2021 · 2021
Cited alongside, same era.
DUE: End-to-End Document Understanding Benchmark
Borchmann, Ł.; Pietruszka, M.; Stanislawek, T.; Jurkiewicz, D.; Turski, M.; Szyndler, K.; and Graliński, F. 2021 · 2021
Cited alongside, same era.
LAMBERT: Layout-Aware Language Modeling for Information Extraction
Garncarek, Ł.; Powalski, R.; Stanisławek, T.; Topolski, B.; Halama, P.; Turski, M.; and Graliński, F. 2021 · 2021
Cited alongside, same era.
Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Mishra, S.; Khashabi, D.; Baral, C.; and Hajishirzi, H. 2022 · 2022
Later among the works it cites.
Training Language Models to Follow Instructions with Human Feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Gray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Later among the works it cites.
ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding
Peng, Q.; Pan, Y.; Wang, W.; Luo, B.; Zhang, Z.; Huang, Z.; Cao, Y.; Yin, W.; Chen, Y.; Zhang, Y.; Feng, S.; Sun, Y.; Tian, H.; Wu, H.; and Wang, H. 2022 · 2022
Later among the works it cites.
Multitask Prompted Training Enables Zero-Shot Task Generalization
Sanh, V.; Webson, A.; Raffel, C.; Bach, S. H.; Sutawika, L.; Alyafeai, Z.; Chaffin, A.; Stiegler, A.; Scao, T. L.; Raja, A.; Dey, M.; Bari, M. S.; Xu, C.; Thakker, U.; Sharma, S. S.; Szczechla, E.; Kim, T.; Chhablani, G.; Nayak, N.; Datta, D.; Chang, J.; Jiang, M. T.-J.; Wang, H.; Manica, M.; Shen, S.; Yong, Z. X.; Pandey, H.; Bawden, R.; Wang, T.; Neeraj, T.; Rozen, J.; Sharma, A.; Santilli, A.; Fevry, T.; Fries, J. A.; Teehan, R.; Bers, T.; Biderman, S.; Gao, L.; Wolf, T.; and Rush, A. M. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spatial Dependency Parsing for Semi-Structured Document Information Extraction
Hwang, W.; Yim, J.; Park, S.; Yang, S.; and Seo, M. 2021 · 2021
Cited alongside, same era.
StructuralLM: Structural Pre-training for Form Understanding
Li, C.; Bi, B.; Yan, M.; Wang, W.; Huang, S.; Huang, F.; and Si, L. 2021a · 2021
Cited alongside, same era.
SelfDoc: Self-Supervised Document Representation Learning
Li, P.; Gu, J.; Kuen, J.; Morariu, V. I.; Zhao, H.; Jain, R.; Manjunatha, V.; and Liu, H. 2021b · 2021
Cited alongside, same era.
StrucTexT: Structured Text Understanding with Multi-Modal Transformers
Li, Y.; Qian, Y.; Yu, Y.; Qin, X.; Zhang, C.; Liu, Y.; Yao, K.; Han, J.; Liu, J.; and Ding, E. 2021c · 2021
Cited alongside, same era.
ViBERTgrid: A Jointly Trained Multi-modal 2D Document Representation for Key Information Extraction from Documents
Lin, W.; Gao, Q.; Sun, L.; Zhong, Z.; Hu, K.; Ren, Q.; and Huo, Q. 2021 · 2021
Cited alongside, same era.
DocVQA: A Dataset for VQA on Document Images
Mathew, M.; Karatzas, D.; and Jawahar, C. V. 2021 · 2021
Cited alongside, same era.
Towards Robust Visual Information Extraction in Real World: New Dataset and Novel Solution
Wang, J.; Liu, C.; Jin, L.; Tang, G.; Zhang, J.; Zhang, S.; Wang, Q.; Wu, Y.; and Cai, M. 2021 · 2021
Cited alongside, same era.
Thoppilan, R.; De Freitas, D.; Hall, J.; Shazeer, N.; Kulshreshtha, A.; Cheng, H.-T.; Jin, A.; Bos, T.; Baker, L.; Du, Y.; Li, Y.; Lee, H.; Zheng, H. S.; Ghafouri, A.; Menegali, M.; Huang, Y.; Krikun, M.; Lepikhin, D.; Qin, J.; Chen, D.; Xu, Y.; Chen, Z.; Roberts, A.; Bosma, M.; Zhao, V.; Zhou, Y.; Chang, C.-C.; Krivokon, I.; Rusch, W.; Pickett, M.; Srinivasan, P.; Man, L.; Meier-Hellstern, K.; Morris, M. R.; Doshi, T.; Santos, R. D.; Duke, T.; Soraker, J.; Zevenbergen, B.; Prabhakaran, V.; Diaz, M.; Hutchinson, B.; Olson, K.; Molina, A.; Hoffman-John, E.; Lee, J.; Aroyo, L.; Rajakumar, R.; Butryna, A.; Lamm, M.; Kuzmina, V.; Fenton, J.; Cohen, A.; Bernstein, R.; Kurzweil, R.; Aguera-Arcas, B.; Cui, C.; Croak, M.; Chi, E.; and Le, Q. 2022 · 2022
Later among the works it cites.
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Wang, Y.; Mishra, S.; Alipoormolabashi, P.; Kordi, Y.; Mirzaei, A.; Arunkumar, A.; Ashok, A.; Dhanasekaran, A. S.; Naik, A.; Stap, D.; Pathak, E.; Karamanolakis, G.; Lai, H. G.; Purohit, I.; Mondal, I.; Anderson, J.; Kuznia, K.; Doshi, K.; Patel, M.; Pal, K. K.; Moradshahi, M.; Parmar, M.; Purohit, M.; Varshney, N.; Kaza, P. R.; Verma, P.; Puri, R. S.; Karia, R.; Sampat, S. K.; Doshi, S.; Mishra, S.; Reddy, S.; Patro, S.; Dixit, T.; Shen, X.; Baral, C.; Choi, Y.; Smith, N. A.; Hajishirzi, H.; and Khashabi, D. 2022c · 2022
Later among the works it cites.
Finetuned Language Models Are Zero-Shot Learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022 · 2022
Later among the works it cites.
MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
Xu, Z.; Shen, Y.; and Huang, L. 2022 · 2022
Later among the works it cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Closest in time.
DocParser: End-to-end OCR-free Information Extraction from Visually Rich Documents
Dhouib, M.; Bettaieb, G.; and Shabou, A. 2023 · 2023
Closest in time.
He, J.; Wang, L.; Hu, Y.; Liu, N.; Liu, H.; Xu, X.; and Shen, H. T. 2023 · 2023
Closest in time.
OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization
Iyer, S.; Lin, X. V.; Pasunuru, R.; Mihaylov, T.; Simig, D.; Yu, P.; Shuster, K.; Wang, T.; Liu, Q.; Koura, P. S.; Li, X.; O’Horo, B.; Pereyra, G.; Wang, J.; Dewan, C.; Celikyilmaz, A.; Zettlemoyer, L.; and Stoyanov, V. 2023 · 2023
Closest in time.
The Flan Collection: Designing Data and Methods for Effective Instruction Tuning
Longpre, S.; Hou, L.; Vu, T.; Webson, A.; Chung, H. W.; Tay, Y.; Zhou, D.; Le, Q. V.; Zoph, B.; Wei, J.; and Roberts, A. 2023 · 2023
Closest in time.
GeoLayoutLM: Geometric Pre-training for Visual Information Extraction
Luo, C.; Cheng, C.; Zheng, Q.; and Yao, C. 2023 · 2023
Closest in time.
GPT-4 Technical Report
Open AI. 2023 · 2023
Closest in time.
Peng, B.; Li, C.; He, P.; Galley, M.; and Gao, J. 2023 · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA Model
Rohan Taori; Ishaan Gulrajani; Tianyi Zhang; Yann Dubois; Xuechen Li; Carlos Guestrin; and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
Hierarchical Multimodal Transformers for Multi-Page DocVQA
Tito, R.; Karatzas, D.; and Valveny, E. 2023 · 2023
Closest in time.
WizardLM: Empowering Large Language Models to Follow Complex Instructions
Xu, C.; Sun, Q.; Zheng, K.; Geng, X.; Zhao, P.; Feng, J.; Tao, C.; and Jiang, D. 2023 · 2023
Closest in time.
Large Language Models Are Human-Level Prompt Engineers
Zhou, Y.; Muresanu, A. I.; Han, Z.; Paster, K.; Pitis, S.; Chan, H.; and Ba, J. 2023 · 2023
Closest in time.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Closest in time.