Fetching the paper…
Reading the bibliography…
Document understanding refers to automatically extract, analyze and comprehend information from various types of digital documents, such as a web page.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
P. Pasupat and P. Liang · 2015
Earlier work this paper cites.
Towards VQA models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Tabfact : A large-scale dataset for table-based fact verification
W. Chen, H. Wang, J. Chen, Y. Zhang, H. Wang, S. Li, X. Zhou, and W. Y. Wang · 2020
Earlier work this paper cites.
Textcaps: A dataset for image captioning with reading comprehension
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh · 2020
Earlier work this paper cites.
Deepform: Understand structured documents at scale, 2020
S. Svetlichnaya · 2020
Earlier work this paper cites.
Layoutlm: Pre-training of text and layout for document image understanding
Y. Xu, M. Li, L. Cui, S. Huang, F. Wei, and M. Zhou · 2020
Earlier work this paper cites.
DUE: end-to-end document understanding benchmark
L. Borchmann, M. Pietruszka, T. Stanislawek, D. Jurkiewicz, M. Turski, K. Szyndler, and F. Gralinski · 2021
Earlier work this paper cites.
Question-controlled text-aware image captioning
A. Hu, S. Chen, and Q. Jin · 2021
Earlier work this paper cites.
Docvqa: A dataset for VQA on document images
M. Mathew, D. Karatzas, and C. V. Jawahar · 2021
Earlier work this paper cites.
Kleister: Key information extraction datasets involving long documents with complex layouts
T. Stanislawek, F. Gralinski, A. Wróblewska, D. Lipinski, A. Kaliska, P. Rosalska, B. Topolski, and P. Biecek · 2021
Cited alongside, same era.
Visualmrc: Machine reading comprehension on document images
R. Tanaka, K. Nishida, and S. Yoshida · 2021
Cited alongside, same era.
TAP: text-aware pre-training for text-vqa and text-caption
Z. Yang, Y. Lu, J. Wang, X. Yin, D. Florêncio, L. Wang, C. Zhang, L. Zhang, and J. Luo · 2021
Cited alongside, same era.
End-to-end document recognition and understanding with dessurt
B. L. Davis, B. S. Morse, B. L. Price, C. Tensmeyer, C. Wigington, and V. I. Morariu · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
Layoutlmv3: Pre-training for document AI with unified text and image masking
Introducing chatgpt
OpenAI · 2022
Later among the works it cites.
BLOOM: A 176b-parameter open-access multilingual language model
T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilic, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé, J. Tow, A. M. Rush, S. Biderman, A. Webson, P. S. Ammanamanchi, T. Wang, B. Sagot, N. Muennighoff, A. V. del Moral, O. Ruwase, R. Bawden, S. Bekman, A. McMillan-Major, I. Beltagy, H. Nguyen, L. Saulnier, S. Tan, P. O. Suarez, V. Sanh, H. Laurençon, Y. Jernite, J. Launay, M. Mitchell, C. Raffel, A. Gokaslan, A. Simhi, A. Soroa, A. F. Aji, A. Alfassy, A. Rogers, A. K. Nitzav, C. Xu, C. Mou, C. Emezue, C. Klamm, C. Leong, D. van Strien, D. I. Adelani, and et al · 2022
Later among the works it cites.
Self-instruct: Aligning language model with self generated instructions
Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi · 2022
Later among the works it cites.
Unifying vision, text, and layout for universal document processing
Z. Tang, Z. Yang, G. Wang, Y. Fang, Y. Liu, C. Zhu, M. Zeng, C. Zhang, and M. Bansal · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Huang, T. Lv, L. Cui, Y. Lu, and F. Wei · 2022
Cited alongside, same era.
Ocr-free document understanding transformer
G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, and S. Park · 2022
Cited alongside, same era.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
K. Lee, M. Joshi, I. Turc, H. Hu, F. Liu, J. Eisenschlos, U. Khandelwal, P. Shaw, M. Chang, and K. Toutanova · 2022
Cited alongside, same era.
mplug: Effective and efficient vision-language learning by cross-modal skip-connections
C. Li, H. Xu, J. Tian, W. Wang, M. Yan, B. Bi, J. Ye, H. Chen, G. Xu, Z. Cao, J. Zhang, S. Huang, F. Huang, J. Zhou, and L. Si · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. R. Joty, and E. Hoque · 2022
Cited alongside, same era.
Infographicvqa
M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. V. Jawahar · 2022
Cited alongside, same era.
H. Liu, C. Li, Q. Wu, and Y. J. Lee
Cited in the paper.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Closest in time.
Vicuna: An open chatbot impressing gpt-4
Vicuna · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan · 2023
Closest in time.
MM-REACT: prompting chatgpt for multimodal reasoning and action
Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, C. Li, Y. Xu, H. Chen, J. Tian, Q. Qi, J. Zhang, and F. Huang · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Closest in time.