Fetching the paper…
Reading the bibliography…
Multimodal document understanding is a challenging task to process and comprehend large amounts of textual and visual information.
Huang, Z., Chen, K., He, J., Bai, X., Karatzas, D., Lu, S., Jawahar, C.: Icdar2019 competition on scanned receipt ocr and information extraction. In: ICDAR. pp. 1516–1520 (2019)
2019
Earlier work this paper cites.
Mishra, A., Shekhar, S., Singh, A.K., Chakraborty, A.: Ocr-vqa: Visual question answering by reading text in images. In: ICDAR. pp. 947–952 (2019)
2019
Earlier work this paper cites.
Park, S., Shin, S., Lee, B., Lee, J., Surh, J., Seo, M., Lee, H.: Cord: A consolidated receipt dataset for post-ocr parsing. In: Workshop on Document Intelligence at NeurIPS (2019)
2019
Earlier work this paper cites.
Zhong, X., Tang, J., Yepes, A.J.: Publaynet: largest dataset ever for document layout analysis. In: 2019 International conference on document analysis and recognition (ICDAR). pp. 1015–1022. IEEE (2019)
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Zaheer, M., Guruganesh, G., Dubey, K.A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al.: Big bird: Transformers for longer sequences. Advances in neural information processing systems 33
2020
Earlier work this paper cites.
Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N.A., Gardner, M.: A dataset of information-seeking questions and answers anchored in research papers. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 4599–4610 (2021)
2021
Earlier work this paper cites.
Huang, L., Cao, S., Parulian, N., Ji, H., Wang, L.: Efficient attentions for long document summarization. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. pp. 1419–1436 (2021)
2021
Earlier work this paper cites.
Mathew, M., Karatzas, D., Jawahar, C.V.: Docvqa: A dataset for vqa on document images. In: WACV. pp. 2200–2209 (2021)
2021
Earlier work this paper cites.
Huang, Y., Lv, T., Cui, L., Lu, Y., Wei, F.: Layoutlmv3: Pre-training for document ai with unified text and image masking. In: Proceedings of the 30th ACM International Conference on Multimedia. pp. 4083–4091 (2022)
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., Jawahar, C.: Infographicvqa. In: WACV. pp. 1697–1706 (2022)
2022
Earlier work this paper cites.
Pfitzmann, B., Auer, C., Dolfi, M., Nassar, A.S., Staar, P.: Doclaynet: a large human-annotated dataset for document-layout segmentation. In: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. pp. 3743–3751 (2022)
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
OpenAI: Gpt-4v(ision) system card. https://openai.com/contributions/gpt-4v (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Šimsa, Š., Šulc, M., Uřičář, M., Patel, Y., Hamdi, A., Kocián, M., Skalickỳ, M., Matas, J., Doucet, A., Coustaty, M., et al.: Docile benchmark for document information localization and extraction pp. 147–166 (2023)
2023
Earlier work this paper cites.
Tanaka, R., Nishida, K., Nishida, K., Hasegawa, T., Saito, I., Saito, K.: Slidevqa: A dataset for document visual question answering on multiple images. In: AAAI. pp. 13636–13645 (2023)
2023
Earlier work this paper cites.
Tito, R., Karatzas, D., Valveny, E.: Hierarchical multimodal transformers for multipage docvqa. Pattern Recognition 144
2023
Earlier work this paper cites.
Tworkowski, S., Staniszewski, K., Pacek, M.a., Wu, Y., Michalewski, H., Mił oś, P.: Focused transformer: Contrastive training for context scaling. In: Advances in Neural Information Processing Systems. vol. 36, pp. 42661–42688 (2023)
2023
Earlier work this paper cites.
Van Landeghem, J., Tito, R., Borchmann, Ł., Pietruszka, M., Joziak, P., Powalski, R., Jurkiewicz, D., Coustaty, M., Anckaert, B., Valveny, E., et al.: Document understanding dataset and evaluation (dude). In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 19528–19540 (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Grobid. https://github.com/kermitt2/grobid (2008–2024)
2024
Cited alongside, same era.
Appalaraju, S., Tang, P., Dong, Q., Sankaran, N., Zhou, Y., Manmatha, R.: Docformerv2: Local features for document understanding. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 38, pp. 709–718 (2024)
2024
Cited alongside, same era.
Blau, T., Fogel, S., Ronen, R., Golts, A., Ganz, R., Ben Avraham, E., Aberdam, A., Tsiper, S., Litman, R.: Gram: Global reasoning for multi-page vqa. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 15598–15607 (2024)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
Wei, H., Kong, L., Chen, J., Zhao, L., Ge, Z., Yang, J., Sun, J., Han, C., Zhang, X.: Vary: Scaling up the vision vocabulary for large vision-language model. In: European Conference on Computer Vision. pp. 408–424. Springer (2024)
2024
Closest in time.
Wu, W., Cai, Y., Shen, C., Zhang, D., Fu, Y., Zhou, H., Luo, P.: End-to-end video text spotting with transformer. International Journal of Computer Vision 132
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, Y., Qian, S., Tang, H., Lai, X., Liu, Z., Han, S., Jia, J.: LongloRA: Efficient fine-tuning of long-context large language models. In: The Twelfth International Conference on Learning Representations (2024)
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Zhang, J., Yu, Y., Zhang, Y.: Cream: Coarse-to-fine retrieval and multi-modal efficient tuning for document vqa. In: Proceedings of the 32nd ACM International Conference on Multimedia. pp. 925–934 (2024)
2024
Closest in time.
Zhao, X., Feng, W., Zhang, Z., Lv, J., Zhu, X., Lin, Z., Hu, J., Shao, J.: Cbnet: A plug-and-play network for segmentation-based scene text detection. International Journal of Computer Vision 132
2024
Closest in time.
Zheng, T., Chen, Z., Fang, S., Xie, H., Jiang, Y.G.: Cdistnet: Perceiving multi-domain character distance for robust text recognition. International Journal of Computer Vision 132
2024
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Chen, J., Zhang, R., Zhou, Y., Yu, T., Dernoncourt, F., Gu, J., Rossi, R.A., Chen, C., Sun, T.: SV-RAG: LoRA-contextualizing adaptation of MLLMs for long document understanding. In: The Thirteenth International Conference on Learning Representations (2025), https://openreview.net/forum?id=FDaHjwInXO
2025
Closest in time.
Faysse, M., Sibille, H., Wu, T., Omrani, B., Viaud, G., HUDELOT, C., Colombo, P.: Colpali: Efficient document retrieval with vision language models. In: The Thirteenth International Conference on Learning Representations (2025), https://openreview.net/forum?id=ogjBpZ8uSi
2025
Closest in time.
Liu, Y., Xie, X., Liu, Y., Bai, X.: Multi-scenario overlapping text segmentation with depth awareness. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 17454–17463 (2025)
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Yin, L., Xie, X., Li, Z., Bai, X., Liu, Y.: Mstar: Box-free multi-query scene text retrieval with attention recycling. In: The Thirty-ninth Annual Conference on Neural Information Processing Systems (2025)
2025
Closest in time.
2025
Closest in time.
Yu, W., Yang, Z., Liu, Y., Bai, X.: Docthinker: Explainable multimodal large language models with rule-based reinforcement learning for document understanding. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 837–847 (2025)
2025
Closest in time.
Zhang, E., Li, Y., Liu, Y., Zhu, Y., Bai, X.: Towards comprehensive lecture slides understanding: Large-scale dataset and effective method. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 4455–4464 (2025)
2025
Closest in time.
Zhang, P., Zhang, J., Cao, J., Li, H., Jin, L.: Smaller but better: Unifying layout generation with smaller large language models. International Journal of Computer Vision pp. 1–27 (2025)
2025
Closest in time.
Liu, Y., Yang, B., Liu, Q., Li, Z., Ma, Z., Zhang, S., Bai, X.: Textmonkey: An ocr-free large multimodal model for understanding document. IEEE Transactions on Pattern Analysis and Machine Intelligence (2026)
2026
Closest in time.
Yan, H., Liu, Y., Liu, X., Zhang, Y., Liao, M., Wu, J., Chen, W., Bai, X.: Docseeker: Structured visual reasoning with evidence grounding for long document understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2026)
2026
Closest in time.