Fetching the paper…
Reading the bibliography…
Document Visual Question Answering (VQA) demands robust integration of text detection, recognition, and spatial reasoning to interpret complex document layouts.
A normalized levenshtein distance metric
Yujian, L. and Bo, L · 2007
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Shi, B., Bai, X., and Yao, C · 2016
Earlier work this paper cites.
Rico: A mobile app dataset for building data-driven design applications
Deka, B., Huang, Z., Franzen, C., Hibschman, J., Afergan, D., Li, Y., Nichols, J., and Kumar, R · 2017
Earlier work this paper cites.
Icdar2019 competition on scanned receipt ocr and information extraction
Huang, Z., Chen, K., He, J., Bai, X., Karatzas, D., Lu, S., and Jawahar, C · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Jaume, G., Ekenel, H. K., and Thiran, J.-P · 2019
Earlier work this paper cites.
Show, attend and read: A simple and strong baseline for irregular text recognition
Li, H., Wang, P., Shen, C., and Zhang, G · 2019
Earlier work this paper cites.
Cord: a consolidated receipt dataset for post-ocr parsing
Park, S., Shin, S., Lee, B., Lee, J., Surh, J., Seo, M., and Lee, H · 2019
Earlier work this paper cites.
Generalized intersection over union: A metric and a loss for bounding box regression
Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S · 2019
Earlier work this paper cites.
Real-time scene text detection with differentiable binarization
Liao, M., Wan, Z., Yao, C., Chen, K., and Bai, X · 2020
Earlier work this paper cites.
On the general value of evidence, and bilingual scene-text visual question answering
Wang, X., Liu, Y., Shen, C., Ng, C. C., Luo, C., Jin, L., Chan, C. S., Hengel, A. v. d., and Wang, L · 2020
Earlier work this paper cites.
Vision transformer for fast and efficient scene text recognition
Atienza, R · 2021
Earlier work this paper cites.
Fast: Faster arbitrarily-shaped text detector with minimalist kernel representation
Chen, Z., Wang, J., Wang, W., Chen, G., Xie, E., Luo, P., and Lu, T · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Master: Multi-aspect non-local network for scene text recognition
Lu, N., Yu, W., Qi, X., Chen, Y., Gong, P., Xiao, R., and Bai, X · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Mathew, M., Karatzas, D., and Jawahar, C · 2021
Earlier work this paper cites.
Scene text recognition with permuted autoregressive sequence models
Bautista, D. and Atienza, R · 2022
Cited alongside, same era.
Layoutlmv3: Pre-training for document ai with unified text and image masking
Huang, Y., Lv, T., Cui, L., Lu, Y., and Wei, F · 2022
Cited alongside, same era.
Ocr-free document understanding transformer
Kim, G., Hong, T., Yim, M., Nam, J., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., and Park, S · 2022
Cited alongside, same era.
Maskocr: Text recognition with masked encoder-decoder pretraining
Lyu, P., Zhang, C., Liu, S., Qiao, M., Xu, Y., Wu, L., Yao, K., Han, J., Ding, E., and Wang, J · 2022
Cited alongside, same era.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., Li, B., Luo, P., Lu, T., Qiao, Y., and Dai, J · 2023
Cited alongside, same era.
Agrawal, P., Antoniak, S., Hanna, E. B., Chaplot, D., Chudnovsky, J., Garg, S., Gervet, T., Ghosh, S., Héliou, A., Jacob, P., et al · 2024
Closest in time.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Dtrocr: Decoder-only transformer for optical character recognition
Fujitake, M · 2024
Closest in time.
Lora+: Efficient low rank adaptation of large models
Hayou, S., Ghosh, N., and Yu, B · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Icl-d3ie: In-context learning with diverse demonstrations updating for document information extraction
He, J., Wang, L., Hu, Y., Liu, N., Liu, H., Xu, X., and Shen, H. T · 2023
Cited alongside, same era.
Kim, G., Lee, H., Kim, D., Jung, H., Park, S., Kim, Y., Yun, S., Kil, T., Lee, B., and Park, S · 2023
Cited alongside, same era.
Trocr: Transformer-based optical character recognition with pre-trained models
Li, M., Lv, T., Chen, J., Cui, L., Lu, Y., Florencio, D., Zhang, C., Li, Z., and Wei, F · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning, 2023
Liu, H., Li, C., Li, Y., and Lee, Y. J · 2023
Cited alongside, same era.
Unifying vision, text, and layout for universal document processing
Tang, Z., Yang, Z., Wang, G., Fang, Y., Liu, Y., Zhu, C., Zeng, M., Zhang, C., and Bansal, M · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Layout and task aware instruction prompt for zero-shot document image question answering
Wang, W., Li, Y., Ou, Y., and Zhang, Y · 2023
Cited alongside, same era.
Huang, Y., Sun, L., Wang, H., Wu, S., Zhang, Q., Li, Y., Gao, C., Huang, Y., Lyu, W., Zhang, Y., et al · 2024
Closest in time.
From image to language: A critical analysis of visual question answering (vqa) approaches, challenges, and opportunities
Ishmam, M. F., Shovon, M. S. H., Mridha, M. F., and Dey, N · 2024
Closest in time.
Llava-onevision: Easy visual task transfer
Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Li, Y., Liu, Z., and Li, C · 2024
Closest in time.
Liao, W., Wang, J., Li, H., Wang, C., Huang, J., and Jin, L · 2024
Closest in time.
Lu, J., Yu, H., Wang, Y., Ye, Y., Tang, J., Yang, Z., Wu, B., Liu, Q., Feng, H., Wang, H., et al · 2024
Closest in time.
Layoutllm: Layout instruction tuning with large language models for document understanding
Luo, C., Shen, Y., Zhu, Z., Zheng, Q., Yu, Z., and Yao, C · 2024
Closest in time.
Instructdoc: A dataset for zero-shot generalization of visual document understanding with instructions
Tanaka, R., Iki, T., Nishida, K., Saito, K., and Suzuki, J · 2024
Closest in time.
Omniparser: A unified framework for text spotting key information extraction and table recognition
Wan, J., Song, S., Yu, W., Liu, Y., Cheng, W., Huang, F., Bai, X., Yao, C., and Yang, Z · 2024
Closest in time.
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et al · 2025
Closest in time.