Fetching the paper…
Reading the bibliography…
Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest.
1909
Earlier work this paper cites.
1912
Earlier work this paper cites.
T.-L. Yuan, Z. Zhu, K. Xu, C.-J. Li, T.-J. Mu, and S.-M. Hu, “A large chinese text dataset in the wild,” Journal of Computer Science and Technology , vol. 34, no. 3, pp. 509–521, 2019. [Online]. Available: https://jcst.ict.ac.cn/en/article/doi/10.1007/s11390-019-1923-y
1923
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of Annual Meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
U.-V. Marti and H. Bunke, “The iam-database: an english sentence database for offline handwriting recognition,” International Journal on Document Analysis and Recognition , vol. 5, pp. 39–46, 2002
2002
Earlier work this paper cites.
S. Banerjee and A. Lavie, “METEOR: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
A. Mishra, K. Alahari, and C. V. Jawahar, “Scene text recognition using higher order language priors,” in British Machine Vision Conference , 2012
2012
Earlier work this paper cites.
C. Yao, X. Bai, W. Liu, Y. Ma, and Z. Tu, “Detecting texts of arbitrary orientations in natural images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . IEEE, 2012, pp. 1083–1090
2012
Earlier work this paper cites.
D. Karatzas, F. Shafait, S. Uchida, M. Iwamura, L. G. i. Bigorda, S. R. Mestre, J. Mas, D. F. Mota, J. A. Almazàn, and L. P. de las Heras, “Icdar 2013 robust reading competition,” in Proceedings of International Conference on Document Analysis and Recognition , 2013, pp. 1484–1493
2013
Earlier work this paper cites.
T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, “Recognizing text with perspective distortion in natural scenes,” in Proceedings of IEEE/CVF International Conference on Computer Vision . IEEE Computer Society, 2013, pp. 569–576. [Online]. Available: https://doi.org/10.1109/ICCV.2013.76
2013
Earlier work this paper cites.
C. Shi, C. Wang, B. Xiao, S. Gao, and J. Hu, “End-to-end scene text recognition using tree-structured models,” Pattern Recognition , vol. 47, pp. 2853–2866, 2014. [Online]. Available: https://api.semanticscholar.org/CorpusID:30201169
2014
Earlier work this paper cites.
A. Risnumawan, P. Shivakumara, C. S. Chan, and C. L. Tan, “A robust arbitrary text detection system for natural scene images,” Expert Systems with Applications , vol. 41, pp. 8027–8048, 2014. [Online]. Available: https://api.semanticscholar.org/CorpusID:15559857
2014
Earlier work this paper cites.
M. Diem, S. Fiel, F. Kleber, R. Sablatnig, J. M. Saavedra, D. Contreras, J. M. Barrios, and L. S. Oliveira, “Proceedings of ieee international conference on frontiers in handwriting recognition,” in 2014 14th International Conference on Frontiers in Handwriting Recognition , 2014, pp. 779–784
2014
Earlier work this paper cites.
D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu et al. , “Icdar 2015 competition on robust reading,” in Proceedings of International Conference on Document Analysis and Recognition . IEEE, 2015, pp. 1156–1160
2015
Earlier work this paper cites.
P. Pasupat and P. Liang, “Compositional semantic parsing on semi-structured tables,” in Proceedings of Annual Meeting of the Association for Computational Linguistics , C. Zong and M. Strube, Eds. Beijing, China: Association for Computational Linguistics, Jul. 2015, pp. 1470–1480. [Online]. Available: https://aclanthology.org/P15-1142
2015
Earlier work this paper cites.
A. W. Harley, A. Ufkes, and K. G. Derpanis, “Evaluation of deep convolutional nets for document image classification and retrieval,” in Proceedings of International Conference on Document Analysis and Recognition , 2015
2015
Earlier work this paper cites.
B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 39, no. 11, pp. 2298–2304, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi, “A diagram is worth a dozen images,” in Proceedings of European Conference on Computer Vision . Springer, 2016, pp. 235–251
2016
Earlier work this paper cites.
C. K. Ch’ng and C. S. Chan, “Total-text: A comprehensive dataset for scene text detection and recognition,” in Proceedings of International Conference on Document Analysis and Recognition , vol. 1. IEEE, 2017, pp. 935–942
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
B. Shi, C. Yao, M. Liao, M. Yang, P. Xu, L. Cui, S. Belongie, S. Lu, and X. Bai, “Icdar2017 competition on reading chinese text in the wild (rctw-17),” in Proceedings of International Conference on Document Analysis and Recognition , vol. 1. IEEE, 2017, pp. 1429–1434
2017
Earlier work this paper cites.
B. Deka, Z. Huang, C. Franzen, J. Hibschman, D. Afergan, Y. Li, J. Nichols, and R. Kumar, “Rico: A mobile app dataset for building data-driven design applications,” in Proceedings of the 30th annual ACM symposium on user interface software and technology , 2017, pp. 845–854
2017
Earlier work this paper cites.
A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi, “Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2017, pp. 4999–5007
2017
Earlier work this paper cites.
B. Shi, M. Yang, X. Wang, P. Lyu, C. Yao, and X. Bai, “Aster: An attentional scene text recognizer with flexible rectification,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 41, no. 9, pp. 2035–2048, 2018
2018
Earlier work this paper cites.
K. Kafle, S. Cohen, B. Price, and C. Kanan, “Dvqa: Understanding data visualizations via question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2018
2018
Earlier work this paper cites.
M. He, Y. Liu, Z. Yang, S. Zhang, C. Luo, F. Gao, Q. Zheng, Y. Wang, X. Zhang, and L. Jin, “Icpr2018 contest on robust reading for multi-type web images,” in Proceedings of the International Conference on Pattern Recognition , 2018, pp. 7–12
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 8317–8326
2019
Earlier work this paper cites.
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas, “Scene text visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 4291–4301
2019
Earlier work this paper cites.
Y. Liu, L. Jin, S. Zhang, C. Luo, and S. Zhang, “Curved scene text detection via transverse and longitudinal sequence connection,” Pattern Recognition , vol. 90, no. C, p. 337–345, Jun. 2019. [Online]. Available: https://doi.org/10.1016/j.patcog.2019.02.002
2019
Earlier work this paper cites.
G. Jaume, H. K. Ekenel, and J.-P. Thiran, “Funsd: A dataset for form understanding in noisy scanned documents,” in Proceedings of International Conference on Document Analysis and Recognition Workshops , vol. 2. IEEE, 2019, pp. 1–6
2019
Earlier work this paper cites.
Z. Huang, K. Chen, J. He, X. Bai, D. Karatzas, S. Lu, and C. Jawahar, “Icdar2019 competition on scanned receipt ocr and information extraction,” in Proceedings of International Conference on Document Analysis and Recognition . IEEE, 2019, pp. 1516–1520
2019
Earlier work this paper cites.
D. Saxton, E. Grefenstette, F. Hill, and P. Kohli, “Analysing mathematical reasoning abilities of neural models,” in Proceedings of the International Conference on Learning Representations . OpenReview.net, 2019
2019
Earlier work this paper cites.
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty, “Ocr-vqa: Visual question answering by reading text in images,” in Proceedings of International Conference on Document Analysis and Recognition . IEEE, 2019, pp. 947–952
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Park, S. Shin, B. Lee, J. Lee, J. Surh, M. Seo, and H. Lee, “Cord: a consolidated receipt dataset for post-ocr parsing,” in Advances in Neural Information Processing Systems Workshop , 2019
2019
Earlier work this paper cites.
X. Zhong, J. Tang, and A. Jimeno-Yepes, “Publaynet: Largest dataset ever for document layout analysis,” in Proceedings of International Conference on Document Analysis and Recognition . IEEE, 2019, pp. 1015–1022
2019
Earlier work this paper cites.
J. Tang, Z. Yang, Y. Wang, Q. Zheng, Y. Xu, and X. Bai, “Seglink++: Detecting dense and arbitrary-shaped scene text by instance-aware component grouping,” Pattern Recognition , vol. 96, p. 106954, 2019. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0031320319302511
2019
Earlier work this paper cites.
C. K. Chng, Y. Liu, Y. Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding et al. , “Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art,” in Proceedings of International Conference on Document Analysis and Recognition . IEEE, 2019, pp. 1571–1576
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in Neural Information Processing Systems , 2020
2020
Earlier work this paper cites.
X. Wang, Y. Liu, C. Shen, C. C. Ng, C. Luo, L. Jin, C. S. Chan, A. v. d. Hengel, and L. Wang, “On the general value of evidence, and bilingual scene-text visual question answering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 126–10 135
2020
Earlier work this paper cites.
X. Zhong, E. ShafieiBavani, and A. Jimeno Yepes, “Image-based table recognition: data, model, and evaluation,” in Proceedings of European Conference on Computer Vision . Springer, 2020, pp. 564–580
2020
Earlier work this paper cites.
Y. Liu, H. Chen, C. Shen, T. He, L. Jin, and L. Wang, “Abcnet: Real-time scene text spotting with adaptive bezier-curve network,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 9809–9818
2020
Earlier work this paper cites.
N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar, “Plotqa: Reasoning over scientific plots,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision , March 2020
2020
Earlier work this paper cites.
M. Mathew, D. Karatzas, and C. Jawahar, “Docvqa: A dataset for vqa on document images,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision , 2021, pp. 2200–2209
2021
Earlier work this paper cites.
S. Fang, H. Xie, Y. Wang, Z. Mao, and Y. Zhang, “Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 7098–7107
2021
Earlier work this paper cites.
N. Lu, W. Yu, X. Qi, Y. Chen, P. Gong, R. Xiao, and X. Bai, “Master: Multi-aspect non-local network for scene text recognition,” Pattern Recognition , vol. 117, p. 107980, 2021
2021
Earlier work this paper cites.
Y. Liu, C. Shen, L. Jin, T. He, P. Chen, C. Liu, and H. Chen, “Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 11, pp. 8048–8064, 2021
2021
Earlier work this paper cites.
Y. Wang, H. Xie, S. Fang, J. Wang, S. Zhu, and Y. Zhang, “From two to one: A new scene text recognizer with visual language modeling network,” in Proceedings of IEEE/CVF International Conference on Computer Vision , 2021, pp. 14 194–14 203
2021
Earlier work this paper cites.
L. Rujiao, W. Wen, X. Nan, G. Feiyu, Y. Zhibo, W. Yongpan, and X. Gui-Song, “Parsing table structures in the wild,” in Proceedings of IEEE/CVF International Conference on Computer Vision , October 2021
2021
Earlier work this paper cites.
M. Mathew, D. Karatzas, and C. V. Jawahar, “Docvqa: A dataset for VQA on document images,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision . IEEE, 2021, pp. 2199–2208
2021
Earlier work this paper cites.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
X. Zheng, D. Burdick, L. Popa, X. Zhong, and N. X. R. Wang, “Global table extractor (GTE): A framework for joint table identification and cell structure recognition using visual context,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision . IEEE, 2021, pp. 697–706. [Online]. Available: https://doi.org/10.1109/WACV48630.2021.00074
2021
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
S. Zhang, B. Yang, Z. Li, Z. Ma, Y. Liu, and X. Bai, “Exploring the Capabilities of Large Multimodal Models on Dense Text,” in Proceedings of International Conference on Document Analysis and Recognition . Springer, 2024, pp. 281–298
2024
Closest in time.
J. Chen, L. Kong, H. Wei, C. Liu, Z. Ge, L. Zhao, J. Sun, C. Han, and X. Zhang, “Onechart: Purify the chart structural extraction via one auxiliary token,” in Proceedings of the ACM International Conference on Multimedia , 2024, pp. 147–155
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar, “Infographicvqa,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision , 2022, pp. 1697–1706
2022
Cited alongside, same era.
Y. Du, Z. Chen, C. Jia, X. Yin, T. Zheng, C. Li, Y. Du, and Y. Jiang, “SVTR: scene text recognition with a single visual model,” in Proceedings of the International Joint Conference on Artificial Intelligence , L. D. Raedt, Ed. ijcai.org, 2022, pp. 884–890. [Online]. Available: https://doi.org/10.24963/ijcai.2022/124
2022
Cited alongside, same era.
X. Zhang, Y. Su, S. Tripathi, and Z. Tu, “Text spotting transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 9519–9528
2022
Cited alongside, same era.
X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, “Toward understanding wordart: Corner-guided transformer for scene text recognition,” in Proceedings of European Conference on Computer Vision , S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner, Eds. Cham: Springer Nature Switzerland, 2022, pp. 303–321
2022
Cited alongside, same era.
S. Long, S. Qin, D. Panteleev, A. Bissacco, Y. Fujii, and M. Raptis, “Towards end-to-end unified scene text detection and layout analysis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 1049–1059
2022
Cited alongside, same era.
Y. Xu, T. Lv, L. Cui, G. Wang, Y. Lu, D. Florencio, C. Zhang, and F. Wei, “XFUND: A benchmark dataset for multilingual visually rich form understanding,” in Proceedings of Annual Meeting of the Association for Computational Linguistics . Dublin, Ireland: Association for Computational Linguistics, May 2022, pp. 3214–3224. [Online]. Available: https://aclanthology.org/2022.findings-acl.253
2022
Cited alongside, same era.
2022
Cited alongside, same era.
L. Ding, M. Zhao, F. Yin, S. Zeng, and C.-L. Liu, “A large-scale database for chemical structure recognition and preliminary evaluation,” in Proceedings of the International Conference on Pattern Recognition , 2022, pp. 1464–1470
2022
Cited alongside, same era.
Closest in time.
2024
Closest in time.
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee, “Llava-next: Improved reasoning, ocr, and world knowledge,” 2024
2024
Closest in time.
2024
Closest in time.
Z. Li, B. Yang, Q. Liu, Z. Ma, S. Zhang, J. Yang, Y. Sun, Y. Liu, and X. Bai, “Monkey: Image resolution and text label are important things for large multi-modal models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 763–26 773
2024
Closest in time.
2024
Closest in time.
S. Tong, E. L. Brown II, P. Wu, S. Woo, A. J. IYER, S. C. Akula, S. Yang, J. Yang, M. Middepogu, Z. Wang et al. , “Cambrian-1: A fully open, vision-centric exploration of multimodal llms,” in Advances in Neural Information Processing Systems , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
OpenAI, “GPT-4o mini: advancing cost-efficient intelligence,” https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence , 2024, accessed: 2024-12-29
2024
Closest in time.
Anthropic, “Claude 3.5 Sonnet,” https://www.anthropic.com/news/claude-3-5-sonnet , 2024, accessed: 2024-12-29
2024
Closest in time.
StepFun, “Step-1V,” https://www.stepfun.com/##step1v , 2024, accessed: 2024-12-29
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Liang, Y. Zhang, C. Ma, Z. Zhang, Y. Zhao, L. Xiang, C. Zong, and Y. Zhou, “Document image machine translation with dynamic multi-pre-trained models assembling,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2024, pp. 7077–7088
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Zhang, D. Jiang, Y. Zhang, H. Lin, Z. Guo, P. Qiu, A. Zhou, P. Lu, K. Chang, Y. Qiao, P. Gao, and H. Li, “MATHVERSE: does your multi-modal LLM truly see the diagrams in visual math problems?” in Proceedings of European Conference on Computer Vision , ser. Lecture Notes in Computer Science, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., vol. 15066. Springer, 2024, pp. 169–186
2024
Closest in time.
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K. Chang, M. Galley, and J. Gao, “Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,” in Proceedings of the International Conference on Learning Representations . OpenReview.net, 2024
2024
Closest in time.
G. Baechler, S. Sunkara, M. Wang, F. Zubach, H. Mansoor, V. Etter, V. Cărbune, J. Lin, J. Chen, and A. Sharma, “Screenai: A vision-language model for ui and infographics understanding,” 2024
2024
Closest in time.
S. Lee, B. Lai, F. Ryan, B. Boote, and J. M. Rehg, “Modeling multimodal social interactions: New challenges and baselines with densely aligned representations,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 585–14 595
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Q. Sun, Y. Cui, X. Zhang, F. Zhang, Q. Yu, Y. Wang, Y. Rao, J. Liu, T. Huang, and X. Wang, “Generative multimodal models are in-context learners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 398–14 409
2024
Closest in time.
2024
Closest in time.
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu et al. , “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 24 185–24 198
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Laurençon, A. Marafioti, V. Sanh, and L. Tronchon, “Building and better understanding vision-language models: insights and future directions,” in Workshop on Responsibly Building the Next Generation of Multimodal Foundational Models , 2024
2024
Closest in time.
H. Duan, J. Yang, Y. Qiao, X. Fang, L. Chen, Y. Liu, X. Dong, Y. Zang, P. Zhang, J. Wang et al. , “Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,” in Proceedings of the ACM International Conference on Multimedia , 2024, pp. 11 198–11 201
2024
Closest in time.
Q. Team, “Qwen2.5-vl,” January 2025. [Online]. Available: https://qwenlm.github.io/blog/qwen2.5-vl/
2025
Closest in time.
J. Zhang, W. Yang, S. Lai, Z. Xie, and L. Jin, “Dockylin: A large multimodal model for visual document understanding with efficient visual slimming,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 9, 2025, pp. 9923–9932
2025
Closest in time.
Z. Zhu, C. Luo, Z. Shao, F. Gao, H. Xing, Q. Zheng, and J. Zhang, “A simple yet effective layout token in large language models for document understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2025
2025
Closest in time.
H. Xiao, Y. Xie, G. Tan, Y. Chen, R. Hu, K. Wang, A. Zhou, H. Li, H. Shao, X. Lu, P. Gao, Y. Wen, X. Chen, S. Ren, and H. Li, “Adaptive markup language generation for contextually-grounded visual document understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2025
2025
Closest in time.
Z. Wang, T. Guan, P. Fu, C. Duan, Q. Jiang, Z. Guo, S. Guo, J. Luo, W. Shen, and X. Yang, “Marten: Visual question answering with mask generation for multi-modal document understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2025
2025
Closest in time.
W. Wang, S. Zhang, Y. Ren, Y. Duan, T. Li, S. Liu, M. Hu, Z. Chen, K. Zhang, L. Lu et al. , “Needle in a multimodal haystack,” Advances in Neural Information Processing Systems , vol. 37, pp. 20 540–20 565, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
W. Yang, Z. Li, D. Peng, L. Jin, M. He, and C. Yao, “Read ten lines at one glance: Line-aware semi-autoregressive transformer for multi-line handwritten mathematical expression recognition,” in Proceedings of the ACM International Conference on Multimedia , A. El-Saddik, T. Mei, R. Cucchiara, M. Bertini, D. P. T. Vallejo, P. K. Atrey, and M. S. Hossain, Eds. ACM, 2023, pp. 2066–2077
2077
Closest in time.