Fetching the paper…
Reading the bibliography…
Text-Centric Visual Question Answering (TEC-VQA) in its proper format not only facilitates human-machine interaction in text-centric visual environments but also serves as a de facto gold proxy to evaluate AI models in the domain of text-centric scene understanding.
Visual question answering dataset for bilingual image understanding: A study of cross-lingual transfer using attention maps
Nobuyuki Shimizu, Na Rong, and Takashi Miyazaki. 2018 · 1928
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. 2015 · 2015
Earlier work this paper cites.
Scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas. 2019 · 2019
Earlier work this paper cites.
Icdar 2019 competition on table detection and recognition (ctdar)
Liangcai Gao, Yilun Huang, Herve Dejean, Jean-Luc Meunier, Qinqin Yan, Yu Fang, Florian Kleber, and Eva Lang. 2019 · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. 2019 · 2019
Earlier work this paper cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019 · 2019
Earlier work this paper cites.
A unified framework for multilingual and code-mixed visual question answering
Deepak Gupta, Pabitra Lenka, Asif Ekbal, and Pushpak Bhattacharyya. 2020 · 2020
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. 2021 · 2021
Earlier work this paper cites.
Towards developing a multilingual and code-mixed visual question answering system by knowledge distillation
Humair Raj Khan, Deepak Gupta, and Asif Ekbal. 2021 · 2021
Earlier work this paper cites.
Infographicvqa
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. 2022 · 2022
Earlier work this paper cites.
xGQA: Cross-lingual visual question answering
Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin Steitz, Stefan Roth, Ivan Vulić, and Iryna Gurevych. 2022 · 2022
Earlier work this paper cites.
DuReader vis \textrm{DuReader}_{\textrm{vis}} : A Chinese dataset for open-domain document visual question answering
Le Qi, Shangwen Lv, Hongyu Li, Jing Liu, Yu Zhang, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ting Liu. 2022 · 2022
Earlier work this paper cites.
Must-vqa: multilingual scene-text vqa
Emanuele Vivoli, Ali Furkan Biten, Andres Mafla, Dimosthenis Karatzas, and Lluis Gomez. 2022 · 2022
Earlier work this paper cites.
OpenAI:Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, et al. 2023 · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023 · 2023
Earlier work this paper cites.
MaXM: Towards multilingual visual question answering
Soravit Changpinyo, Linting Xue, Michal Yarom, Ashish Thapliyal, Idan Szpektor, Julien Amelot, Xi Chen, and Radu Soricut. 2023 · 2023
Earlier work this paper cites.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Zhong Muyan, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. 2023 · 2023
Earlier work this paper cites.
Large multilingual models pivot zero-shot multimodal learning across languages
Jinyi Hu, Yuan Yao, Chongyi Wang, Shan Wang, Yinxu Pan, Qianyu Chen, Tianyu Yu, Hanghao Wu, Yue Zhao, Haoye Zhang, Xu Han, Yankai Lin, Jiao Xue, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Cited alongside, same era.
An empirical study of multilingual scene-text visual question answering
Lin Li, Haohan Zhang, and Zeqin Fang. 2023 · 2023
Cited alongside, same era.
Spts v2: single-point scene text spotting
Yuliang Liu, Jiaxin Zhang, Dezhi Peng, Mingxin Huang, Xinyu Wang, Jingqun Tang, Can Huang, Dahua Lin, Chunhua Shen, Xiang Bai, et al. 2023 · 2023
Cited alongside, same era.
Vlsp2022-evjvqa challenge: Multilingual visual question answering
Ngan Luu-Thuy Nguyen, Nghia Hieu Nguyen, Duong TD Vo, Khanh Quoc Tran, and Kiet Van Nguyen. 2023 · 2023
Cited alongside, same era.
Character recognition competition for street view shop signs
Jingqun Tang, Weidong Du, Bin Wang, Wenyang Zhou, Shuqi Mei, Tao Xue, Xing Xu, and Hai Zhang. 2023 · 2023
Learn from global correlations: Enhancing evolutionary algorithm via spectral gnn
Kaichen Ouyang, Shengwei Fu, and Zong Ke. 2024 · 2024
Closest in time.
Mctbench: Multimodal cognition towards text-rich visual scenes benchmark
Bin Shan, Xiang Fei, Wei Shi, An-Lan Wang, Guozhi Tang, Lei Liao, Jingqun Tang, Xiang Bai, and Can Huang. 2024 · 2024
Closest in time.
Imagpose: A unified conditional framework for pose-guided person generation
Fei Shen and Jinhui Tang. 2024 · 2024
Closest in time.
Textsquare: Scaling up text-centric visual instruction tuning
Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, et al. 2024 · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023 · 2023
Cited alongside, same era.
mPLUG-DocOwl: Modularized multimodal large language model for document understanding
Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Yuhao Dan, Chenlin Zhao, Guohai Xu, Chenliang Li, Junfeng Tian, et al. 2023 · 2023
Cited alongside, same era.
Llavar: Enhanced visual instruction tuning for text-rich image understanding
Yanzhe Zhang, Ruiyi Zhang, Jiuxiang Gu, Yufan Zhou, Nedim Lipka, Diyi Yang, and Tong Sun. 2023 · 2023
Cited alongside, same era.
Glm-4v main page
ZhiPu AI. 2024 · 2024
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic. 2024 · 2024
Cited alongside, same era.
Common crawl main page
Common Crawl. 2024 · 2024
Cited alongside, same era.
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, et al. 2024 · 2024
Cited alongside, same era.
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024 · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. 2024 · 2024
Closest in time.
Advancing sequential numerical prediction in autoregressive models
Xiang Fei, Jinghui Lu, Qi Sun, Hao Feng, Yanjie Wang, Wei Shi, An-Lan Wang, Jingqun Tang, and Can Huang. 2025 · 2025
Closest in time.
Dolphin: Document image parsing via heterogeneous anchor prompting
Hao Feng, Shu Wei, Xiang Fei, Wei Shi, Yingdong Han, Lei Liao, Jinghui Lu, Binghong Wu, Qi Liu, Chunhui Lin, et al. 2025 · 2025
Closest in time.
Dong Guo, Faming Wu, Feida Zhu, Fuxing Leng, Guang Shi, Haobin Chen, Haoqi Fan, Jian Wang, Jianyu Jiang, Jiawei Wang, et al. 2025 · 2025
Closest in time.
Enhancing intent understanding for ambiguous prompts through human-machine co-adaptation
Yangfan He, Jianhui Wang, Kun Li, Yijin Wang, Li Sun, Jun Yin, Miao Zhang, and Xueqian Wang. 2025 · 2025
Closest in time.
Detection of ai deepfake and fraud in online payments using gan-based models
Zong Ke, Shicheng Zhou, Yining Zhou, Chia Hong Chang, and Rong Zhang. 2025 · 2025
Closest in time.
Jinghui Lu, Haiyang Yu, Siliang Xu, Shiwei Ran, Guozhi Tang, Siqi Wang, Bin Shan, Teng Fu, Hao Feng, Jingqun Tang, et al. 2025 · 2025
Closest in time.
A generative adversarial network-based investor sentiment indicator: Superior predictability for the stock market
Shiqing Qiu, Yang Wang, Zong Ke, Qinyan Shen, Zichao Li, Rong Zhang, and Kaichen Ouyang. 2025 · 2025
Closest in time.
Qwen2.5-vl main page
Team Qwen. 2025 · 2025
Closest in time.
Imagdressing-v1: Customizable virtual dressing
Fei Shen, Xin Jiang, Xin He, Hu Ye, Cong Wang, Xiaoyu Du, Zechao Li, and Jinhui Tang. 2025 · 2025
Closest in time.
Attentive eraser: Unleashing diffusion model’s object removal potential via self-attention redirection guidance
Wenhao Sun, Xue-Mei Dong, Benlei Cui, and Jingqun Tang. 2025 · 2025
Closest in time.