Fetching the paper…
Reading the bibliography…
TextVQA requires models to read and reason about text in images to answer questions about them.
Synthetic data for text localisation in natural images
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman · 2016
Earlier work this paper cites.
Tap: Text-aware pre-training for text-vqa and text-caption, 2020
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florencio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo · 2020
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models, 2021
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…