Fetching the paper…
Reading the bibliography…
The open-ended question answering task of Text-VQA often requires reading and reasoning about rarely seen or completely unseen scene-text content of an image.
J. Bigham, C. Jayant, H. Ji, G. Little, A. Miller, R. Miller, R. Miller, A. Tatarowicz, B. White, S. White, and T. Yeh, “Vizwiz: Nearly real-time answers to visual questions,” 01 2010, pp. 333–342
2010
Earlier work this paper cites.
R. Speer, C. Havasi
2012
Earlier work this paper cites.
D. Karatzas, F. Shafait, S. Uchida, M. Iwamura, L. G. i Bigorda, S. R. Mestre, J. Mas, D. F. Mota, J. A. Almazan, and L. P. De Las Heras, “Icdar 2013 robust reading competition,” in
2013
Earlier work this paper cites.
A. Mishra, A. Karteek, and C. V. Jawahar, “Image retrieval using textual cues,”
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
D. Vrandečić and M. Krötzsch, “Wikidata: a free collaborative knowledgebase,”
2014
Earlier work this paper cites.
J. Almazán, A. Gordo, A. Fornés, and E. Valveny, “Word spotting and recognition with embedded attributes,”
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in
2015
Earlier work this paper cites.
M. Ren, R. Kiros, and R. Zemel, “Image question answering: A visual semantic embedding model and a new dataset,”
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” in
2015
Earlier work this paper cites.
D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu, F. Shafait, S. Uchida, and E. Valveny, “Icdar 2015 competition on robust reading,” in
2015
Earlier work this paper cites.
Q. Wu, P. Wang, C. Shen, A. Dick, and A. Van Den Hengel, “Ask me anything: Free-form visual question answering based on knowledge from external sources,” in
2016
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. Shamma, M. S. Bernstein, and L. Fei-Fei, “Visual genome: Connecting language and vision using crowdsourced dense image annotations,”
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Karaoglu, R. Tao, T. Gevers, and A. W. Smeulders, “Words matter: Scene text for image classification and retrieval,”
2017
Earlier work this paper cites.
Z. Hussain, M. Zhang, X. Zhang, K. Ye, C. Thomas, Z. Agha, N. Ong, and A. Kovashka, “Automatic understanding of image and video advertisements,”
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching word vectors with subword information,”
2017
Earlier work this paper cites.
B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,”
2017
Cited alongside, same era.
K. Kafle and C. Kanan, “Visual question answering: Datasets, algorithms, and future challenges,”
2017
Cited alongside, same era.
2017
Cited alongside, same era.
S. Bai and S. An, “A survey on automatic image caption generation,”
2018
Cited alongside, same era.
M. Narasimhan and A. G. Schwing, “Straight to the facts: Learning knowledge base retrieval for factual visual question answering,” in
2018
Cited alongside, same era.
M. Liao, P. Lyu, M. He, C. Yao, W. Wu, and X. Bai, “Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes,”
2019
Later among the works it cites.
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh, “Textcaps: a dataset for image captioning with reading comprehension,” in
2020
Later among the works it cites.
R. Hu, A. Singh, T. Darrell, and M. Rohrbach, “Iterative answer prediction with pointer-augmented multimodal transformers for textvqa,”
2020
Later among the works it cites.
Y. Kant, D. Batra, P. Anderson, A. Schwing, D. Parikh, J. Lu, and H. Agrawal, “Spatially aware multimodal transformers for textvqa,” in
2020
Later among the works it cites.
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Wang, Q. Wu, C. Shen, A. Dick, and A. van den Hengel, “Fvqa: Fact-based visual question answering,”
2018
Cited alongside, same era.
M. Narasimhan, S. Lazebnik, and A. G. Schwing, “Out of the box: Reasoning with graph convolution nets for factual visual question answering,” in
2018
Cited alongside, same era.
D. Deng, H. Liu, X. Li, and D. Cai, “Pixellink: Detecting scene text via instance segmentation,”
2018
Cited alongside, same era.
A. Singh, V. Natarajan, Y. Jiang, X. Chen, M. Shah, M. Rohrbach, D. Batra, and D. Parikh, “Pythia-a platform for vision & language research,” in
2018
Cited alongside, same era.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in
2019
Cited alongside, same era.
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas, “Scene text visual question answering,” in
2019
Cited alongside, same era.
S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar, “Kvqa: Knowledge-aware visual question answering,” in
2019
Cited alongside, same era.
Later among the works it cites.
G. Li, X. Wang, and W. Zhu, “Boosting visual question answering with context-aware knowledge aggregation,”
2020
Later among the works it cites.
N. García, M. Otani, C. Chu, and Y. Nakashima, “Knowit vqa: Answering knowledge-based questions about videos,” in
2020
Later among the works it cites.
J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee, “12-in-1: Multi-task vision and language representation learning,” in
2020
Later among the works it cites.
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, and et al., “The open images dataset v4,”
2020
Later among the works it cites.
F. Liu, G. Xu, Q. Wu, Q. Du, W. Jia, and M. Tan, “Cascade reasoning network for text-based visual question answering,” in
2020
Later among the works it cites.
W. Han, H. Huang, and T. Han, “Finding the evidence: Localization-aware answer prediction for text visual question answering,” in
2020
Later among the works it cites.
A. U. Dey, S. K. Ghosh, E. Valveny, and G. Harit, “Beyond visual semantics: Exploring the role of scene text in image understanding,”
2021
Closest in time.
X. Chen, L. Jin, Y. Zhu, C. Luo, and T. Wang, “Text recognition in the wild: A survey,”
2021
Closest in time.
K. Ye, M. Zhang, and A. Kovashka, “Breaking shortcuts by masking for robust visual reasoning,” in
2021
Closest in time.
Z. Yang, Y. Lu, J. Wang, X. Yin, D. Florencio, L. Wang, C. Zhang, L. Zhang, and J. Luo, “Tap: Text-aware pre-training for text-vqa and text-caption,” in
2021
Closest in time.
2021
Closest in time.
J. Wu, J. Lu, A. Sabharwal, and R. Mottaghi, “Multi-modal answer validation for knowledge-based vqa,” in
2022
Closest in time.