Gqa: A new dataset for real-world visual reasoning and compositional question answering
Original
D. Hudson and C. Manning · 1902
Earlier work this paper cites.
Docvqa: A dataset for VQA on document images
Original
M. Mathew, D. Karatzas, R. Manmatha, and C. V. Jawahar · 2007
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
T. Kudo and J. Richardson · 2012
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Original
T. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Doll’a r, and C. L. Zitnick · 2014
Earlier work this paper cites.
A diagram is worth a dozen images
E. K. M. S. H. H. A. F. Aniruddha Kembhavi, Mike Salvato · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
Tallyqa: Answering complex counting questions
M. Acharya, K. Kafle, and C. Kanan · 2019
Earlier work this paper cites.
nocaps: novel object captioning at scale
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson · 2019
Earlier work this paper cites.
Scene text visual question answering
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, C. Jawahar, E. Valveny, and D. Karatzas · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty · 2019
Earlier work this paper cites.
Big transfer (BiT): General visual representation learning
A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby · 2020
Earlier work this paper cites.
Widget captioning: Generating natural language description for mobileuser interface elements
Y. Li, G. Li, L. He, J. Zheng, H. Li, and Z. Guan · 2020
Earlier work this paper cites.
Rsvqa: Visual question answering for remote sensing data
S. Lobry, D. Marcos, J. Murray, and D. Tuia · 2020
Earlier work this paper cites.
Unifying vision-and-language tasks via text generation
J. Cho, J. Lei, H. Tan, and M. Bansal · 2021
Earlier work this paper cites.
Virtex: Learning visual representations from textual annotations
K. Desai and J. Johnson · 2021
Earlier work this paper cites.
Initializing new word embeddings for pretrained language models
J. Hewitt · 2021
Earlier work this paper cites.
Scicap: Generating captions for scientific figures
Original
T.-Y. Hsu, C. L. Giles, and T.-H. Huang · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig · 2021
Earlier work this paper cites.