Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) has been primarily studied through the lens of the English language.
ROUGE: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Im2Text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T · 2011
Earlier work this paper cites.
Language models for machine translation: Original vs. translated texts
Lembersky, G., Ordan, N., and Wintner, S · 2012
Earlier work this paper cites.
On the features of translationese
Volansky, V., Ordan, N., and Wintner, S · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L., and Dollár, P · 2014
Earlier work this paper cites.
VQA: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Microsoft COCO Captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Ren, M., Kiros, R., and Zemel, R · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Multi30K: Multilingual English-German image descriptions
Elliott, D., Frank, S., Sima’an, K., and Specia, L · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Visual7W: Grounded question answering in images
Zhu, Y., Groth, O., Bernstein, M., and Li, F.-F · 2016
Earlier work this paper cites.
Findings of the second shared task on multimodal machine translation and multilingual image description
Elliott, D., Frank, S., Barrault, L., Bougares, F., and Specia, L · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
An analysis of visual question answering algorithms
Kafle, K. and Kanan, C · 2017
Earlier work this paper cites.
OpenImages: A public dataset for large-scale multi-label and multi-class image classification
Krasin, I., Duerig, T., Alldrin, N., Ferrari, V., Abu-El-Haija, S., Kuznetsova, A., Rom, H., Uijlings, J., Popov, S., Kamali, S., Malloci, M., Pont-Tuset, J., Veit, A., Belongie, S., Gomes, V., Gupta, A., Sun, C., Chechik, G., Cai, D., Feng, Z., Narayanan, D., and Murphy, K · 2017
Earlier work this paper cites.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., Bernstein, M., and Fei-Fei, L · 2017
Earlier work this paper cites.
STAIR captions: Constructing a large-scale Japanese image caption dataset
Yoshikawa, Y., Shigeto, Y., and Takeuchi, A · 2017
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Agrawal, A., Batra, D., Parikh, D., and Kembhavi, A · 2018
Earlier work this paper cites.
Findings of the third shared task on multimodal machine translation
Barrault, L., Bougares, F., Specia, L., Lala, C., Elliott, D., and Frank, S · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Earlier work this paper cites.
VizWiz Grand Challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for SQuAD
Rajpurkar, P., Jia, R., and Liang, P · 2018
Cited alongside, same era.
Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Cited alongside, same era.
Synthetic QA corpora generation with roundtrip consistency
Alberti, C., Andor, D., Pitler, E., Devlin, J., and Collins, M · 2019
Cited alongside, same era.
Scene text visual question answering
Biten, A. F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D · 2019
RedCaps: Web-curated image-text data created by the people, for the people
Desai, K., Kaul, G., Aysola, Z., and Johnson, J · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Q 2 Q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Honovich, O., Choshen, L., Aharoni, R., Neeman, E., Szpektor, I., and Abend, O · 2021
Later among the works it cites.
QACE: Asking questions to evaluate an image caption
Lee, H., Scialom, T., Yoon, S., Dernoncourt, F., and Jung, K · 2021
Later among the works it cites.
Adversarial VQA: A new benchmark for evaluating the robustness of vqa models
Li, L., Lei, J., Gan, Z., and Liu, J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Cited alongside, same era.
Natural Questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., , Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Cited alongside, same era.
COCO-CN for cross-lingual image tagging, captioning, and retrieval
Li, X., Xu, C., Wang, X., Lan, W., Jia, Z., Yang, G., and Xu, J · 2019
Cited alongside, same era.
OK-VQA: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Cited alongside, same era.
Towards VQA models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Cited alongside, same era.
VaTeX: A large-scale, high-quality multilingual dataset for video-and-language research
Wang, X., Wu, J., Chen, J., Li, L., Wang, Y.-F., and Wang, W. Y · 2019
Cited alongside, same era.
Visually grounded reasoning across languages and cultures
Liu, F., Bugliarello, E., Ponti, E. M., Reddy, S., Collier, N., and Elliott, D · 2021
Later among the works it cites.
LAION-400M: Open dataset of CLIP-filtered 400 million image-text pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Later among the works it cites.
Human-adversarial visual question answering
Sheng, S., Singh, A., Goswami, V., Magana, J. A. L., Galuba, W., Parikh, D., and Kiela, D · 2021
Later among the works it cites.
WIT: Wikipedia-based image text dataset for multimodal multilingual machine learning
Srinivasan, K., Raman, K., Chen, J., Bendersky, M., and Najork, M · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Closest in time.
IGLUE: A benchmark for transfer learning across modalities, tasks, and languages
Bugliarello, E., Liu, F., Pfeiffer, J., Reddy, S., Elliott, D., Ponti, E. M., and Vulić, I · 2022
Closest in time.
All you may need for VQA are image captions
Changpinyo, S., Kukliansky, D., Szpektor, I., Chen, X., Ding, N., and Soricut, R · 2022
Closest in time.
Wukong: 100 million large-scale chinese cross-modal pre-training dataset and a foundation framework
Gu, J., Meng, X., Lu, G., Hou, L., Niu, M., Xu, H., Liang, X., Zhang, W., Jiang, X., and Xu, C · 2022
Closest in time.
Quality at a glance: An audit of web-crawled multilingual datasets
Kreutzer, J., Caswell, I., Wang, L., Wahab, A., van Esch, D., Ulzii-Orshikh, N., Tapo, A., Subramani, N., Sokolov, A., Sikasote, C., Setyawan, M., Sarin, S., Samb, S., Sagot, B., Rivera, C., Rios, A., Papadimitriou, I., Osei, S., Suarez, P. O., Orife, I., Ogueji, K., Rubungo, A. N., Nguyen, T. Q., Müller, M., Müller, A., Muhammad, S. H., Muhammad, N., Mnyakeni, A., Mirzakhalov, J., Matangira, T., Leong, C., Lawson, N., Kudugunta, S., Jernite, Y., Jenny, M., Firat, O., Dossou, B. F. P., Dlamini, S., de Silva, N., Çabuk Ballı, S., Biderman, S., Battisti, A., Baruwa, A., Bapna, A., Baljekar, P., Azime, I. A., Awokoya, A., Ataman, D., Ahia, O., Ahia, O., Agrawal, S., and Adeyemi, M · 2022
Closest in time.
xGQA: Cross-lingual visual question answering
Pfeiffer, J., Geigle, G., Kamath, A., Steitz, J.-M. O., Roth, S., Vulić, I., and Gurevych, I · 2022
Closest in time.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Scao, T. L., Biderman, S., Gao, L., Wolf, T., and Rush, A. M · 2022
Closest in time.
A-OKVQA: A benchmark for visual question answering using world knowledge
Schwenk, D., Khandelwal, A., Clark, C., Marino, K., and Mottaghi, R · 2022
Closest in time.
Crossmodal-3600: A massively multilingual multimodal evaluation dataset
Thapliyal, A. V., Pont-Tuset, J., Chen, X., and Soricut, R · 2022
Closest in time.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Closest in time.