Fetching the paper…
Reading the bibliography…
Teaching Visual Question Answering (VQA) models to refrain from answering unanswerable questions is necessary for building a trustworthy AI system.
K. Kafle and C. Kanan, “An analysis of visual question answering algorithms,” in ICCV . IEEE, 2017, pp. 1983–1991
1991
Earlier work this paper cites.
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. B. Girshick, “CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning,” in CVPR . IEEE, 2017, pp. 1988–1997
1997
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in EMNLP . ACL, 2014, pp. 1532–1543
2014
Earlier work this paper cites.
M. Malinowski and M. Fritz, “A multi-world approach to question answering about real-world scenes based on uncertain input,” in NIPS , 2014, pp. 1682–1690
2014
Earlier work this paper cites.
T. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: common objects in context,” in ECCV . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: visual question answering,” in ICCV . IEEE, 2015, pp. 2425–2433
2015
Earlier work this paper cites.
M. Ren, R. Kiros, and R. S. Zemel, “Exploring models and data for image question answering,” in NIPS , 2015, pp. 2953–2961
2015
Earlier work this paper cites.
L. Yu, E. Park, A. C. Berg, and T. L. Berg, “Visual madlibs: Fill in the blank description generation and question answering,” in ICCV . IEEE, 2015, pp. 2461–2469
2015
Earlier work this paper cites.
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu, “Are you talking to a machine? dataset and methods for multilingual image question answering,” in NIPS , 2015
2015
Earlier work this paper cites.
Y. Zhu, O. Groth, M. S. Bernstein, and L. Fei-Fei, “Visual7w: Grounded question answering in images,” in CVPR . IEEE, 2016, pp. 4995–5004
2016
Earlier work this paper cites.
P. Zhang, Y. Goyal, D. Summers-Stay, D. Batra, and D. Parikh, “Yin and yang: Balancing and answering binary visual questions,” in CVPR . IEEE, 2016, pp. 5014–5022
2016
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei, “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio-temporal reasoning in visual question answering,” in CVPR . IEEE, 2017, pp. 2758–2766
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi, “Don’t just assume; look and answer: Overcoming priors for visual question answering,” in CVPR . IEEE, 2018, pp. 4971–4980
2018
Earlier work this paper cites.
P. Wang, Q. Wu, C. Shen, A. R. Dick, and A. van den Hengel, “FVQA: fact-based visual question answering,” TPAMI , vol. 40, no. 10, pp. 2413–2427, 2018
2018
Earlier work this paper cites.
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham, “Vizwiz grand challenge: Answering visual questions from blind people,” in CVPR . IEEE, 2018, pp. 3608–3617
2018
Earlier work this paper cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR . IEEE, 2018, pp. 6077–6086
2018
Earlier work this paper cites.
Q. Li, Q. Tao, S. R. Joty, J. Cai, and J. Luo, “VQA-E: explaining, elaborating, and enhancing your answers for visual questions,” in ECCV . Springer, 2018, pp. 570–586
2018
Earlier work this paper cites.
D. H. Park, L. A. Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, and M. Rohrbach, “Multimodal explanations: Justifying decisions and pointing to the evidence,” in CVPR . IEEE, 2018, pp. 8779–8788
2018
Earlier work this paper cites.
P. Rajpurkar, R. Jia, and P. Liang, “Know what you don’t know: Unanswerable questions for squad,” in ACL . ACL, 2018, pp. 784–789
2018
Earlier work this paper cites.
E. Choi, H. He, M. Iyyer, M. Yatskar, W. Yih, Y. Choi, P. Liang, and L. Zettlemoyer, “Quac: Question answering in context,” in EMNLP . ACL, 2018, pp. 2174–2184
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL . ACL, 2018, pp. 2556–2565
2018
Earlier work this paper cites.
Y. Goyal, T. Khot, A. Agrawal, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,” IJCV , vol. 127, no. 4, pp. 398–414, 2019
2019
Earlier work this paper cites.
D. A. Hudson and C. D. Manning, “GQA: A new dataset for real-world visual reasoning and compositional question answering,” in CVPR . IEEE, 2019, pp. 6700–6709
2019
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “OK-VQA: A visual question answering benchmark requiring external knowledge,” in CVPR . IEEE, 2019, pp. 3195–3204
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL . ACL, 2019, pp. 4171–4186
2019
Cited alongside, same era.
H. Tan and M. Bansal, “LXMERT: learning cross-modality encoder representations from transformers,” in EMNLP . ACL, 2019, pp. 5099–5110
2019
Cited alongside, same era.
A. F. Biten, R. Tito, A. Mafla, L. G. i Bigorda, M. Rusiñol, C. V. Jawahar, E. Valveny, and D. Karatzas, “Scene text visual question answering,” in ICCV . IEEE, 2019, pp. 4290–4300
2019
Cited alongside, same era.
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi, “From recognition to cognition: Visual commonsense reasoning,” in CVPR . IEEE, 2019, pp. 6720–6731
2019
Cited alongside, same era.
S. Reddy, D. Chen, and C. D. Manning, “Coqa: A conversational question answering challenge,” TASL , vol. 7, pp. 249–266, 2019
Z. Gao, S. Chen, Y. Guo, W. Guan, J. Nie, and A. Liu, “Generic image manipulation localization through the lens of multi-scale spatial inconsistence,” in MM . ACM, 2022, pp. 6146–6154
2022
Later among the works it cites.
Y. Guo, L. Nie, Z. Cheng, Q. Tian, and M. Zhang, “Loss re-scaling VQA: revisiting the language prior problem from a class-imbalance view,” TIP , vol. 31, pp. 227–238, 2022
2022
Later among the works it cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev, “LAION-5B: an open large-scale dataset for training next generation image-text models,” in NeurIPS , 2022
2022
Later among the works it cites.
V. Raina and M. J. F. Gales, “Answer uncertainty and unanswerability in multiple-choice machine reading comprehension,” in Findings of ACL . ACL, 2022, pp. 1020–1034
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
P. S. H. Lewis, L. Denoyer, and S. Riedel, “Unsupervised question answering by cloze translation,” in ACL . ACL, 2019, pp. 4896–4910
2019
Cited alongside, same era.
M. Hu, F. Wei, Y. Peng, Z. Huang, N. Yang, and D. Li, “Read + verify: Machine reading comprehension with unanswerable questions,” in AAAI . AAAI Press, 2019, pp. 6529–6537
2019
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in NeurIPS , 2019, pp. 13–23
2019
Cited alongside, same era.
C. Alberti, J. Ling, M. Collins, and D. Reitter, “Fusion of detected objects in text for visual question answering,” in EMNLP . ACL, 2019, pp. 2131–2140
2019
Cited alongside, same era.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in ICCV . IEEE, 2019, pp. 7463–7472
2019
Cited alongside, same era.
T. Gokhale, P. Banerjee, C. Baral, and Y. Yang, “VQA-LOL: visual question answering under the lens of logic,” in ECCV . Springer, 2020, pp. 379–396
2020
Cited alongside, same era.
S. Back, S. C. Chinthakindi, A. Kedia, H. Lee, and J. Choo, “Neurquri: Neural question requirement inspector for answerability prediction in machine reading comprehension,” in ICLR . OpenReview.net, 2020
2020
Cited alongside, same era.
E. Sulem, J. Hay, and D. Roth, “Yes, no or IDK: the challenge of unanswerable yes/no questions,” in NAACL , M. Carpuat, M. de Marneffe, and I. V. M. Ruíz, Eds. ACL, 2022, pp. 1075–1085
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Closest in time.
OpenAI, “GPT-4 technical report,” CoRR , vol. abs/2303.08774, 2023
2023
Closest in time.
Z. Yin, Q. Sun, Q. Guo, J. Wu, X. Qiu, and X. Huang, “Do large language models know what they don’t know?” in Findings of ACL . ACL, 2023, pp. 8653–8665
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” in NeurIPS , 2023
2023
Closest in time.
S. Tascon-Morales, P. Márquez-Neila, and R. Sznitman, “Logical implications for visual question answering consistency,” in CVPR . IEEE, 2023, pp. 6725–6735
2023
Closest in time.
2023
Closest in time.
M. N. Team. (2023) Introducing mpt-7b: A new standard for open-source, commercially usable llms. Accessed: 2023-05-05. [Online]. Available: www.mosaicml.com/blog/mpt-7b
2023
Closest in time.
2023
Closest in time.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
C. Li and J. Flanigan, “Task contamination: Language models may not be few-shot anymore,” in AAAI . AAAI Press, 2024, pp. 18 471–18 480
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: Enhancing vision-language understanding with advanced large language models,” in ICLR , 2024
2024
Closest in time.
L. Li, J. Lei, Z. Gan, and J. Liu, “Adversarial VQA: A new benchmark for evaluating the robustness of VQA models,” in ICCV . IEEE, 2021, pp. 2022–2031
2031
Closest in time.