Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) has the potential to make the Internet more accessible in an interactive way, allowing people who cannot see images to ask questions about them.
Describing images on the web: A survey of current practice and prospects for the future
Helen Petrie, Chandra Harrison, and Sundeep Dev · 2005
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Visual Dialog, 2016
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, and Dhruv Batra · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Caption Crawler: Enabling Reusable Alternative Text Descriptions using Reverse Image Search
Darren Guinness, Edward Cutrell, and Meredith Ringel Morris · 2018
Earlier work this paper cites.
VizWiz Grand Challenge: Answering Visual Questions from Blind People
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham · 2018
Cited alongside, same era.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Cited alongside, same era.
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Cited alongside, same era.
Person, Shoes, Tree. Is the Person Naked? What People with Vision Impairments Want in Image Descriptions
Abigale Stangl, Meredith Ringel Morris, and Danna Gurari · 2020
Cited alongside, same era.
Going Beyond One-Size-Fits-All Image Descriptions to Satisfy the Information Wants of People Who are Blind or Have Low Vision
Abigale Stangl, Nitin Verma, Kenneth Fleischmann, Meredith Ringel Morris, and Danna Gurari · 2021
Cited alongside, same era.
Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge
Pierre Dognin, Igor Melnyk, Youssef Mroueh, Inkit Padhi, Mattia Rigotti, Jarret Ross, Yair Schiff, Richard A Young, and Brian Belgodere · 2022
Later among the works it cites.
Context matters for image descriptions for accessibility: Challenges for referenceless evaluation metrics
Elisa Kreiss, Cynthia Bennett, Shayan Hooshmand, Eric Zelikman, Meredith Ringel Morris, and Christopher Potts · 2022
Later among the works it cites.
Concadia: Towards image-based text generation with a purpose
Elisa Kreiss, Fei Fang, Noah Goodman, and Christopher Potts · 2022
Later among the works it cites.
What’s in an ALT Tag? Exploring Caption Content Priorities through Collaborative Captioning
Annika Muehlbradt and Shaun K Kane · 2022
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What’s Different between Visual Question Answering for Machine “Understanding” Versus for Accessibility?
Yang Trista Cao, Kyle Seelman, Kyungjun Lee, and Hal III Daumé · 2022
Cited alongside, same era.
Closest in time.
GPT-4 Technical Report, 2023
OpenAI · 2023
Closest in time.