Fetching the paper…
Reading the bibliography…
Methodologies for training visual question answering (VQA) models assume the availability of datasets with human-annotated \textit{Image-Question-Answer} (I-Q-A) triplets.
Trends in integration of vision and language research: A survey of tasks, datasets, and methods
Aditya Mogadala, Marimuthu Kalimuthu, and Dietrich Klakow. 2019 · 1907
Earlier work this paper cites.
Uniter: Learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Actively seeking and learning from live data
Damien Teney and Anton van den Hengel. 2019 · 1949
Earlier work this paper cites.
An analysis of visual question answering algorithms
Kushal Kafle and Christopher Kanan. 2017 · 1991
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick. 2017 · 1997
Earlier work this paper cites.
Unshuffling data for improved generalization
Damien Teney, Ehsan Abbasnejad, and Anton van den Hengel. 2020b · 2002
Earlier work this paper cites.
Learning what makes a difference from counterfactual examples and gradient supervision
Damien Teney, Ehsan Abbasnedjad, and Anton van den Hengel. 2020a · 2004
Earlier work this paper cites.
On the value of out-of-distribution testing: An example of goodhart’s law
Damien Teney, Kushal Kafle, Robik Shrestha, Ehsan Abbasnejad, Christopher Kanan, and Anton van den Hengel. 2020c · 2005
Earlier work this paper cites.
Weak supervision and referring attention for temporal-textual association learning
Zhiyuan Fang, Shu Kong, Zhe Wang, Charless Fowlkes, and Yezhou Yang. 2020b · 2006
Earlier work this paper cites.
Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories
Svetlana Lazebnik, Cordelia Schmid, and Jean Ponce. 2006 · 2006
Earlier work this paper cites.
When training and test sets are different: characterizing learning transfer
Amos Storkey. 2009 · 2009
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. 2011 · 2011
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross B. Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Towards a visual turing challenge
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
On learning to localize objects with minimal supervision
Hyun Oh Song, Ross B. Girshick, Stefanie Jegelka, Julien Mairal, Zaïd Harchaoui, and Trevor Darrell. 2014 · 2014
Earlier work this paper cites.
VQA: visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Question-answer driven semantic role labeling: Using natural language to annotate natural language
Luheng He, Mike Lewis, and Luke Zettlemoyer. 2015 · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard S. Zemel. 2015a · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015b · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Zero-shot visual question answering
Damien Teney and Anton van den Hengel. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alexander J. Smola. 2016 · 2016
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016 · 2016
Cited alongside, same era.
Adversarial regularization for visual question answering: Strengths, shortcomings, and side effects
Gabriel Grand and Yonatan Belinkov. 2019 · 2019
Later among the works it cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Clevr-ref+: Diagnosing visual reasoning with referring expressions
Runtao Liu, Chenxi Liu, Yutong Bai, and Alan L. Yuille. 2019 · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Weakly supervised video moment retrieval from text queries
Niluthpol Chowdhury Mithun, Sujoy Paul, and Amit K. Roy-Chowdhury. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Àgata Lapedriza, Aude Oliva, and Antonio Torralba. 2016 · 2016
Cited alongside, same era.
C-vqa: A compositional split of the visual question answering (vqa) v1.0 dataset
Aishwarya Agrawal, Aniruddha Kembhavi, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Learning to ask: Neural question generation for reading comprehension
Xinya Du, Junru Shao, and Claire Cardie. 2017 · 2017
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell. 2017 · 2017
Cited alongside, same era.
Simple does it: Weakly supervised instance and semantic segmentation
Anna Khoreva, Rodrigo Benenson, Jan Hendrik Hosang, Matthias Hein, and Bernt Schiele. 2017 · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017 · 2017
Cited alongside, same era.
Timothy Niven and Hung-Yu Kao. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Later among the works it cites.
Sunny and dark outside?! improving answer consistency in VQA through entailed question generation
Arijit Ray, Karan Sikka, Ajay Divakaran, Stefan Lee, and Giedrius Burachas. 2019 · 2019
Later among the works it cites.
Are red roses red? evaluating consistency of question-answering models
Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019 · 2019
Later among the works it cites.
Answer them all! toward universal visual question answering models
Robik Shrestha, Kushal Kafle, and Christopher Kanan. 2019 · 2019
Later among the works it cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Meta-learning to detect rare objects
Yu-Xiong Wang, Deva Ramanan, and Martial Hebert. 2019 · 2019
Later among the works it cites.
Self-critical reasoning for robust visual question answering
Jialin Wu and Raymond J. Mooney. 2019 · 2019
Later among the works it cites.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019 · 2019
Later among the works it cites.
Self-supervised knowledge triplet learning for zero-shot question answering
Pratyay Banerjee and Chitta Baral. 2020 · 2020
Closest in time.
Counterfactual samples synthesizing for robust visual question answering
Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang. 2020 · 2020
Closest in time.
Template-based question generation from retrieved sentences for improved unsupervised question answering
Alexander Fabbri, Patrick Ng, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2020 · 2020
Closest in time.
Video2Commonsense: Generating commonsense descriptions to enrich video captioning
Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. 2020a · 2020
Closest in time.
MUTANT: A training paradigm for out-of-distribution generalization in visual question answering
Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. 2020a · 2020
Closest in time.
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik G. Learned-Miller, and Xinlei Chen. 2020 · 2020
Closest in time.
Learning the difference that makes A difference with counterfactually-augmented data
Divyansh Kaushik, Eduard H. Hovy, and Zachary Chase Lipton. 2020 · 2020
Closest in time.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Closest in time.
Harvesting and refining question-answer pairs for unsupervised QA
Zhongli Li, Wenhui Wang, Li Dong, Furu Wei, and Ke Xu. 2020 · 2020
Closest in time.
Unsupervised adaptation of question answering systems via generative self-training
Steven Rennie, Etienne Marcheret, Neil Mallinar, David Nahamoo, and Vaibhava Goel. 2020 · 2020
Closest in time.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Closest in time.
Squinting at VQA models: Introspecting VQA models with sub-questions
Ramprasaath R. Selvaraju, Purva Tendulkar, Devi Parikh, Eric Horvitz, Marco Túlio Ribeiro, Besmira Nushi, and Ece Kamar. 2020 · 2020
Closest in time.
Open-ended visual question answering by multi-modal domain adaptation
Yiming Xu, Lin Chen, Zhongwei Cheng, Lixin Duan, and Jiebo Luo. 2020 · 2020
Closest in time.
Self-supervised test-time learning for reading comprehension
Pratyay Banerjee, Tejas Gokhale, and Chitta Baral. 2021 · 2021
Closest in time.
Large-scale QA-SRL parsing
Nicholas FitzGerald, Julian Michael, Luheng He, and Luke Zettlemoyer. 2018 · 2060
Closest in time.