Fetching the paper…
Reading the bibliography…
Despite Visual Question Answering (VQA) has realized impressive progress over the last few years, today's VQA models tend to capture superficial linguistic correlations in the train set and fail to generalize to the test set with different QA distributions.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens Van Der Maaten · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Earlier work this paper cites.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning
Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Attention correctness in neural image captioning
Chenxi Liu, Junhua Mao, Fei Sha, and Alan Yuille · 2017
Earlier work this paper cites.
Right for the right reasons: Training differentiable models by constraining their explanations
Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
Video question answering via attribute-augmented attention network learning
Yunan Ye, Zhou Zhao, Yimeng Li, Long Chen, Jun Xiao, and Yueting Zhuang · 2017
Cited alongside, same era.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Zero-shot visual recognition using semantics-preserving adversarial embedding networks
Long Chen, Hanwang Zhang, Jun Xiao, Wei Liu, and Shih-Fu Chang · 2018
Cited alongside, same era.
Counterfactual critic multi-agent training for scene graph generation
Long Chen, Hanwang Zhang, Jun Xiao, Xiangnan He, Shiliang Pu, and Shih-Fu Chang · 2019
Later among the works it cites.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer · 2019
Later among the works it cites.
Adversarial regularization for visual question answering: Strengths, shortcomings, and side effects
Gabriel Grand and Yonatan Belinkov · 2019
Later among the works it cites.
Learning by abstraction: The neural state machine
Drew A Hudson and Christopher D Manning · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C Wallace · 2019
Later among the works it cites.
Relation-aware graph attention network for visual question answering
Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bilinear attention networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Cited alongside, same era.
Learning visual question answering by bootstrapping hard attention
Mateusz Malinowski, Carl Doersch, Adam Santoro, and Peter Battaglia · 2018
Cited alongside, same era.
Exploring human-like attention supervision in visual question answering
Tingting Qiao, Jianfeng Dong, and Duanqing Xu · 2018
Cited alongside, same era.
Overcoming language priors in visual question answering with adversarial regularization
Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee · 2018
Cited alongside, same era.
Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing
Vedika Agarwal, Rakshith Shetty, and Mario Fritz · 2019
Cited alongside, same era.
Don’t take the premise for granted: Mitigating artifacts in natural language inference
Yonatan Belinkov, Adam Poliak, Stuart M Shieber, Benjamin Van Durme, and Alexander M Rush · 2019
Cited alongside, same era.
Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection
Hedi Ben-Younes, Rémi Cadene, Nicolas Thome, and Matthieu Cord · 2019
Cited alongside, same era.
Later among the works it cites.
Simple but effective techniques to reduce biases
Rabeeh Karimi Mahabadi and James Henderson · 2019
Later among the works it cites.
Recursive visual attention in visual dialog
Yulei Niu, Hanwang Zhang, Manli Zhang, Jianhong Zhang, Zhiwu Lu, and Ji-Rong Wen · 2019
Later among the works it cites.
Question-conditioned counterfactual image generation for vqa
Jingjing Pan, Yash Goyal, and Stefan Lee · 2019
Later among the works it cites.
Taking a hint: Leveraging explanations to make vision and language models more grounded
Ramprasaath R Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Dhruv Batra, and Devi Parikh · 2019
Later among the works it cites.
Cycle-consistency for robust visual question answering
Meet Shah, Xinlei Chen, Marcus Rohrbach, and Devi Parikh · 2019
Later among the works it cites.
Learning to compose dynamic tree structures for visual contexts
Kaihua Tang, Hanwang Zhang, Baoyuan Wu, Wenhan Luo, and Wei Liu · 2019
Later among the works it cites.
Self-critical reasoning for robust visual question answering
Jialin Wu and Raymond J Mooney · 2019
Later among the works it cites.
Interpretable visual question answering by visual grounding from attention supervision mining
Yundong Zhang, Juan Carlos Niebles, and Alvaro Soto · 2019
Later among the works it cites.