Fetching the paper…
Reading the bibliography…
Many recent datasets contain a variety of different data modalities, for instance, image, question, and answer data in visual question answering (VQA).
Analytic inequalities, isoperimetric inequalities and logarithmic sobolev inequalities
OS Rothaus · 1985
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
An introduction to variational methods for graphical models
Michael Jordan, Zoubin Ghahramani, Tommi Jaakkola, and Lawrence Saul · 1999
Earlier work this paper cites.
The Concentration of Measure Phenomenon
Michel Ledoux · 2001
Earlier work this paper cites.
Analysis and geometry of Markov diffusion operators
Dominique Bakry, Ivan Gentil, and Michel Ledoux · 2013
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross B. Girshick · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Visual Dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M.F. Moura, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
High-order attention models for visual question answering
Idan Schwartz, Alexander Schwing, and Tamir Hazan · 2017
Earlier work this paper cites.
Creativity: Generating Diverse Questions using Variational Autoencoders
U. Jain ∗ , Z. Zhang ∗ , and A. G. Schwing · 2017
Earlier work this paper cites.
The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task
Roy Schwartz, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi, and Noah A. Smith · 2017
Earlier work this paper cites.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Cited alongside, same era.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Cited alongside, same era.
Learning not to learn: Training deep neural networks with biased data
Byungju Kim, Hyunwoo Kim, Kyungsu Kim, Sungjin Kim, and Junmo Kim · 2018
Cited alongside, same era.
Bilinear Attention Networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Factor graph attention
Idan Schwartz, Seunghak Yu, Tamir Hazan, and Alexander G. Schwing · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning
J. Aneja ∗ , H. Agrawal ∗ , D. Batra, and A. G. Schwing · 2019
Later among the works it cites.
TAB-VCR: Tags and Attributes based VCR Baselines
J. Lin, U. Jain, and A. G. Schwing · 2019
Later among the works it cites.
A simple baseline for audio-visual scene-aware dialog
Idan Schwartz, Alexander G Schwing, and Tamir Hazan · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Drew A Hudson and Christopher D Manning · 2018
Cited alongside, same era.
Two can play this Game: Visual Dialog with Discriminative Question Generation and Answering
U. Jain, S. Lazebnik, and A. G. Schwing · 2018
Cited alongside, same era.
Diverse and Coherent Paragraph Generation from Images
M. Chatterjee and A. G. Schwing · 2018
Cited alongside, same era.
Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering
M. Narasimhan, S. Lazebnik, and A. G. Schwing · 2018
Cited alongside, same era.
Straight to the Facts: Learning Knowledge Base Retrieval for Factual Visual Question Answering
M. Narasimhan and A. G. Schwing · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith · 2018
Cited alongside, same era.
Overcoming language priors in visual question answering with adversarial regularization
Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee · 2018
Cited alongside, same era.
Later among the works it cites.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Later among the works it cites.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer · 2019
Later among the works it cites.
Rubi: Reducing unimodal biases for visual question answering
Remi Cadene, Corentin Dancette, Hedi Ben younes, Matthieu Cord, and Devi Parikh · 2019
Later among the works it cites.
Information maximizing visual question generation
Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2019
Later among the works it cites.
Taking a hint: Leveraging explanations to make vision and language models more grounded
Ramprasaath R Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Shalini Ghosh, Larry Heck, Dhruv Batra, and Devi Parikh · 2019
Later among the works it cites.
Self-critical reasoning for robust visual question answering
Jialin Wu and Raymond Mooney · 2019
Later among the works it cites.
Adversarial regularization for visual question answering: Strengths, shortcomings, and side effects
Gabriel Grand and Yonatan Belinkov · 2019
Later among the works it cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Closest in time.
Approximate fisher information matrix to characterize the training of deep neural networks
Zhibin Liao, Tom Drummond, Ian Reid, and Gustavo Carneiro · 2020
Closest in time.
A negative case analysis of visual grounding methods for vqa
Kushal Kafle Robik Shrestha and Christopher Kanan · 2020
Closest in time.