Overcoming language priors in visual question answering with adversarial regularization
Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee · 2018
Later among the works it cites.
Fast and accurate reading comprehension by combining self-attention and convolution
Adams Wei Yu, David Dohan, Quoc Le, Thang Luong, Rui Zhao, and Kai Chen · 2018
Later among the works it cites.
Audio visual scene-aware dialog
Huda AlAmri, Vincent Cartillier, Abhishek Das, Jue Wang, Stefan Lee, Pip Anderson, Irfan Essa, Devi Parikh, Dhruv Batra, Anoop Cherian, Tim K. Marks, and Chiori Hori · 2019
Later among the works it cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2019
Later among the works it cites.
Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection
Hedi Ben-Younes, Remi Cadene, Nicolas Thome, and Matthieu Cord · 2019
Later among the works it cites.
Rubi: Reducing unimodal biases in visual question answering
Rémi Cadène, Corentin Dancette, Hedi Ben-younes, Matthieu Cord, and Devi Parikh · 2019
Later among the works it cites.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, Stefan Lee, José M. F. Moura, Devi Parikh, and Dhruv Batra · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Egovqa - an egocentric video question answering benchmark dataset
Chenyou Fan · 2019
Later among the works it cites.
Deep multimodal representation learning: A survey
W. Guo, J. Wang, and S. Wang · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Original
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Later among the works it cites.
Counterfactual samples synthesizing for robust visual question answering
Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang · 2020
Closest in time.
CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
Rohit Girdhar and Deva Ramanan · 2020
Closest in time.
TVQA+: Spatio-temporal grounding for video question answering
Jie Lei, Licheng Yu, Tamara Berg, and Mohit Bansal · 2020
Closest in time.
Bert representations for video question answering
Zekun Yang, Noa Garcia, Chenhui Chu, Mayu Otani, Yuta Nakashima, and Haruo Takemura · 2020
Closest in time.
Clevrer: Collision events for video representation and reasoning
Kexin Yi*, Chuang Gan*, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum · 2020
Closest in time.
Multimodal intelligence: Representation learning, information fusion, and applications
Chao Zhang, Zichao Yang, Xiaodong He, and Li Deng · 2020
Closest in time.