Fetching the paper…
Reading the bibliography…
Existing methods for video question answering (VideoQA) often suffer from spurious correlations between different modalities, leading to a failure in identifying the dominant visual evidence and the intended question.
Heterogeneous memory enhanced multimodal attention model for video question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 1999–2007
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang. 2019 · 2007
Earlier work this paper cites.
Controlling selection bias in causal inference. In Artificial Intelligence and Statistics . PMLR, 100–108
Elias Bareinboim and Judea Pearl. 2012 · 2012
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision . 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
End-to-end memory networks
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2016 · 2016
Earlier work this paper cites.
Causal inference in statistics: A primer
Judea Pearl, Madelyn Glymour, and Nicholas P Jewell. 2016 · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 21–29
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. 2016 · 2016
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2758–2766
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Right for the right reasons: training differentiable models by constraining their explanations. In Proceedings of the 26th International Joint Conference on Artificial Intelligence . 2662–2670
Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1492–1500
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion. In Proceedings of the 25th ACM international conference on Multimedia . 1645–1653
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017 · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 6077–6086
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Motion-appearance co-memory networks for video question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6576–6585
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia. 2018 · 2018
Cited alongside, same era.
Counterfactual critic multi-agent training for scene graph generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4613–4623
Long Chen, Hanwang Zhang, Jun Xiao, Xiangnan He, Shiliang Pu, and Shih-Fu Chang. 2019 · 2019
Cited alongside, same era.
Modularized textual grounding for counterfactual resilience. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6378–6388
Zhiyuan Fang, Shu Kong, Charless Fowlkes, and Yezhou Yang. 2019 · 2019
Cited alongside, same era.
Counterfactual visual explanations. In International Conference on Machine Learning . PMLR, 2376–2384
Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
Video question answering with spatio-temporal reasoning
Scout: Self-aware discriminant counterfactual explanations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8981–8990
Pei Wang and Nuno Vasconcelos. 2020 · 2020
Later among the works it cites.
Causal Intervention for Weakly-Supervised Semantic Segmentation
Dong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua, and Qianru Sun. 2020 · 2020
Later among the works it cites.
Adaptive spatio-temporal graph enhanced vision-language representation for video QA
Weike Jin, Zhou Zhao, Xiaochun Cao, Jieming Zhu, Xiuqiang He, and Yueting Zhuang. 2021 · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7331–7341
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. 2021 · 2021
Later among the works it cites.
Question-guided erasing-based spatiotemporal attention learning for video question answering
Fei Liu, Jing Liu, Richang Hong, and Hanqing Lu. 2021a · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yunseok Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2019 · 2019
Cited alongside, same era.
Multimodal explanations by predicting counterfactuality in videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8594–8602
Atsushi Kanehira, Kentaro Takemoto, Sho Inayoshi, and Tatsuya Harada. 2019 · 2019
Cited alongside, same era.
Beyond rnns: Positional self-attention with co-attention for video question answering. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 8658–8665
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan. 2019 · 2019
Cited alongside, same era.
Counterfactual vision and language learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10044–10054
Ehsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi, and Anton van den Hengel. 2020 · 2020
Cited alongside, same era.
Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9690–9698
Vedika Agarwal, Rakshith Shetty, and Mario Fritz. 2020 · 2020
Cited alongside, same era.
Counterfactuals uncover the modular structure of deep generative models. In Eighth International Conference on Learning Representations (ICLR 2020)
M Besserve, A Mehrjou, R Sun, and B Schölkopf. 2020 · 2020
Cited alongside, same era.
Counterfactual samples synthesizing for robust visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10800–10809
Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang. 2020 · 2020
Cited alongside, same era.
Location-aware graph convolutional networks for video question answering. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 11021–11028
Deng Huang, Peihao Chen, Runhao Zeng, Qing Du, Mingkui Tan, and Chuang Gan. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Interventional Video Grounding with Dual Contrastive Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2765–2775
Guoshun Nan, Rui Qiao, Yao Xiao, Jun Liu, Sicong Leng, Hao Zhang, and Wei Lu. 2021 · 2021
Later among the works it cites.
Counterfactual vqa: A cause-effect look at language bias. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12700–12710
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen. 2021 · 2021
Later among the works it cites.
Bridge to answer: Structure-aware graph interaction network for video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 15526–15535
Jungin Park, Jiyoung Lee, and Kwanghoon Sohn. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision. In International conference on machine learning . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Dualvgr: A dual-visual graph reasoning unit for video question answering
Jianyu Wang, Bing-Kun Bao, and Changsheng Xu. 2021a · 2021
Later among the works it cites.
Next-qa: Next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9777–9786
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua. 2021 · 2021
Later among the works it cites.
Deconfounded image captioning: A causal retrospect
Xu Yang, Hanwang Zhang, and Jianfei Cai. 2021a · 2021
Later among the works it cites.
Revisiting the" video" in video-language understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2917–2927
Shyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu, Li Fei-Fei, and Juan Carlos Niebles. 2022 · 2022
Later among the works it cites.
Causal Reasoning Meets Visual Representation Learning: A Prospective Study
Yang Liu, Yu-Shen Wei, Hong Yan, Guan-Bin Li, and Liang Lin. 2022a · 2022
Later among the works it cites.
Cross-Attentional Spatio-Temporal Semantic Graph Networks for Video Question Answering
Yun Liu, Xiaoming Zhang, Feiran Huang, Bo Zhang, and Zhoujun Li. 2022b · 2022
Later among the works it cites.
Video graph transformer for video question answering. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVI . Springer, 39–58
Junbin Xiao, Pan Zhou, Tat-Seng Chua, and Shuicheng Yan. 2022b · 2022
Later among the works it cites.
Visual-Linguistic Causal Intervention for Radiology Report Generation
Weixing Chen, Yang Liu, Ce Wang, Guanbin Li, Jiarui Zhu, and Liang Lin. 2023 · 2023
Closest in time.
Cross-modal causal relational reasoning for event-level visual question answering
Yang Liu, Guanbin Li, and Liang Lin. 2023 · 2023
Closest in time.