Fetching the paper…
Reading the bibliography…
A key challenge in video question answering is how to realize the cross-modal semantic alignment between textual concepts and corresponding visual objects.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang. 2019 · 2007
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Y. Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Earlier work this paper cites.
A regularized framework for sparse and structured neural attention
Vlad Niculae and Mathieu Blondel. 2017 · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017 · 2017
Earlier work this paper cites.
Fusing temporally distributed multi-modal semantic clues for video question answering
Fuwei Zhang, Ruomei Wang, Songhua Xu, and Fan Zhou. 2021 · 2017
Earlier work this paper cites.
Motion-appearance co-memory networks for video question answering
J. Gao, Runzhou Ge, Kan Chen, and Ramakant Nevatia. 2018 · 2018
Earlier work this paper cites.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018 · 2018
Cited alongside, same era.
Multimodal dual attention memory for video story question answering
Kyung min Kim, Seongho Choi, Jin-Hwa Kim, and Byoung-Tak Zhang. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Hypergraph neural networks
Yifan Feng, Haoxuan You, Zizhao Zhang, Rongrong Ji, and Yue Gao. 2019 · 2019
Cited alongside, same era.
Video question answering with spatio-temporal reasoning
Y. Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2019 · 2019
Cited alongside, same era.
Beyond rnns: Positional self-attention with co-attention for video question answering
Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering
Jian wen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao, and Yue Gao. 2020 · 2020
Later among the works it cites.
A fast proximal point method for computing exact wasserstein distance
Yujia Xie, Xiangfeng Wang, Ruijia Wang, and Hongyuan Zha. 2020 · 2020
Later among the works it cites.
Memory augmented deep recurrent neural network for video question answering
Chengxiang Yin, Jian Tang, Zhiyuan Xu, and Yanzhi Wang. 2020 · 2020
Later among the works it cites.
A universal quaternion hypergraph network for multimodal video question answering
Zhicheng Guo, Jiaxuan Zhao, Licheng Jiao, Xu Liu, and Fang Liu. 2021 · 2021
Later among the works it cites.
Relation-aware hierarchical attention framework for video question answering
Fangtao Li, Ting Bai, Chenyu Cao, Zihe Liu, Chenghao Yan, and Bin Wu. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan. 2019 · 2019
Cited alongside, same era.
Question-aware tube-switch network for video question answering
Tianhao Yang, Zhengjun Zha, Hongtao Xie, Meng Wang, and Hanwang Zhang. 2019 · 2019
Cited alongside, same era.
Improving textual network embedding with global attention via optimal transport
Liqun Chen, Guoyin Wang, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang, Wenlin Wang, Yizhe Zhang, and Lawrence Carin. 2020 · 2020
Cited alongside, same era.
Location-aware graph convolutional networks for video question answering
Deng Huang, Peihao Chen, Runhao Zeng, Qing Du, Mingkui Tan, and Chuang Gan. 2020 · 2020
Cited alongside, same era.
Reasoning with heterogeneous graph alignment for video question answering
Pin Jiang and Yahong Han. 2020 · 2020
Cited alongside, same era.
Hierarchical conditional relation networks for video question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and T. Tran. 2020 · 2020
Cited alongside, same era.
Hair: Hierarchical visual-semantic relational reasoning for video question answering
Fei Liu, Jing Liu, Weining Wang, and Hanqing Lu. 2021 · 2021
Later among the works it cites.
Bridge to answer: Structure-aware graph interaction network for video question answering
Jungin Park, Jiyoung Lee, and Kwanghoon Sohn. 2021 · 2021
Later among the works it cites.
Attend what you need: Motion-appearance synergistic networks for video question answering
Ahjeong Seo, Gi-Cheon Kang, Joon Ki Park, and Byoung-Tak Zhang. 2021 · 2021
Later among the works it cites.
Dualvgr: A dual-visual graph reasoning unit for video question answering
Jianyu Wang, Bingkun Bao, and Changsheng Xu. 2021 · 2021
Later among the works it cites.
Syntax-enhanced pre-trained model
Zenan Xu, Daya Guo, Duyu Tang, Qinliang Su, Linjun Shou, Ming Gong, Wanjun Zhong, Xiaojun Quan, Daxin Jiang, and Nan Duan. 2021 · 2021
Later among the works it cites.