Fetching the paper…
Reading the bibliography…
Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language processing.
Linear hinge loss and average margin
Claudio Gentile and Manfred K. Warmuth · 1998
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Hierarchical relational attention for video question answering
Muhammad Iqbal Hasan Chowdhury, Kien Nguyen, Sridha Sridharan, and Clinton Fookes · 2018
Earlier work this paper cites.
Motion-appearance co-memory networks for video question answering
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Cited alongside, same era.
Video action transformer network
Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019
Cited alongside, same era.
Progressive attention memory network for movie story question answering
Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, and Chang D Yoo · 2019
Cited alongside, same era.
Beyond rnns: Positional self-attention with co-attention for video question answering
Hierarchical conditional relation networks for video question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Later among the works it cites.
Polytransform: Deep polygon transformer for instance segmentation
Justin Liang, Namdar Homayounfar, Wei-Chiu Ma, Yuwen Xiong, Rui Hu, and Raquel Urtasun · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Later among the works it cites.
Up-detr: Unsupervised pre-training for object detection with transformers
Zhigang Dai, Bolun Cai, Yugeng Lin, and Junying Chen · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan · 2019
Cited alongside, same era.
Spatiotemporal-textual co-attention network for video question answering
Zheng-Jun Zha, Jiawei Liu, Tianhao Yang, and Yongdong Zhang · 2019
Cited alongside, same era.
Feature augmented memory with global attention network for videoqa
Jiayin Cai, Chun Yuan, Cheng Shi, Lei Li, Yangyang Cheng, and Ying Shan · 2020
Cited alongside, same era.
Location-aware graph convolutional networks for video question answering
Deng Huang, Peihao Chen, Runhao Zeng, Qing Du, Mingkui Tan, and Chuang Gan · 2020
Cited alongside, same era.
Reasoning with heterogeneous graph alignment for video question answering
Pin Jiang and Yahong Han · 2020
Cited alongside, same era.
Hierarchical object-oriented spatio-temporal reasoning for video question answering
Long Hoang Dang, Thao Minh Le, Vuong Le, and Truyen Tran · 2021
Later among the works it cites.
Chop chop bert: Visual question answering by chopping visualbert’s heads
Chenyu Gao, Qi Zhu, Peng Wang, and Qi Wu · 2021
Later among the works it cites.
Hair: Hierarchical visual-semantic relational reasoning for video question answering
Fei Liu, Jing Liu, Weining Wang, and Hanqing Lu · 2021
Later among the works it cites.
Bridge to answer: Structure-aware graph interaction network for video question answering
Jungin Park, Jiyoung Lee, and Kwanghoon Sohn · 2021
Later among the works it cites.