Fetching the paper…
Reading the bibliography…
Video question answering (VideoQA) is challenging as it requires modeling capacity to distill dynamic visual artifacts and distant relations and to associate them with linguistic concepts.
Abstracting home video automatically
Rainer Lienhart · 1999
Earlier work this paper cites.
Uncovering the temporal context for video question answering
Linchao Zhu, Zhongwen Xu, Yi Yang, and Alexander G Hauptmann · 2002
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Hierarchical recurrent neural encoder for video representation with application to captioning
Pingbo Pan, Zhongwen Xu, Yi Yang, Fei Wu, and Yueting Zhuang · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Hierarchical boundary-aware neural encoder for video captioning
Lorenzo Baraldi, Costantino Grana, and Rita Cucchiara · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
DeepStory: video story QA by deep embedded memory networks
Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang · 2017
Earlier work this paper cites.
Temporal modeling approaches for large-scale youtube-8m video understanding
Fu Li, Chuang Gan, Xiao Liu, Yunlong Bian, Xiang Long, Yandong Li, Zhichao Li, Jie Zhou, and Shilei Wen · 2017
Earlier work this paper cites.
A read-write memory network for movie story understanding
Seil Na, Sangho Lee, Jisung Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Cited alongside, same era.
Video question answering via attribute-augmented attention network learning
Yunan Ye, Zhou Zhao, Yimeng Li, Long Chen, Jun Xiao, and Yueting Zhuang · 2017
Cited alongside, same era.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao · 2017
Cited alongside, same era.
Leveraging video descriptions to learn video question answering
Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Carlos Niebles, and Min Sun · 2017
Cited alongside, same era.
Video question answering via hierarchical spatio-temporal attention networks
Zhou Zhao, Qifan Yang, Deng Cai, Xiaofei He, and Yueting Zhuang · 2017
Cited alongside, same era.
Non-local netvlad encoding for video classification
Yongyi Tang, Xing Zhang, Lin Ma, Jingwen Wang, Shaoxiang Chen, and Yu-Gang Jiang · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Later among the works it cites.
Movie question answering: Remembering the textual cues for layered visual contents
Bo Wang, Youjiang Xu, Yahong Han, and Richang Hong · 2018
Later among the works it cites.
Multi-turn video question answering via multi-stream hierarchical attention context network
Zhou Zhao, Xinghua Jiang, Deng Cai, Jun Xiao, Xiaofei He, and Shiliang Pu · 2018
Later among the works it cites.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Later among the works it cites.
Heterogeneous memory enhanced multimodal attention model for video question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical relational attention for video question answering
Muhammad Iqbal Hasan Chowdhury, Kien Nguyen, Sridha Sridharan, and Clinton Fookes · 2018
Cited alongside, same era.
Learning deep matrix representations
Kien Do, Truyen Tran, and Svetha Venkatesh · 2018
Cited alongside, same era.
Motion-appearance co-memory networks for video question answering
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
Bilinear attention networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Cited alongside, same era.
Multimodal dual attention memory for video story question answering
Kyung-Min Kim, Seong-Ho Choi, Jin-Hwa Kim, and Byoung-Tak Zhang · 2018
Cited alongside, same era.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Cited alongside, same era.
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Later among the works it cites.
Multi-interaction network with object relation for video question answering
Weike Jin, Zhou Zhao, Mao Gu, Jun Yu, Jun Xiao, and Yueting Zhuang · 2019
Later among the works it cites.
Progressive attention memory network for movie story question answering
Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, and Chang D Yoo · 2019
Later among the works it cites.
Learning to reason with relational video representation for question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2019
Later among the works it cites.
Learnable aggregating net with diversity learning for video question answering
Xiangpeng Li, Lianli Gao, Xuanhan Wang, Wu Liu, Xing Xu, Heng Tao Shen, and Jingkuan Song · 2019
Later among the works it cites.
Beyond RNNs: Positional Self-Attention with Co-Attention for Video Question Answering
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan · 2019
Later among the works it cites.
Spatio-temporal relational reasoning for video question answering
Gursimran Singh, Leonid Sigal, and James J Little · 2019
Later among the works it cites.
Holistic multi-modal memory network for movie question answering
Anran Wang, Anh Tuan Luu, Chuan-Sheng Foo, Hongyuan Zhu, Yi Tay, and Vijay Chandrasekhar · 2019
Later among the works it cites.
Long-term feature banks for detailed video understanding
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick · 2019
Later among the works it cites.
Question-aware tube-switch network for video question answering
Tianhao Yang, Zheng-Jun Zha, Hongtao Xie, Meng Wang, and Hanwang Zhang · 2019
Later among the works it cites.
Long-form video question answering via dynamic hierarchical reinforced networks
Zhou Zhao, Zhu Zhang, Shuwen Xiao, Zhenxin Xiao, Xiaohui Yan, Jun Yu, Deng Cai, and Fei Wu · 2019
Later among the works it cites.