Fetching the paper…
Reading the bibliography…
Video Question Answering (QA) is an important task in understanding video temporal structure.
Two-frame motion estimation based on polynomial expansion
G. Farnebäck · 2003
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Abc-cnn: An attention based convolutional neural network for visual question answering
K. Chen, J. Wang, L.-C. Chen, H. Gao, W. Xu, and R. Nevatia · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Earlier work this paper cites.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
A focused dynamic attention model for visual question answering
I. Ilievski, S. Yan, and J. Feng · 2016
Earlier work this paper cites.
Multimodal residual learning for visual qa
J.-H. Kim, S.-W. Lee, D. Kwak, M.-O. Heo, J. Kim, J.-W. Ha, and B.-T. Zhang · 2016
Cited alongside, same era.
Ask me anything: Dynamic memory networks for natural language processing
A. Kumar, O. Irsoy, P. Ondruska, M. Iyyer, J. Bradbury, I. Gulrajani, V. Zhong, R. Paulus, and R. Socher · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2016
Cited alongside, same era.
Temporal action localization in untrimmed videos via multi-stage cnns
Z. Shou, D. Wang, and S.-F. Chang · 2016
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Cascaded boundary regression for temporal action detection
J. Gao, Z. Yang, and R. Nevatia · 2017
Later among the works it cites.
Red: Reinforced encoder-decoder networks for action anticipation
J. Gao, Z. Yang, and R. Nevatia · 2017
Later among the works it cites.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Later among the works it cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim · 2017
Later among the works it cites.
Deepstory: video story qa by deep embedded memory networks
K.-M. Kim, M.-O. Heo, S.-H. Choi, and B.-T. Zhang · 2017
Later among the works it cites.
Marioqa: Answering questions by watching gameplay videos
J. Mun, P. Hongsuck Seo, I. Jung, and B. Han · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Cited alongside, same era.
Cuhk & ethz & siat submission to activitynet challenge 2016
Y. Xiong, L. Wang, Z. Wang, B. Zhang, H. Song, W. Li, D. Lin, Y. Qiao, L. Van Gool, and X. Tang · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
Localizing moments in video with natural language
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell · 2017
Cited alongside, same era.
Later among the works it cites.
A read-write memory network for movie story understanding
S. Na, S. Lee, J. Kim, and G. Kim · 2017
Later among the works it cites.
Decomposing motion and content for natural video sequence prediction
R. Villegas, J. Yang, S. Hong, X. Lin, and H. Lee · 2017
Later among the works it cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Later among the works it cites.
Multi-level attention networks for visual question answering
D. Yu, J. Fu, T. Mei, and Y. Rui · 2017
Later among the works it cites.
End-to-end concept word detection for video captioning, retrieval, and question answering
Y. Yu, H. Ko, J. Choi, and G. Kim · 2017
Later among the works it cites.
Structured attentions for visual question answering
C. Zhu, Y. Zhao, S. Huang, K. Tu, and Y. Ma · 2017
Later among the works it cites.
Visual semantic planning using deep successor representations
Y. Zhu, D. Gordon, E. Kolve, D. Fox, L. Fei-Fei, A. Gupta, R. Mottaghi, and A. Farhadi · 2017
Later among the works it cites.
Knowledge aided consistency for weakly supervised phrase grounding
K. Chen, J. Gao, and R. Nevatia · 2018
Closest in time.