Fetching the paper…
Reading the bibliography…
What does it take to design a machine that learns to answer natural questions about a video? A Video QA system must simultaneously understand language, represent visual content over space-time, and iteratively transform these representations in response to lingual content in the query, and finally arriving at a sensible answer.
The modularity of mind
Jerry A Fodor · 1983
Earlier work this paper cites.
The symbol grounding problem
Stevan Harnad · 1990
Earlier work this paper cites.
Dual-processing accounts of reasoning, judgment, and social cognition
Jonathan St BT Evans · 2008
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
From machine learning to machine reasoning
Léon Bottou · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Live repetition counting
Ofir Levy and Lior Wolf · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel · 2015
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning to reason: End-to-end module networks for visual question answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
DeepStory: video story QA by deep embedded memory networks
Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi, and Byoung-Tak Zhang · 2017
Cited alongside, same era.
A read-write memory network for movie story understanding
Seil Na, Sangho Lee, Jisung Kim, and Gunhee Kim · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Tim Lillicrap · 2017
Cited alongside, same era.
End-to-end concept word detection for video captioning, retrieval, and question answering
Youngjae Yu, Hyungjin Ko, Jongwook Choi, and Gunhee Kim · 2017
Cited alongside, same era.
Leveraging video descriptions to learn video question answering
Explore multi-step reasoning in video question answering
Xiaomeng Song, Yucheng Shi, Xin Chen, and Yahong Han · 2018
Later among the works it cites.
Interpretable counting for visual question answering
Alexander Trott, Caiming Xiong, and Richard Socher · 2018
Later among the works it cites.
Movie question answering: Remembering the textual cues for layered visual contents
Bo Wang, Youjiang Xu, Yahong Han, and Richang Hong · 2018
Later among the works it cites.
Neural-symbolic VQA: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum · 2018
Later among the works it cites.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Later among the works it cites.
End-to-end video-level representation learning for action recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kuo-Hao Zeng, Tseng-Hung Chen, Ching-Yao Chuang, Yuan-Hong Liao, Juan Carlos Niebles, and Min Sun · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Motion-appearance co-memory networks for video question answering
Jiyang Gao, Runzhou Ge, Kan Chen, and Ram Nevatia · 2018
Cited alongside, same era.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning · 2018
Cited alongside, same era.
Multimodal dual attention memory for video story question answering
Kyung-Min Kim, Seong-Ho Choi, Jin-Hwa Kim, and Byoung-Tak Zhang · 2018
Cited alongside, same era.
Focal visual-text attention for visual question answering
Junwei Liang, Lu Jiang, Liangliang Cao, Li-Jia Li, and Alexander G Hauptmann · 2018
Cited alongside, same era.
Jiagang Zhu, Zheng Zhu, and Wei Zou · 2018
Later among the works it cites.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Closest in time.
Artur d’Avila Garcez, Marco Gori, Luis C Lamb, Luciano Serafini, Michael Spranger, and Son N Tran · 2019
Closest in time.
Learning by abstraction: The neural state machine
Drew Hudson and Christopher D Manning · 2019
Closest in time.
Progressive attention memory network for movie story question answering
Junyeong Kim, Minuk Ma, Kyungsu Kim, Sungjin Kim, and Chang D Yoo · 2019
Closest in time.
Beyond RNNs: Positional Self-Attention with Co-Attention for Video Question Answering
Xiangpeng Li, Jingkuan Song, Lianli Gao, Xianglong Liu, Wenbing Huang, Xiangnan He, and Chuang Gan · 2019
Closest in time.