Fetching the paper…
Reading the bibliography…
We study the problem of dynamic visual reasoning on raw videos.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman · 1960
Earlier work this paper cites.
Bidirectional lstm networks for improved phoneme classification and recognition
Alex Graves, Santiago Fernández, and Jürgen Schmidhuber · 2005
Earlier work this paper cites.
Simulation as an engine of physical scene understanding
Peter W Battaglia, Jessica B Hamrick, and Joshua B Tenenbaum · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Imagining the unseen: Stability-based cuboid arrangements for scene understanding
Tianjia Shao, Aron Monszpart, Youyi Zheng, Bongjin Koo, Weiwei Xu, Kun Zhou, and Niloy J Mitra · 2014
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Learning perceptual causality from video
Amy Fire and Song-Chun Zhu · 2015
Earlier work this paper cites.
Devnet: A deep event network for multimedia event detection and evidence recounting
Chuang Gan, Naiyan Wang, Yi Yang, Dit-Yan Yeung, and Alex G Hauptmann · 2015
Earlier work this paper cites.
Finding action tubes
Georgia Gkioxari and Jitendra Malik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Simple online and realtime tracking
Alex Bewley, Zongyuan Ge, Lionel Ott, Fabio Ramos, and Ben Upcroft · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning physical intuition of block towers by example
Adam Lerer, Sam Gross, and Rob Fergus · 2016
Earlier work this paper cites.
“what happens if…” learning to predict the effect of forces in images
Roozbeh Mottaghi, Mohammad Rastegari, Abhinav Gupta, and Ali Farhadi · 2016
Earlier work this paper cites.
MovieQA: Understanding Stories in Movies through Question-Answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Cited alongside, same era.
What value do explicit high level concepts have in vision to language problems?
Qi Wu, Chunhua Shen, Lingqiao Liu, Anthony Dick, and Anton van den Hengel · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Vqs: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation
Chuang Gan, Yandong Li, Haoxiang Li, Chen Sun, and Boqing Gong · 2017
Cited alongside, same era.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Cited alongside, same era.
Phyre: A new benchmark for physical reasoning
Anton Bakhtin, Laurens van der Maaten, Justin Johnson, Laura Gustafson, and Ross Girshick · 2019
Later among the works it cites.
Weakly-supervised spatio-temporally grounding natural sentence in video
Zhenfang Chen, Lin Ma, Wenhan Luo, and Kwan-Yee Kenneth Wong · 2019
Later among the works it cites.
Heterogeneous memory enhanced multimodal attention model for video question answering
Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang, Chi Zhang, and Heng Huang · 2019
Later among the works it cites.
Learning by abstraction: The neural state machine
Drew Hudson and Christopher D Manning · 2019
Later among the works it cites.
The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision
Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B Tenenbaum, and Jiajun Wu · 2019
Later among the works it cites.
Grounded video description
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simple online and realtime tracking with a deep association metric
Nicolai Wojke, Alex Bewley, and Dietrich Paulus · 2017
Cited alongside, same era.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Cited alongside, same era.
Video question answering via attribute-augmented attention network learning
Yunan Ye, Zhou Zhao, Yimeng Li, Long Chen, Jun Xiao, and Yueting Zhuang · 2017
Cited alongside, same era.
Explainable neural computation via stack neural module networks
Ronghang Hu, Jacob Andreas, Trevor Darrell, and Kate Saenko · 2018
Cited alongside, same era.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning · 2018
Cited alongside, same era.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Cited alongside, same era.
Transparency by design: Closing the gap between performance and interpretability in visual reasoning
David Mascharka, Philip Tran, Ryan Soklaski, and Arjun Majumdar · 2018
Cited alongside, same era.
Luowei Zhou, Yannis Kalantidis, Xinlei Chen, Jason J Corso, and Marcus Rohrbach · 2019
Later among the works it cites.
Neuro-symbolic visual reasoning: Disentangling” visual” from” reasoning”
Saeed Amizadeh, Hamid Palangi, Oleksandr Polozov, Yichen Huang, and Kazuhito Koishida · 2020
Later among the works it cites.
Cophy: Counterfactual learning of physical dynamics
Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, and Christian Wolf · 2020
Later among the works it cites.
Fine-grained video-text retrieval with hierarchical graph reasoning
Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu · 2020
Later among the works it cites.
Threedworld: A platform for interactive multi-modal physical simulation
Chuang Gan, Jeremy Schwartz, Seth Alter, Martin Schrimpf, James Traer, Julian De Freitas, Jonas Kubilius, Abhishek Bhandwaldar, Nick Haber, Megumi Sano, et al · 2020
Later among the works it cites.
Cater: A diagnostic dataset for compositional actions and temporal reasoning
Rohit Girdhar and Deva Ramanan · 2020
Later among the works it cites.
Location-aware graph convolutional networks for video question answering
Deng Huang, Peihao Chen, Runhao Zeng, Qing Du, Mingkui Tan, and Chuang Gan · 2020
Later among the works it cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Jingwei Ji, Ranjay Krishna, Li Fei-Fei, and Juan Carlos Niebles · 2020
Later among the works it cites.
Hierarchical conditional relation networks for video question answering
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran · 2020
Later among the works it cites.
Closed loop neural-symbolic learning via integrating neural perception, grammar parsing, and symbolic reasoning
Qing Li, Siyuan Huang, Yining Hong, Yixin Chen, Ying Nian Wu, and Song-Chun Zhu · 2020
Later among the works it cites.
Something-else: Compositional action recognition with spatial-temporal interaction networks
Joanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu, Xiaolong Wang, and Trevor Darrell · 2020
Later among the works it cites.
Clevrer: Collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B Tenenbaum · 2020
Later among the works it cites.