Fetching the paper…
Reading the bibliography…
Computer vision has undergone a dramatic revolution in performance, driven in large part through deep features trained on large-scale supervised datasets.
Maintaining knowledge about temporal intervals
James F Allen · 1983
Earlier work this paper cites.
Reasoning about change: time and causation from the standpoint of artificial intelligence
Yoav Shoham · 1987
Earlier work this paper cites.
Movement, activity and action: the role of knowledge in the perception of motion
Aaron F Bobick · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Video-based event recognition: activity representation and probabilistic recognition methods
Somboon Hongeng, Ram Nevatia, and Francois Bremond · 2004
Earlier work this paper cites.
Recognizing human actions: a local svm approach
Christian Schuldt, Ivan Laptev, and Barbara Caputo · 2004
Earlier work this paper cites.
On space-time interest points
Ivan Laptev · 2005
Earlier work this paper cites.
Vidmap: video monitoring of activity with prolog
Vinay D. Shet, David Harwood, and Larry S. Davis · 2005
Earlier work this paper cites.
A new method for real time abandoned object detection and owner tracking
Silvia Ferrando, Gianluca Gera, Massimo Massa, and Carlo Regazzoni · 2006
Earlier work this paper cites.
Evaluation of an ivs system for abandoned object detection on pets 2006 datasets
Liyuan Li, Ruijiang Luo, Ruihua Ma, Weimin Huang, and Karianto Leman · 2006
Earlier work this paper cites.
Understanding video events: A survey of methods for automatic interpretation of semantic occurrences in video
Gal Lavee, Ehud Rivlin, and Michael Rudzsky · 2009
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Robust detection of abandoned and removed objects in complex surveillance videos
YingLi Tian, Rogerio Schmidt Feris, Haowei Liu, Arun Hampapur, and Ming-Ting Sun · 2011
Earlier work this paper cites.
Action Recognition by Dense Trajectories
Heng Wang, Alexander Kläser, Cordelia Schmid, and Liu Cheng-Lin · 2011
Earlier work this paper cites.
A naturalistic open source movie for optical flow evaluation
David J. Butler, Jonas Wulff, Garrett B. Stanley, and Michael J. Black · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Learning to detect carried objects with minimal supervision
Radu Dondera, Vlad Morariu, and Larry Davis · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal · 2013
Earlier work this paper cites.
TV-L1 Optical Flow Estimation
Javier Sánchez Pérez, Enric Meinhardt-Llopis, and Gabriele Facciolo · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Heng Wang and Cordelia Schmid · 2013
Earlier work this paper cites.
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Trevor Darrell · 2015
Earlier work this paper cites.
First-person pose recognition using egocentric workspaces
Grégory Rogez, James S Supančič, and Deva Ramanan · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Render for cnn: Viewpoint estimation in images using cnns trained with rendered 3d model views
Hao Su, Charles R. Qi, Yangyan Li, and Leonidas J. Guibas · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Cited alongside, same era.
Spatiotemporal residual networks for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Richard Wildes · 2016
Cited alongside, same era.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Cited alongside, same era.
Learning a predictable and generative vector representation for objects
Rohit Girdhar, David Fouhey, Mikel Rodriguez, and Abhinav Gupta · 2016
Cited alongside, same era.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Playing for benchmarks
Stephan R Richter, Zeeshan Hayder, and Vladlen Koltun · 2017
Later among the works it cites.
What actions are needed for understanding human actions in videos?
Gunnar A. Sigurdsson, Olga Russakovsky, and Abhinav Gupta · 2017
Later among the works it cites.
Semantic scene completion from a single depth image
Shuran Song, Fisher Yu, Andy Zeng, Angel X Chang, Manolis Savva, and Thomas Funkhouser · 2017
Later among the works it cites.
Human pose forecasting via deep markov models
Sam Toyer, Anoop Cherian, Tengda Han, and Stephen Gould · 2017
Later among the works it cites.
The daily home life activity dataset: A high semantic activity dataset for online recognition
Geoffrey Vaquette, Astrid Orcesi, Laurent Lucat, and Catherine Achard · 2017
Later among the works it cites.
Rethinking spatiotemporal feature learning for video understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski · 2016
Cited alongside, same era.
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Häusser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox · 2016
Cited alongside, same era.
NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, and Vincent Vanhoucke · 2016
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Cited alongside, same era.
Single image 3d interpreter network
Jiajun Wu, Tianfan Xue, Joseph J Lim, Yuandong Tian, Joshua B Tenenbaum, Antonio Torralba, and William T Freeman · 2016
Cited alongside, same era.
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy · 2017
Later among the works it cites.
Allen’s interval algebra
Thomas A. Alspaugh · 2018
Later among the works it cites.
From lifestyle VLOGs to everyday interactions
David F. Fouhey, Weicheng Kuo, Alexei A. Efros, and Jitendra Malik · 2018
Later among the works it cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions
Chunhui Gu, Chen Sun, David Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik · 2018
Later among the works it cites.
Resound: Towards action recognition without representation bias
Yingwei Li, Yi Li, and Nuno Vasconcelos · 2018
Later among the works it cites.
Attention clusters: Purely attention based local feature integration for video classification
Xiang Long, Chuang Gan, Gerard de Melo, Jiajun Wu, Xiao Liu, and Shilei Wen · 2018
Later among the works it cites.
Learning visual question answering by bootstrapping hard attention
Mateusz Malinowski, Carl Doersch, Adam Santoro, and Peter Battaglia · 2018
Later among the works it cites.
Intphys: A framework and benchmark for visual intuitive physics reasoning
Ronan Riochet, Mario Ynocente Castro, Mathieu Bernard, Adam Lerer, Rob Fergus, Véronique Izard, and Emmanuel Dupoux · 2018
Later among the works it cites.
Airsim: High-fidelity visual and physical simulation for autonomous vehicles
Shital Shah, Debadeepta Dey, Chris Lovett, and Ashish Kapoor · 2018
Later among the works it cites.
Explore multi-step reasoning in video question answering
Xiaomeng Song, Yucheng Shi, Xin Chen, and Yahong Han · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Later among the works it cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Later among the works it cites.
Building generalizable agents with a realistic and rich 3d environment
Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian · 2018
Later among the works it cites.
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Yuanjun Xiong · 2018
Later among the works it cites.
TSN Pretrained Models on Kinetics Dataset
Yuanjun Xiong · 2018
Later among the works it cites.
Spatial temporal graph convolutional networks for skeleton-based action recognition
Sijie Yan, Yuanjun Xiong, and Dahua Lin · 2018
Later among the works it cites.
Distractor-aware siamese networks for visual object tracking
Zheng Zhu, Qiang Wang, Li Bo, Wei Wu, Junjie Yan, and Weiming Hu · 2018
Later among the works it cites.
Adopting abstract images for semantic scene understanding
C Lawrence Zitnick, Ramakrishna Vedantam, and Devi Parikh · 2018
Later among the works it cites.
Phyre: A new benchmark for physical reasoning
Anton Bakhtin, Laurens van der Maaten, Justin Johnson, Laura Gustafson, and Ross Girshick · 2019
Closest in time.
Cophy: Counterfactual learning of physical dynamics
Fabien Baradel, Natalia Neverova, Julien Mille, Greg Mori, and Christian Wolf · 2020
Closest in time.
Sideways: Depth-parallel training of video models
Mateusz Malinowski, Grzegorz Swirszcz, Joao Carreira, and Viorica Patraucean · 2020
Closest in time.
Clevrer: Collision events for video representation and reasoning
Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, and Joshua B. Tenenbaum · 2020
Closest in time.