Fetching the paper…
Reading the bibliography…
High level understanding of sequential visual input is important for safe and stable autonomy, especially in localization and object detection.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Best practices for convolutional neural networks applied to visual document analysis
Patrice Y Simard, David Steinkraus, John C Platt, et al · 2003
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Recurrent models of visual attention
Volodymyr Mnih, Nicolas Heess, Alex Graves, et al · 2014
Earlier work this paper cites.
Keras, 2015
François Chollet et al · 2015
Earlier work this paper cites.
Deformable part models are convolutional neural networks
Ross Girshick, Forrest Iandola, Trevor Darrell, and Jitendra Malik · 2015
Earlier work this paper cites.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al · 2015
Cited alongside, same era.
An elastic deformation field model for object detection and tracking
Marco Pedersoli, Radu Timofte, Tinne Tuytelaars, and Luc Van Gool · 2015
Cited alongside, same era.
Action recognition using visual attention
Shikhar Sharma, Ryan Kiros, and Ruslan Salakhutdinov · 2015
Cited alongside, same era.
Recurrent spatial transformer networks
Søren Kaae Sønderby, Casper Kaae Sønderby, Lars Maaløe, and Ole Winther · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Cited alongside, same era.
Inverse compositional spatial transformer networks
Chen-Hsuan Lin and Simon Lucey · 2016
Later among the works it cites.
Spatial transformer networks, 2017
Octavio Arriaga · 2017
Closest in time.
Deformable convolutional networks
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei · 2017
Closest in time.
Five video classification methods implemented in keras and tensorflow, 2017
Matt Harvey · 2017
Closest in time.
Tracking and classification of in-air hand gesture based on thermal guided joint filter
Seongwan Kim, Yuseok Ban, and Sangyoun Lee · 2017
Closest in time.
Notes on “deformable convolutional networks”, 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hao Su, Charles R Qi, Yangyan Li, and Leonidas J Guibas · 2015
Cited alongside, same era.
Attend, infer, repeat: Fast scene understanding with generative models. arxiv preprint arxiv:…, 2016
SM Eslami, N Heess, and T Weber · 2016
Cited alongside, same era.
Lenet-5, convolutional neural networks
Yann LeCun et al
Cited in the paper.
Felix Lau · 2017
Closest in time.