Fetching the paper…
Reading the bibliography…
The objective of this paper is self-supervised learning of spatio-temporal embeddings from video, suitable for human action recognition.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A. Efros · 1909
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Laurenz Wiskott and Terrence Sejnowski · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin · 2003
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael U. Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
Hilde Kuehne, Huei-han Jhuang, Estibaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Learning word embeddings efficiently with noise-contrastive estimation
Andriy Mnih and Koray Kavukcuoglu · 2013
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Learning to see by moving
Pulkit Agrawal, Joao Carreira, and Jitendra Malik · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A. Efros · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Learning visual groups from co-occurrences in space and time
Phillip Isola, Daniel Zoran, Dilip Krishnan, and Edward H Adelson · 2015
Earlier work this paper cites.
Learning image representations tied to ego-motion
Dinesh Jayaraman and Kristen Grauman · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Unsupervised learning of video representations using LSTMs
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhutdinov · 2015
Cited alongside, same era.
Unsupervised learning of visual representations using videos
Xiaolong Wang and Abhinav Gupta · 2015
Cited alongside, same era.
Slow and steady feature analysis: higher order temporal coherence in video
Dinesh Jayaraman and Kristen Grauman · 2016
Cited alongside, same era.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2016
Cited alongside, same era.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David D. Cox · 2017
Later among the works it cites.
Representation learning by learning to count
Mehdi Noroozi, Hamed Pirsiavash, and Paolo Favaro · 2017
Later among the works it cites.
Learning features by watching objects move
Deepak Pathak, Ross B. Girshick, Piotr Dollár, Trevor Darrell, and Bharath Hariharan · 2017
Later among the works it cites.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Later among the works it cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2018
Later among the works it cites.
Unsupervised representation learning by predicting image rotations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shuffle and learn: Unsupervised learning using temporal order verification
Ishan Misra, C. Lawrence Zitnick, and Martial Hebert · 2016
Cited alongside, same era.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Cited alongside, same era.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krähenbühl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros · 2016
Cited alongside, same era.
Learning a metric embedding for face recognition using the multibatch method
Oren Tadmor, Yonatan Wexler, Tal Rosenwein, Shai Shalev-Shwartz, and Amnon Shashua · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor S. Lempitsky · 2016
Cited alongside, same era.
Anticipating visual representations from unlabelled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Spyros Gidaris, Praveer Singh, and Nikos Komodakis · 2018
Later among the works it cites.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Later among the works it cites.
Self-supervised spatiotemporal feature learning by video geometric transformations
Longlong Jing and Yingli Tian · 2018
Later among the works it cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Tracking emerges by colorizing videos
Carl Vondrick, Abhinav Shrivastava, Alireza Fathi, Sergio Guadarrama, and Kevin Murphy · 2018
Later among the works it cites.
Learning and using the arrow of time
Donglai Wei, Joseph Lim, Andrew Zisserman, and William T. Freeman · 2018
Later among the works it cites.
Unsupervised state representation learning in atari
Ankesh Anand, Evan Racah, Sherjil Ozair, Yoshua Bengio, Marc-Alexandre Côté, and R. Devon Hjelm · 2019
Closest in time.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Closest in time.
Self-supervised learning for video correspondence flow
Zihang Lai and Weidi Xie · 2019
Closest in time.
Learning correspondence from the cycle-consistency of time
Xiaolong Wang, Allan Jabri, and Alexei A. Efros · 2019
Closest in time.