Fetching the paper…
Reading the bibliography…
Deep neural networks require collecting and annotating large amounts of data to train successfully.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, M. Shah, K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space, 2013
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Earlier work this paper cites.
Skip-thought vectors, 2015
R. Kiros, Y. Zhu, R. Salakhutdinov, R. S. Zemel, A. Torralba, R. Urtasun, and S. Fidler · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
X. Wang and A. Gupta · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Unsupervised learning using sequential verification for action recognition
I. Misra, C. L. Zitnick, and M. Hebert · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
D. Pathak, P. Krähenbühl, J. Donahue, T. Darrell, and A. Efros · 2016
Cited alongside, same era.
Generating videos with scene dynamics
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Cited alongside, same era.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Cited alongside, same era.
Quo vadis, action recognition? A new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Unsupervised representation learning by sorting sequences
H. Lee, J. Huang, M. Singh, and M. Yang · 2017
Cited alongside, same era.
Unsupervised representation learning by predicting image rotations
Self-supervised spatiotemporal feature learning via video rotation prediction, 2018
L. Jing, X. Yang, J. Liu, and Y. Tian · 2018
Later among the works it cites.
Self-supervised video representation learning with space-time cubic puzzles
D. Kim, D. Cho, and I. S. Kweon · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Later among the works it cites.
Tracking emerges by colorizing videos
C. M. Vondrick, A. Shrivastava, A. Fathi, S. Guadarrama, and K. Murphy · 2018
Later among the works it cites.
Learning and using the arrow of time
D. Wei, J. Lim, A. Zisserman, and W. T. Freeman · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Gidaris, P. Singh, and N. Komodakis · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
K. Hara, H. Kataoka, and Y. Satoh · 2018
Cited alongside, same era.
Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics, 2019
J. Wang, J. Jiao, L. Bao, S. He, Y. Liu, and W. Liu · 2019
Closest in time.
Self-supervised spatiotemporal learning via video clip order prediction
D. Xu, J. Xiao, Z. Zhao, J. Shao, D. Xie, and Y. Zhuang · 2019
Closest in time.