Fetching the paper…
Reading the bibliography…
This paper presents TCE: Temporally Coherent Embeddings for self-supervised video representation learning.
M. Gutmann and A. Hyvärinen, “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics , 2010, pp. 297–304
2010
Earlier work this paper cites.
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “HMDB: a large video database for human motion recognition,” in 2011 International Conference on Computer Vision . IEEE, 2011, pp. 2556–2563
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
A. Mnih and K. Kavukcuoglu, “Learning word embeddings efficiently with noise-contrastive estimation,” in Advances in neural information processing systems , 2013, pp. 2265–2273
2013
Earlier work this paper cites.
N. Srivastava, E. Mansimov, and R. Salakhudinov, “Unsupervised learning of video representations using lstms,” in International conference on machine learning , 2015, pp. 843–852
2015
Earlier work this paper cites.
X. Wang and A. Gupta, “Unsupervised learning of visual representations using videos,” in Proceedings of the IEEE International Conference on Computer Vision , 2015, pp. 2794–2802
2015
Earlier work this paper cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 815–823
2015
Earlier work this paper cites.
M. Mathieu, C. Couprie, and Y. LeCun, “Deep multi-scale video prediction beyond mean square error,” ICLR , 2016
2016
Earlier work this paper cites.
C. Vondrick, H. Pirsiavash, and A. Torralba, “Generating videos with scene dynamics,” in Advances in neural information processing systems , 2016, pp. 613–621
2016
Earlier work this paper cites.
C. Vondrick, H. Pirsiavash, and A. Torralba, “Anticipating visual representations from unlabeled video,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 98–106
2016
Earlier work this paper cites.
I. Misra, C. L. Zitnick, and M. Hebert, “Shuffle and learn: unsupervised learning using temporal order verification,” in European Conference on Computer Vision . Springer, 2016, pp. 527–544
2016
Earlier work this paper cites.
D. Jayaraman and K. Grauman, “Slow and steady feature analysis: higher order temporal coherence in video,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 3852–3861
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
W. Lotter, G. Kreiman, and D. Cox, “Deep predictive coding networks for video prediction and unsupervised learning,” in International Conference on Learning Representations , 2017
2017
Cited alongside, same era.
B. Fernando, H. Bilen, E. Gavves, and S. Gould, “Self-supervised video representation learning with odd-one-out networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 3636–3645
2017
Cited alongside, same era.
H.-Y. Lee, J.-B. Huang, M. Singh, and M.-H. Yang, “Unsupervised representation learning by sorting sequences,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 667–676
2017
Cited alongside, same era.
K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3D CNNs retrace the history of 2D cnns and imagenet?” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2018, pp. 6546–6555
2018
Later among the works it cites.
2018
Later among the works it cites.
D. Wei, J. J. Lim, A. Zisserman, and W. T. Freeman, “Learning and using the arrow of time,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8052–8060
2018
Later among the works it cites.
T. Han, W. Xie, and A. Zisserman, “Video representation learning by dense predictive coding,” in Proceedings of the IEEE International Conference on Computer Vision Workshops , 2019, pp. 0–0
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Zhang, P. Isola, and A. A. Efros, “Split-brain autoencoders: Unsupervised learning by cross-channel prediction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 1058–1067
2017
Cited alongside, same era.
T. Milbich, M. Bautista, E. Sutter, and B. Ommer, “Unsupervised video understanding by reconciliation of posture similarities,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 4394–4404
2017
Cited alongside, same era.
B. Harwood, V. Kumar BG, G. Carneiro, I. Reid, and T. Drummond, “Smart mining for deep metric learning,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 2821–2829
2017
Cited alongside, same era.
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman, “The kinetics human action video dataset,” 2017
2017
Cited alongside, same era.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2018, pp. 6450–6459
2018
Cited alongside, same era.
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain, “Time-contrastive networks: Self-supervised learning from video,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 1134–1141
2018
Cited alongside, same era.
C. Gan, B. Gong, K. Liu, H. Su, and L. J. Guibas, “Geometry guided convolutional neural networks for self-supervised video representation learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 5589–5597
2018
Cited alongside, same era.
S. Gidaris, P. Singh, and N. Komodakis, “Unsupervised representation learning by predicting image rotations,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
D. Kim, D. Cho, and I. S. Kweon, “Self-supervised video representation learning with space-time cubic puzzles,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, 2019, pp. 8545–8552
2019
Later among the works it cites.
D. Xu, J. Xiao, Z. Zhao, J. Shao, D. Xie, and Y. Zhuang, “Self-supervised spatiotemporal learning via video clip order prediction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 10 334–10 343
2019
Later among the works it cites.
P. Bachman, R. D. Hjelm, and W. Buchwalter, “Learning representations by maximizing mutual information across views,” in Advances in Neural Information Processing Systems , 2019, pp. 15 535–15 545
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Wang, J. Jiao, L. Bao, S. He, Y. Liu, and W. Liu, “Self-supervised spatio-temporal representation learning for videos by predicting motion and appearance statistics,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 4006–4015
2019
Later among the works it cites.
2020
Closest in time.
M. Tschannen, J. Djolonga, P. K. Rubenstein, S. Gelly, and M. Lucic, “On mutual information maximization for representation learning,” in International Conference on Learning Representations , 2020
2020
Closest in time.