Fetching the paper…
Reading the bibliography…
We present a large-scale study on unsupervised spatiotemporal representation learning from videos.
Learning temporally persistent hierarchical representations
Suzanna Becker · 1997
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Laurenz Wiskott and Terrence Sejnowski · 2002
Earlier work this paper cites.
Deep learning from temporal coherence in video
Hossein Mobahi, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Learning to see by moving
Pulkit Agrawal, João Carreira, and Jitendra Malik · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei Efros · 2015
Earlier work this paper cites.
Fast R-CNN
R. B. Girshick · 2015
Earlier work this paper cites.
Unsupervised learning of spatiotemporally coherent metrics
Ross Goroshin, Joan Bruna, Jonathan Tompson, David Eigen, and Yann LeCun · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Learning visual groups from co-occurrences in space and time
Phillip Isola, Daniel Zoran, Dilip Krishnan, and Edward H Adelson · 2015
Earlier work this paper cites.
Learning image representations tied to ego-motion
Dinesh Jayaraman and Kristen Grauman · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
Xiaolong Wang and Abhinav Gupta · 2015
Earlier work this paper cites.
Discriminative unsupervised feature learning with exemplar convolutional neural networks
A. Dosovitskiy, P. Fischer, J. T. Springenberg, M. Riedmiller, and T. Brox · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2016
Earlier work this paper cites.
Shuffle and learn: Unsupervised learning using temporal order verification
Ishan Misra, C. Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Earlier work this paper cites.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh H. McDermott, Antonio Torralba, Edward H. Adelson, and William T. Freeman · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
Deepak Pathak, Philipp Krähenbühl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros · 2016
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Anticipating visual representations from unlabelled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A. Efros · 2016
Cited alongside, same era.
Look, listen and learn
Relja Arandjelović and Andrew Zisserman · 2017
Cited alongside, same era.
Self-supervised video representation learning with odd-one-out networks
Basura Fernando, Hakan Bilen, Efstratios Gavves, and Stephen Gould · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
The “Something Something” video database for learning and evaluating visual common sense
Invariant information clustering for unsupervised image classification and segmentation
Xu Ji, João F. Henriques, and Andrea Vedaldi · 2019
Later among the works it cites.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon · 2019
Later among the works it cites.
Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Later among the works it cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Later among the works it cites.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Cited alongside, same era.
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, João Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017
Cited alongside, same era.
Unsupervised representation learning by sorting sequence
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang · 2017
Cited alongside, same era.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2017
Cited alongside, same era.
Learning features by watching objects move
Deepak Pathak, Ross Girshick, Piotr Dollár, Trevor Darrell, and Bharath Hariharan · 2017
Cited alongside, same era.
Transitive invariance for self-supervised visual representation learning
Xiaolong Wang, Kaiming He, and Abhinav Gupta · 2017
Cited alongside, same era.
Learning correspondence from the cycle-consistency of time
Xiaolong Wang, Allan Jabri, and Alexei A. Efros · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang · 2019
Later among the works it cites.
Local aggregation for unsupervised learning of visual embeddings
Chengxu Zhuang, Alex Lin Zhai, and Daniel Yamins · 2019
Later among the works it cites.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Later among the works it cites.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Lorenzo Torresani, Bernard Ghanem, and Du Tran · 2020
Later among the works it cites.
SpeedNet: Learning the Speediness in Videos
Sagie Benaim, Ariel Ephrat, Oran Lang, Inbar Mosseri, William T. Freeman, Michael Rubinstein, Michal Irani, and Tali Dekel · 2020
Later among the works it cites.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Later among the works it cites.
Big self-supervised models are strong semi-supervised learners
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton · 2020
Later among the works it cites.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Later among the works it cites.
PySlowFast
Haoqi Fan, Yanghao Li, Bo Xiong, Wan-Yen Lo, and Christoph Feichtenhofer · 2020
Later among the works it cites.
Watching the world go by: Representation learning from unlabeled videos
Daniel Gordon, Kiana Ehsani, Dieter Fox, and Ali Farhadi · 2020
Later among the works it cites.
Bootstrap your own latent: A new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko · 2020
Later among the works it cites.
Self-supervised co-training for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Video representation learning by recognizing temporal transformations
Simon Jenni, Givi Meishvili, and Paolo Favaro · 2020
Later among the works it cites.
The ava-kinetics localized human actions video dataset
Ang Li, Meghana Thotakuri, David A Ross, João Carreira, Alexander Vostrikov, and Andrew Zisserman · 2020
Later among the works it cites.
Learning spatiotemporal features via video and text pair discrimination
Tianhao Li and Limin Wang · 2020
Later among the works it cites.
Audio-visual instance discrimination with cross-modal agreement
Pedro Morgado, Nuno Vasconcelos, and Ishan Misra · 2020
Later among the works it cites.
Multi-modal self-supervision from generalized data transformations
Mandela Patrick, Yuki M. Asano, Ruth Fong, João F. Henriques, Geoffrey Zweig, and Andrea Vedaldi · 2020
Later among the works it cites.
Evolving losses for unsupervised video representation learning
AJ Piergiovanni, Anelia Angelova, and Michael S. Ryoo · 2020
Later among the works it cites.
Demystifying contrastive self-supervised learning: Invariances, augmentations and dataset biases
Senthil Purushwalkam and Abhinav Gupta · 2020
Later among the works it cites.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui · 2020
Later among the works it cites.
Byol works even without batch statistics
Pierre H Richemond, Jean-Bastien Grill, Florent Altché, Corentin Tallec, Florian Strub, Andrew Brock, Samuel Smith, Soham De, Razvan Pascanu, Bilal Piot, et al · 2020
Later among the works it cites.
Contrastive multiview coding
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Later among the works it cites.
Video representation learning with visual tempo consistency
Ceyuan Yang, Yinghao Xu, Bo Dai, and Bolei Zhou · 2020
Later among the works it cites.