Fetching the paper…
Reading the bibliography…
The objective of this paper is visual-only self-supervised video representation learning.
Combining labeled and unlabeled data with co-training
A. Blum and T. Mitchell · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
A duality based approach for realtime TV-L1 optical flow
C. Zach, T. Pock, and H. Bischof · 2007
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. U. Gutmann and A. Hyvärinen · 2010
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Learning to see by moving
P. Agrawal, J. Carreira, and J. Malik · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. Efros · 2015
Earlier work this paper cites.
Learning visual groups from co-occurrences in space and time
P. Isola, D. Zoran, D. Krishnan, and E. H. Adelson · 2015
Earlier work this paper cites.
Learning image representations tied to ego-motion
D. Jayaraman and K. Grauman · 2015
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Earlier work this paper cites.
Exploring the limits of language modeling
R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu · 2016
Earlier work this paper cites.
Shuffle and learn: Unsupervised learning using temporal order verification
I. Misra, C. L. Zitnick, and M. Hebert · 2016
Earlier work this paper cites.
Unsupervised learning of visual representations by solving jigsaw puzzles
M. Noroozi and P. Favaro · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
D. Pathak, P. Krähenbühl, J. Donahue, T. Darrell, and A. A. Efros · 2016
Earlier work this paper cites.
Anticipating visual representations from unlabelled video
C. Vondrick, H. Pirsiavash, and A. Torralba · 2016
Earlier work this paper cites.
Colorful image colorization
R. Zhang, P. Isola, and A. A. Efros · 2016
Earlier work this paper cites.
Look, listen and learn
R. Arandjelović and A. Zisserman · 2017
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
Multi-task self-supervised visual learning
C. Doersch and A. Zisserman · 2017
Earlier work this paper cites.
Self-supervised video representation learning with odd-one-out networks
B. Fernando, H. Bilen, E. Gavves, and S. Gould · 2017
Earlier work this paper cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Earlier work this paper cites.
Unsupervised representation learning by sorting sequence
H. Lee, J. Huang, M. Singh, and M. Yang · 2017
Cited alongside, same era.
Objects that sound
R. Arandjelović and A. Zisserman · 2018
Cited alongside, same era.
Improving spatiotemporal self-supervision by deep reinforcement learning
U. Büchler, B. Brattoli, and B. Ommer · 2018
Cited alongside, same era.
Deep clustering for unsupervised learning of visual features
M. Caron, P. Bojanowski, A. Joulin, and M. Douze · 2018
Cited alongside, same era.
Self-supervised spatiotemporal feature learning by video geometric transformations
L. Jing and Y. Tian · 2018
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
B. Korbar, D. Tran, and L. Torresani · 2018
Contrastive multiview coding
Y. Tian, D. Krishnan, and P. Isola · 2019
Later among the works it cites.
Learning correspondence from the cycle-consistency of time
X. Wang, A. Jabri, and A. A. Efros · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
D. Xu, J. Xiao, Z. Zhao, J. Shao, D. Xie, and Y. Zhuang · 2019
Later among the works it cites.
Dance with flow: Two-in-one stream action detection
J. Zhao and C. Snoek · 2019
Later among the works it cites.
Local aggregation for unsupervised learning of visual embeddings
C. Zhuang, A. L. Zhai, and D. Yamins · 2019
Later among the works it cites.
Self-labelling via simultaneous clustering and representation learning
Y. M. Asano, C. Rupprecht, and A. Vedaldi · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D3D: distilled 3d networks for video action recognition
J. C. Stroud, D. A. Ross, C. Sun, J. Deng, and R. Sukthankar · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Tracking emerges by colorizing videos
C. Vondrick, A. Shrivastava, A. Fathi, S. Guadarrama, and K. Murphy · 2018
Cited alongside, same era.
Learning and using the arrow of time
D. Wei, J. Lim, A. Zisserman, and W. T. Freeman · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning for video understanding
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy · 2018
Cited alongside, same era.
Closest in time.
SpeedNet: Learning the Speediness in Videos
S. Benaim, A. Ephrat, O. Lang, I. Mosseri, W. T. Freeman, M. Rubinstein, M. Irani, and T. Dekel · 2020
Closest in time.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Closest in time.
Improved baselines with momentum contrastive learning
X. Chen, H. Fan, R. Girshick, and K. He · 2020
Closest in time.
Oops! predicting unintentional action in video
D. Epstein, B. Chen, and C. Vondrick · 2020
Closest in time.
X3D: Expanding Architectures for Efficient Video Recognitionion
C. Feichtenhofer · 2020
Closest in time.
Bootstrap your own latent: A new approach to self-supervised learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. A. Pires, Z. D. Guo, M. G. Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. Valko · 2020
Closest in time.
Memory-augmented dense predictive coding for video representation learning
T. Han, W. Xie, and A. Zisserman · 2020
Closest in time.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, A. Wu, S. Xie, and R. Girshick · 2020
Closest in time.
Supervised contrastive learning
P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan · 2020
Closest in time.
MAST: A memory-augmented self-supervised tracker
Z. Lai, E. Lu, and W. Xie · 2020
Closest in time.
Video cloze procedure for self-supervised spatio-temporal learning
D. Luo, C. Liu, Y. Zhou, D. Yang, C. Ma, Q. Ye, and W. Wang · 2020
Closest in time.
End-to-end learning of visual representations from uncurated instructional videos
A. Miech, J.-B. Alayrac, L. Smaira, I. Laptev, J. Sivic, and A. Zisserman · 2020
Closest in time.
Self-supervised learning of pretext-invariant representations
I. Misra and L. van der Maaten · 2020
Closest in time.
Multi-modal self-supervision from generalized data transformations
M. Patrick, Y. M. Asano, R. Fong, J. F. Henriques, G. Zweig, and A. Vedaldi · 2020
Closest in time.
Evolving losses for unsupervised video representation learning
A. Piergiovanni, A. Angelova, and M. S. Ryoo · 2020
Closest in time.
Spatiotemporal contrastive video representation learning
R. Qian, T. Meng, B. Gong, M.-H. Yang, H. Wang, S. Belongie, and Y. Cui · 2020
Closest in time.
Self-supervised video representation learning by pace prediction
J. Wang, J. Jiao, and Y.-H. Liu · 2020
Closest in time.