Fetching the paper…
Reading the bibliography…
Recent advances in self-supervised representation learning have enabled more efficient and robust model performance without relying on extensive labeled data.
Unsupervised deep learning by neighbourhood discovery
Jiabo Huang, Qi Dong, Shaogang Gong, and Xiatian Zhu. 2019 · 1904
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le. 2019 · 1905
Earlier work this paper cites.
Learning representations by maximizing mutual information across views
Philip Bachman, R. Devon Hjelm, and William Buchwalter. 2019 · 1906
Earlier work this paper cites.
Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2019 · 1906
Earlier work this paper cites.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman. 2019 · 1909
Earlier work this paper cites.
Self-supervised learning by cross-modal audio-video clustering
Humam Alwassel, Dhruv Mahajan, Lorenzo Torresani, Bernard Ghanem, and Du Tran. 2019 · 1911
Earlier work this paper cites.
Self-labelling via simultaneous clustering and representation learning
Yuki Markus Asano, Christian Rupprecht, and Andrea Vedaldi. 2019 · 1911
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick. 2019 · 1911
Earlier work this paper cites.
Learning representations by predicting bags of visual words
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, and Matthieu Cord. 2020 · 2002
Earlier work this paper cites.
Spatiotemporal relationship reasoning for pedestrian intent prediction
Bingbin Liu, Ehsan Adeli, Zhangjie Cao, Kuan-Hui Lee, Abhijeet Shenoi, Adrien Gaidon, and Juan Carlos Niebles. 2020 · 2002
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross B. Girshick, and Kaiming He. 2020b · 2003
Earlier work this paper cites.
Self-supervised spatio-temporal representation learning using variable playback speed prediction
Hyeon Cho, Taehoon Kim, Hyung Jin Chang, and Wonjun Hwang. 2020 · 2003
Earlier work this paper cites.
What makes for good views for contrastive learning
Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. 2020 · 2005
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. 2020 · 2006
Earlier work this paper cites.
Bootstrap your own latent: A new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Ávila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. 2020 · 2006
Earlier work this paper cites.
Video playback rate perception for self-supervisedspatio-temporal representation learning
Yuan Yao, Chang Liu, Dezhao Luo, Yu Zhou, and Qixiang Ye. 2020 · 2006
Earlier work this paper cites.
Representation learning with video deep infomax
R. Devon Hjelm and Philip Bachman. 2020 · 2007
Earlier work this paper cites.
Video representation learning by recognizing temporal transformations
Simon Jenni, Givi Meishvili, and Paolo Favaro. 2020 · 2007
Earlier work this paper cites.
Memory-augmented dense predictive coding for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman. 2020a · 2008
Earlier work this paper cites.
Spatiotemporal contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge J. Belongie, and Yin Cui. 2020 · 2008
Earlier work this paper cites.
Self-supervised video representation learning using inter-intra contrastive framework
Li Tao, Xueting Wang, and Toshihiko Yamasaki. 2020b · 2008
Earlier work this paper cites.
Self-supervised video representation learning by pace prediction
Jiangliu Wang, Jianbo Jiao, and Yun-Hui Liu. 2020 · 2008
Earlier work this paper cites.
Slow, decorrelated features for pretraining complex cell-like networks
Yoshua Bengio and James Bergstra. 2009 · 2009
Earlier work this paper cites.
Cifar-10 (canadian institute for advanced research)
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. 2009 · 2009
Earlier work this paper cites.
Deep learning from temporal coherence in video
Hossein Mobahi, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2020 · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Cited alongside, same era.
Self-supervised co-training for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman. 2020b · 2010
Cited alongside, same era.
Self-supervised video representation using pretext-contrastive learning
Li Tao, Xueting Wang, and T. Yamasaki. 2020a · 2010
Cited alongside, same era.
Hmdb: A large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. 2011 · 2011
Cited alongside, same era.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012 · 2012
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2017 · 2017
Later among the works it cites.
Unsupervised representation learning by sorting sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang. 2017 · 2017
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from multi-view observation
Pierre Sermanet, Corey Lynch, Jasmine Hsu, and Sergey Levine. 2017 · 2017
Later among the works it cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Deep clustering for unsupervised learning of visual features
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. 2014 · 2014
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei-Fei. 2014 · 2014
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2014 · 2014
Cited alongside, same era.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A. Efros. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Spatio-temporal video autoencoder with differentiable memory
Viorica Patraucean, Ankur Handa, and Roberto Cipolla. 2015 · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. 2018 · 2018
Later among the works it cites.
Unsupervised representation learning by predicting image rotations
Spyros Gidaris, Praveer Singh, and Nikos Komodakis. 2018 · 2018
Later among the works it cites.
Self-supervised spatiotemporal feature learning by video geometric transformations
Longlong Jing and Yingli Tian. 2018 · 2018
Later among the works it cites.
Self-supervised video representation learning with space-time cubic puzzles
Dahun Kim, Donghyeon Cho, and In So Kweon. 2018 · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Later among the works it cites.
Tracking emerges by colorizing videos
Carl Vondrick, Abhinav Shrivastava, Alireza Fathi, Sergio Guadarrama, and Kevin Murphy. 2018 · 2018
Later among the works it cites.
Unsupervised feature learning via non-parametric instance-level discrimination
Zhirong Wu, Yuanjun Xiong, Stella X. Yu, and Dahua Lin. 2018 · 2018
Later among the works it cites.
Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio. 2019 · 2019
Later among the works it cites.
Self-supervised spatiotemporal learning via video clip order prediction
Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, and Yueting Zhuang. 2019 · 2019
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. 2020a · 2020
Later among the works it cites.
Byol works even without batch statistics
Pierre H. Richemond, Jean-Bastien Grill, Florent Altché, Corentin Tallec, Florian Strub, Andrew Brock, Samuel Smith, Soham De, Razvan Pascanu, Bilal Piot, and Michal Valko. 2020 · 2020
Later among the works it cites.
Twins: Revisiting spatial attention design in vision transformers
Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. 2021 · 2021
Later among the works it cites.
Vector neurons: A general framework for so(3)-equivariant networks
Congyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard, Andrea Tagliasacchi, and Leonidas J. Guibas. 2021 · 2021
Later among the works it cites.
Levit: a vision transformer in convnet’s clothing for faster inference
Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze. 2021 · 2021
Later among the works it cites.
Rethinking spatial dimensions of vision transformers
Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh. 2021 · 2021
Later among the works it cites.
Going deeper with image transformers
Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. 2021 · 2021
Later among the works it cites.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis E. H. Tay, Jiashi Feng, and Shuicheng Yan. 2021 · 2021
Later among the works it cites.
Deepvit: Towards deeper vision transformer
Daquan Zhou, Bingyi Kang, Xiaojie Jin, Linjie Yang, Xiaochen Lian, Qibin Hou, and Jiashi Feng. 2021 · 2021
Later among the works it cites.
Sepvit: Separable vision transformer
Wei Li, Xing Wang, Xin Xia, Jie Wu, Xuefeng Xiao, Min Zheng, and Shiping Wen. 2022 · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022 · 2022
Later among the works it cites.
Learning multi-view visual correspondences with self-supervision
Pengcheng Zhang, Lei Zhou, Xiao Bai, Chen Wang, Jun Zhou, Liang Zhang, and Jin Zheng. 2022 · 2022
Later among the works it cites.
Multi-view action recognition using contrastive learning
Ketul Shah, Anshul Shah, Chun Pong Lau, Celso M. de Melo, and Rama Chellapp. 2023 · 2023
Closest in time.