Fetching the paper…
Reading the bibliography…
While the empirical success of self-supervised learning (SSL) heavily relies on the usage of deep nonlinear models, existing theoretical works on SSL understanding still focus on linear ones.
A theoretical analysis of contrastive unsupervised representation learning
Sanjeev Arora, Hrishikesh Khandeparkar, Mikhail Khodak, Orestis Plevrakis, and Nikunj Saunshi · 1902
Earlier work this paper cites.
Nonparametric discrimination: consistency properties
Evelyn Fix and Joseph Lawson Hodges · 1951
Earlier work this paper cites.
Rank-one modification of the symmetric eigenproblem
James R Bunch, Christopher P Nielsen, and Danny C Sorensen · 1978
Earlier work this paper cites.
Mathematical methods in the physical sciences, 1984
Mary L Boas and Philip Peters · 1984
Earlier work this paper cites.
Nonlinear dynamics and chaos
John Michael Tutill Thompson, H Bruce Stewart, and Rick Turner · 1990
Earlier work this paper cites.
On the strong universal consistency of nearest neighbor regression function estimates
Luc Devroye, Laszlo Gyorfi, Adam Krzyzak, and Gábor Lugosi · 1994
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Training dynamics and neural network performance
Charles L Wilson, James L Blue, and Omid M Omidvar · 1997
Earlier work this paper cites.
Finite-time blow-up in dynamical systems
Alain Goriely and Craig Hyde · 1998
Earlier work this paper cites.
Introduction to perturbation theory in quantum mechanics
Francisco M Fernández · 2000
Earlier work this paper cites.
A note on the universal approximation capability of support vector machines
Barbara Hammer and Kai Gersmann · 2003
Earlier work this paper cites.
Collected matrix derivative results for forward and reverse mode algorithmic differentiation
Mike B Giles · 2008
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Li Deng · 2012
Earlier work this paper cites.
Matrix computations
Gene H Golub and Charles F Van Loan · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
An introduction to the theory of reproducing kernel Hilbert spaces , volume 152
Vern I Paulsen and Mrinal Raghupathi · 2016
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Zhou Lu, Hongming Pu, Feicheng Wang, Zhiqiang Hu, and Liwei Wang · 2017
Cited alongside, same era.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Bootstrap your own latent: A new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al · 2020
Later among the works it cites.
Expressivity of deep neural networks
Ingo Gühring, Mones Raslan, and Gitta Kutyniok · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Understanding self-supervised learning with dual deep networks
Yuandong Tian, Lantao Yu, Xinlei Chen, and Surya Ganguli · 2020
Later among the works it cites.
Playing the lottery with rewards and multiple languages: lottery tickets in rl and nlp
Haonan Yu, Sergey Edunov, Yuandong Tian, and Ari S. Morcos · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Simon S Du, Wei Hu, and Jason D Lee · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Cited alongside, same era.
On the impact of the activation function on deep neural networks training
Soufiane Hayou, Arnaud Doucet, and Judith Rousseau · 2019
Cited alongside, same era.
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
Ari Morcos, Haonan Yu, Michela Paganini, and Yuandong Tian · 2019
Cited alongside, same era.
Luck matters: Understanding training dynamics of deep relu networks
Yuandong Tian, Tina Jiang, Qucheng Gong, and Ari Morcos · 2019
Cited alongside, same era.
Later among the works it cites.
The lottery tickets hypothesis for supervised and self-supervised pre-training in computer vision models
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Michael Carbin, and Zhangyang Wang · 2021
Later among the works it cites.
Benyamin Ghojogh, Ali Ghodsi, Fakhri Karray, and Mark Crowley · 2021
Later among the works it cites.
Provable guarantees for self-supervised deep learning with spectral contrastive loss
Jeff Z HaoChen, Colin Wei, Adrien Gaidon, and Tengyu Ma · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Later among the works it cites.
The power of contrast for feature learning: A theoretical analysis
Wenlong Ji, Zhun Deng, Ryumei Nakada, James Zou, and Linjun Zhang · 2021
Later among the works it cites.
The principles of deep learning theory
Daniel A. Roberts, Sho Yaida, and Boris Hanin · 2021
Later among the works it cites.
Understanding self-supervised learning dynamics without contrastive pairs
Yuandong Tian, Xinlei Chen, and Surya Ganguli · 2021
Later among the works it cites.
Towards demystifying representation learning with non-contrastive self-supervision
Xiang Wang, Xinlei Chen, Simon S Du, and Yuandong Tian · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel · 2022
Closest in time.
Understanding dimensional collapse in contrastive self-supervised learning
Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian · 2022
Closest in time.
Understanding contrastive learning requires incorporating inductive biases
Nikunj Saunshi, Jordan Ash, Surbhi Goel, Dipendra Misra, Cyril Zhang, Sanjeev Arora, Sham Kakade, and Akshay Krishnamurthy · 2022
Closest in time.
Understanding deep contrastive learning via coordinate-wise optimization
Yuandong Tian · 2022
Closest in time.