Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) has emerged as a powerful framework to learn representations from raw data without supervision.
On Hilbert’s thirteenth problem
Vitushkin, A · 1954
Earlier work this paper cites.
ε \varepsilon -entropy and ε \varepsilon -capacity of sets in functional spaces
Kolmogorov, A. and Tikhomirov, V · 1959
Earlier work this paper cites.
The rotation of eigenvectors by a perturbation
Davis, C. and Kahan, W · 1970
Earlier work this paper cites.
Remarks on inequalities for large deviation probabilities
Pinelis, I. and Sakhanenko, A · 1986
Earlier work this paper cites.
Tangent prop - a formalism for specifying selected invariances in an adaptive network
Simard, P., Victorri, B., LeCun, Y., and Denker, J · 1991
Earlier work this paper cites.
Perturbation Theory for Linear Operators
Kato, T · 1995
Earlier work this paper cites.
Regularization with dot-product kernels
Smola, A., Ovári, Z., and Williamson, R. C · 2000
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Scholkopf, B. and Smola, A · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. and Mendelson, S · 2002
Earlier work this paper cites.
Generalization error bounds for bayesian mixture algorithms
Meir, R. and Zhang, T · 2003
Earlier work this paper cites.
Convexity, classification, and risk bounds
Bartlett, P., Jordan, M., and Mcauliffe, J · 2006
Earlier work this paper cites.
Diffusion maps
Coifman, R. and Lafon, S · 2006
Earlier work this paper cites.
Universal kernels
Micchelli, C., Xu, Y., and Zhang, H · 2006
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Caponnetto, A. and De Vito, E · 2007
Earlier work this paper cites.
Generalization error bounds in semi-supervised classification under the cluster assumption
Rigollet, P · 2007
Earlier work this paper cites.
Learning theory estimates via integral operators and their approximations
Smale, S. and Zhou, D.-X · 2007
Earlier work this paper cites.
Reproducing kernel hilbert spaces associated with analytic translation-invariant mercer kernels
Sun, H.-W. and Zhou, D.-X · 2008
Earlier work this paper cites.
Spherical harmonics in p dimensions
Efthimiou, C. and Frye, C · 2014
Earlier work this paper cites.
Analysis of boolean functions
O’Donnell, R · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Earlier work this paper cites.
The geometry of kernelized spectral clustering
Schiebinger, G., Wainwright, M., and Yu, B · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Tropp, J · 2015
Earlier work this paper cites.
A vector-contraction inequality for Rademacher complexities
Maurer, A · 2016
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Bach, F · 2017
Cited alongside, same era.
Cleaning large correlation matrices: Tools from random matrix theory
Bun, J., Bouchaud, J.-P., and Potters, M · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Cited alongside, same era.
Finite-sample analysis of m-estimators using self-concordance
Ostrovskii, D. and Bach, F · 2018
Cited alongside, same era.
A theoretical analysis of contrastive unsupervised representation learning
Arora, S., Khandeparkar, H., Khodak, M., Plevrakis, O., and Saunshi, N · 2019
Cited alongside, same era.
Provable guarantees for self-supervised deep learning with spectral contrastive loss
HaoChen, J., Wei, C., Gaidon, A., and Ma, T · 2021
Later among the works it cites.
Predicting what you already know helps: Provable self-supervised learning
Lee, J., Lei, Q., Saunshi, N., and Zhuo, J · 2021
Later among the works it cites.
Learning with invariances in random features and kernel models
Mei, S., Misiakiewicz, T., and Montanari, A · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Understanding self-supervised learning with dual deep networks, 2021
Tian, Y., Yu, L., Chen, X., and Ganguli, S · 2021
Later among the works it cites.
Toward understanding the feature learning process of self-supervised contrastive learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the inductive bias of neural tangent kernels
Bietti, A. and Mairal, J · 2019
Cited alongside, same era.
Nonconvex optimization meets low-rank matrix factorization: An overview
Chi, Y., Lu, Y., and Chen, Y · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J · 2019
Cited alongside, same era.
A fine-grained spectral perspective on neural networks
Yang, G. and Salman, H · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Wen, Z. and Li, Y · 2021
Later among the works it cites.
Contrastive and non-contrastive self-supervised learning recover global and local spectral embedding methods
Balestriero, R. and LeCun, Y · 2022
Later among the works it cites.
Vicreg: Variance-invariance-covariance regularization for self-supervised learning
Bardes, A., Ponce, J., and LeCun, Y · 2022
Later among the works it cites.
Approximation and learning with deep convolutional models: a kernel perspective
Bietti, A · 2022
Later among the works it cites.
Garrido, Q., Balestriero, R., Najman, L., and Lecun, Y · 2022
Later among the works it cites.
A theoretical study of inductive biases in contrastive learning
HaoChen, J. Z. and Ma, T · 2022
Later among the works it cites.
Joint embedding self-supervised learning in the kernel regime
Kiani, B. T., Balestriero, R., Chen, Y., Lloyd, S., and LeCun, Y · 2022
Later among the works it cites.
Learning with convolution and pooling operations in kernel methods
Misiakiewicz, T. and Mei, S · 2022
Later among the works it cites.
An elementary analysis of ridge regression with random design
Mourtada, J. and Rosasco, L · 2022
Later among the works it cites.
Distribution-free robust linear regression
Mourtada, J., skevičius, T. V., and Zhivotovskiy, N · 2022
Later among the works it cites.
Understanding contrastive learning requires incorporating inductive biases
Saunshi, N., Ash, J., Goel, S., Misra, D., Zhang, C., Arora, S., Kakade, S., and Krishnamurthy, A · 2022
Later among the works it cites.
Understanding the role of nonlinearity in training dynamics of contrastive learning
Tian, Y · 2022
Later among the works it cites.
Learning Theory from First Principles
Bach, F · 2023
Closest in time.
On minimal variations for unsupervised representation learning
Cabannes, V., Bietti, A., and Balestriero, R · 2023
Closest in time.
Kernelized diffusion maps
Pillaud-Vivien, L. and Bach, F · 2023
Closest in time.
On the stepwise nature of self-supervised learning
Simon, J., Knutins, M., Ziyin, L., Geisz, D., Fetterman, A., and Albrecht, J · 2023
Closest in time.