Fetching the paper…
Reading the bibliography…
Recently the surprising discovery of the Bootstrap Your Own Latent (BYOL) method by Grill et al.
A theoretical analysis of contrastive unsupervised representation learning
Sanjeev Arora, Hrishikesh Khandeparkar, Mikhail Khodak, Orestis Plevrakis, and Nikunj Saunshi · 1902
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2001
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2012
Earlier work this paper cites.
Learning polynomials with neural networks
Alexandr Andoni, Rina Panigrahy, Gregory Valiant, and Li Zhang · 2014
Earlier work this paper cites.
Globally optimal training of generalized polynomial neural networks with nonlinear spectral methods
Antoine Gautier, Quynh N Nguyen, and Matthias Hein · 2016
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Earlier work this paper cites.
When is a convolutional filter easy to learn?
Simon S Du, Jason D Lee, and Yuandong Tian · 2018
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Earlier work this paper cites.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Simon S. Du, Jason D. Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Earlier work this paper cites.
Limitations of lazy training of two-layers neural network
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio · 2019
Earlier work this paper cites.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
On the expressive power of deep polynomial neural networks
Joe Kileel, Matthew Trager, and Joan Bruna · 2019
Earlier work this paper cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Earlier work this paper cites.
For self-supervised learning, rationality implies generalization, provably
Yamini Bansal, Gal Kaplun, and Boaz Barak · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Cited alongside, same era.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Cited alongside, same era.
Towards the generalization of contrastive self-supervised learning
Weiran Huang, Mingyang Yi, and Xuyang Zhao · 2021
Later among the works it cites.
The power of contrast for feature learning: A theoretical analysis
Wenlong Ji, Zhun Deng, Ryumei Nakada, James Zou, and Linjun Zhang · 2021
Later among the works it cites.
Understanding dimensional collapse in contrastive self-supervised learning
Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian · 2021
Later among the works it cites.
Local signal adaptivity: Provable feature learning in neural networks beyond kernels
Stefani Karp, Ezra Winston, Yuanzhi Li, and Aarti Singh · 2021
Later among the works it cites.
Predicting what you already know helps: Provable self-supervised learning
Jason D Lee, Qi Lei, Nikunj Saunshi, and Jiacheng Zhuo · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-supervised adversarial robustness for the low-label, high-data regime
Sven Gowal, Po-Sen Huang, Aaron van den Oord, Timothy Mann, and Pushmeet Kohli · 2020
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Cited alongside, same era.
Making method of moments great again?–how can gans learn distributions
Yuanzhi Li and Zehao Dou · 2020
Cited alongside, same era.
Learning over-parametrized two-layer relu neural networks beyond ntk
Yuanzhi Li, Tengyu Ma, and Hongyang R. Zhang · 2020
Cited alongside, same era.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Cited alongside, same era.
BYOL works even without batch statistics
Pierre H. Richemond, Jean-Bastien Grill, Florent Altché, Corentin Tallec, Florian Strub, Andrew Brock, Samuel Smith, Soham De, Razvan Pascanu, Bilal Piot, and Michal Valko · 2020
Cited alongside, same era.
Later among the works it cites.
Analyzing and improving the optimization landscape of noise-contrastive estimation
Bingbin Liu, Elan Rosenfeld, Pradeep Ravikumar, and Andrej Risteski · 2021
Later among the works it cites.
Byol for audio: Self-supervised learning for general-purpose audio representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, and Kunio Kashino · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Can contrastive learning avoid shortcut solutions?
Joshua Robinson, Li Sun, Ke Yu, Kayhan Batmanghelich, Stefanie Jegelka, and Suvrit Sra · 2021
Later among the works it cites.
Can pretext-based self-supervised learning be boosted by downstream data? a theoretical analysis
Jiaye Teng, Weiran Huang, and Haowei He · 2021
Later among the works it cites.
Understanding self-supervised learning dynamics without contrastive pairs
Yuandong Tian, Xinlei Chen, and Surya Ganguli · 2021
Later among the works it cites.
Contrastive learning, multi-view redundancy, and linear models
Christopher Tosh, Akshay Krishnamurthy, and Daniel Hsu · 2021
Later among the works it cites.
Self-supervised learning with data augmentations provably isolates content from style
Julius Von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel, Bernhard Schölkopf, Michel Besserve, and Francesco Locatello · 2021
Later among the works it cites.
Towards demystifying representation learning with non-contrastive self-supervision
Xiang Wang, Xinlei Chen, Simon S Du, and Yuandong Tian · 2021
Later among the works it cites.
Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning
Colin Wei, Sang Michael Xie, and Tengyu Ma · 2021
Later among the works it cites.
Toward understanding the feature learning process of self-supervised contrastive learning
Zixin Wen and Yuanzhi Li · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny · 2021
Later among the works it cites.
Understanding the generalization of adam in learning neural networks with proper regularization
Difan Zou, Yuan Cao, Yuanzhi Li, and Quanquan Gu · 2021
Later among the works it cites.
Learning polynomial transformations
Sitan Chen, Jerry Li, Yuanzhi Li, and Anru R Zhang · 2022
Closest in time.
Towards understanding how momentum improves generalization in deep learning, 2022
Samy Jelassi and Yuanzhi Li · 2022
Closest in time.
Adam is no better than normalized SGD: Dissecting how adaptivity improves GAN performance, 2022
Samy Jelassi, Arthur Mensch, Gauthier Gidel, and Yuanzhi Li · 2022
Closest in time.
Masked prediction tasks: a parameter identifiability view
Bingbin Liu, Daniel Hsu, Pradeep Ravikumar, and Andrej Risteski · 2022
Closest in time.
One objective for all models–self-supervised learning for topic models
Zeping Luo, Cindy Weng, Shiyou Wu, Mo Zhou, and Rong Ge · 2022
Closest in time.
Contrasting the landscape of contrastive and non-contrastive learning
Ashwini Pokle, Jinjin Tian, Yuchen Li, and Andrej Risteski · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
Understanding contrastive learning requires incorporating inductive biases
Nikunj Saunshi, Jordan Ash, Surbhi Goel, Dipendra Misra, Cyril Zhang, Sanjeev Arora, Sham Kakade, and Akshay Krishnamurthy · 2022
Closest in time.
Chaoning Zhang, Kang Zhang, Chenshuang Zhang, Trung X Pham, Chang D Yoo, and In So Kweon · 2022
Closest in time.