Fetching the paper…
Reading the bibliography…
We identify incremental learning dynamics in transformers, where the difference between trained and initial weights progressively increases in rank.
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., Language models are few-shot learners , Advances in neural information processing systems 33
1901
Earlier work this paper cites.
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2017
Earlier work this paper cites.
Sanjeev Arora, Nadav Cohen, and Elad Hazan, On the optimization of deep networks: Implicit acceleration by overparameterization , International Conference on Machine Learning, PMLR, 2018, pp. 244–253
2018
Earlier work this paper cites.
Lénaïc Chizat, Edouard Oyallon, and Francis R. Bach, On lazy training in differentiable programming , Neural Information Processing Systems, 2018
2018
Earlier work this paper cites.
Simon S Du, Wei Hu, and Jason D Lee, Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced , Advances in Neural Information Processing Systems 31
2018
Earlier work this paper cites.
Arthur Jacot, Franck Gabriel, and Clément Hongler, Neural tangent kernel: Convergence and generalization in neural networks , Advances in neural information processing systems 31
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo, Implicit regularization in deep matrix factorization , Advances in Neural Information Processing Systems 32
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien, Implicit regularization of discrete gradient dynamics in linear neural networks , Advances in Neural Information Processing Systems 32
2019
Earlier work this paper cites.
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Limitations of lazy training of two-layers neural network , Advances in Neural Information Processing Systems 32
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
Francis Bach, Effortless optimization through gradient flows , Machine Learning Research Blog. https://francisbach. com/gradient-flows (2020)
2020
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi, Neural networks can learn representations with gradient descent , Conference on Learning Theory, PMLR, 2022, pp. 5413–5452
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Noam Razin and Nadav Cohen, Implicit regularization in deep learning may not be explainable by norms , Advances in neural information processing systems 33
2020
Cited alongside, same era.
2021
Cited alongside, same era.
Arthur Jacot, Francois Gaston Ged, Berfin Simsek, Clément Hongler, and Franck Gabriel, Saddle-to-saddle dynamics in deep linear networks: Small initialization training, symmetry, and sparsity , 2021
2021
Cited alongside, same era.
Paolo Milanesi, Hachem Kadri, S. Ayache, and Thierry Artières, Implicit regularization in deep tensor factorization , 2021 International Joint Conference on Neural Networks (IJCNN) (2021), 1–8
2021
Cited alongside, same era.
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro, Quantifying the benefit of using differentiable learning over tangent kernels , Proceedings of the 38th International Conference on Machine Learning (Marina Meila and Tong Zhang, eds.), Proceedings of Machine Learning Research, vol. 139, PMLR, 18–24 Jul 2021, pp. 7379–7389
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Closest in time.
Dylan Patel and Afzal Ahmad, Google “we have no moat, and neither does openai” , May 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto, Alpaca: A strong, replicable instruction-following model , Stanford Center for Research on Foundation Models. https://crfm. stanford. edu/2023/03/13/alpaca. html (2023)
2023
Closest in time.
Nadav Timor, Gal Vardi, and Ohad Shamir, Implicit regularization towards rank minimization in relu networks , International Conference on Algorithmic Learning Theory, PMLR, 2023, pp. 1429–1459
2023
Closest in time.