Fetching the paper…
Reading the bibliography…
The effect of regularizers such as weight decay when training deep neural networks is not well understood.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 1905
Earlier work this paper cites.
Sur certains processus stochastiques homogènes
Paul Lévy · 1940
Earlier work this paper cites.
Free states of the canonical anticommutation relations
Robert T Powers and Erling Størmer · 1970
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A. Hertz · 1991
Earlier work this paper cites.
Probable networks and plausible predictions — a review of practical bayesian methods for supervised neural networks
David J C Mackay · 1995
Earlier work this paper cites.
Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit
Richard H. R. Hahnloser, Rahul Sarpeshkar, Misha A. Mahowald, Rodney J. Douglas, and H. Sebastian Seung · 2000
Earlier work this paper cites.
Rank, trace-norm and max-norm
Nathan Srebro and Adi Shraibman · 2005
Earlier work this paper cites.
The power of convex relaxation: Near-optimal matrix completion, 2009
Emmanuel J. Candes and Terence Tao · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, 2013
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Adam: a method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Guaranteed matrix completion via non-convex factorization
Ruoyu Sun and Zhi-Quan Luo · 2016
Earlier work this paper cites.
L2 regularization versus batch and weight normalization, 2017
Twan van Laarhoven · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization, 2017
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2017
Earlier work this paper cites.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Three mechanisms of weight decay regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Implicit regularization in deep learning may not be explainable by norms, 2020
Noam Razin and Nadav Cohen · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Visual transformers: Token-based image representation and processing for computer vision, 2020
Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Zhicheng Yan, Masayoshi Tomizuka, Joseph Gonzalez, Kurt Keutzer, and Peter Vajda · 2020
Cited alongside, same era.
Low rank regularization: A review
Zhanxuan Hu, Feiping Nie, Rong Wang, and Xuelong Li · 2020
Initialization and regularization of factorized neural layers, 2022
Mikhail Khodak, Neil Tenenholtz, Lester Mackey, and Nicolò Fusi · 2022
Later among the works it cites.
Exact solutions of a deep linear network
Liu Ziyin, Botao Li, and Xiangming Meng · 2022
Later among the works it cites.
Training a vision transformer from scratch in less than 24 hours with 1 gpu, 2022
Saghar Irandoust, Thibaut Durand, Yunduz Rakhmangulova, Wenjie Zi, and Hossein Hajimirsadeghi · 2022
Later among the works it cites.
On the overlooked pitfalls of weight decay and how to mitigate them: A gradient-norm perspective
Zeke Xie, zhiqiang xu, Jingzhao Zhang, Issei Sato, and Masashi Sugiyama · 2023
Later among the works it cites.
Why do we need weight decay in modern deep learning?, 2023
Maksym Andriushchenko, Francesco D’Angelo, Aditya Varre, and Nicolas Flammarion · 2023
Later among the works it cites.
spred: Solving l1 penalty with SGD
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Low-rank bottleneck in multi-head attention models, 2020
Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi, and Sanjiv Kumar · 2020
Cited alongside, same era.
The pile: an 800GB dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Cited alongside, same era.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2021
Cited alongside, same era.
Equivalences between sparse models and neural networks
Ryan J Tibshirani · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
A geometric analysis of neural collapse with unconstrained features, 2021
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
Cited alongside, same era.
Liu Ziyin and Zihao Wang · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Implicit bias of large depth networks: a notion of rank for nonlinear functions, 2023
Arthur Jacot · 2023
Later among the works it cites.
Characterizing the implicit bias of regularized sgd in rank minimization, 2023
Tomer Galanti, Zachary S. Siegel, Aparna Gupte, and Tomaso Poggio · 2023
Later among the works it cites.
Implicit bias of sgd in l 2 l_{2} -regularized linear dnns: One-way jumps from high to low rank, 2023
Zihan Wang and Arthur Jacot · 2023
Later among the works it cites.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction, 2023
Pratyusha Sharma, Jordan T. Ash, and Dipendra Misra · 2023
Later among the works it cites.
Hungry hungry hippos: Towards language modeling with state space models, 2023
Daniel Y. Fu, Tri Dao, Khaled K. Saab, Armin W. Thomas, Atri Rudra, and Christopher Ré · 2023
Later among the works it cites.
Hyena hierarchy: Towards larger convolutional language models, 2023
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré · 2023
Later among the works it cites.
Uncovering mesa-optimization algorithms in transformers, 2023
Johannes von Oswald, Eyvind Niklasson, Maximilian Schlegel, Seijin Kobayashi, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Blaise Agüera y Arcas, Max Vladymyrov, Razvan Pascanu, and João Sacramento · 2023
Later among the works it cites.
The transient nature of emergent in-context learning in transformers, 2023
Aaditya K. Singh, Stephanie C. Y. Chan, Ted Moskovitz, Erin Grant, Andrew M. Saxe, and Felix Hill · 2023
Later among the works it cites.