Fetching the paper…
Reading the bibliography…
Existing analyses of neural network training often operate under the unrealistic assumption of an extremely small learning rate.
Sparse coding with an overcomplete basis set: a strategy employed by v1?
Bruno A Olshausen and David J Field · 1997
Earlier work this paper cites.
Sparse coding and decorrelation in primary visual cortex during natural vision
William E. Vinje and Jack L. Gallant · 2000
Earlier work this paper cites.
Sparse coding of sensory inputs
Bruno A Olshausen and David J Field · 2004
Earlier work this paper cites.
Linear spatial pyramid matching using sparse coding for image classification
Jianchao Yang, Kai Yu, Yihong Gong, and Thomas Huang · 2009
Earlier work this paper cites.
Simple, efficient, and neural algorithms for sparse coding
Sanjeev Arora, Rong Ge, Tengyu Ma, and Ankur Moitra · 2015
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Earlier work this paper cites.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Width of minima reached by stochastic gradient descent is influenced by learning rate to batch size ratio
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Earlier work this paper cites.
The comparative power of ReLU networks and polynomial kernels in the presence of sparse latent structure
Frederic Koehler and Andrej Risteski · 2018
Earlier work this paper cites.
How SGD selects the global minima in over-parameterized learning: a dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Earlier work this paper cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Earlier work this paper cites.
On the relation between the sharpest directions of DNN loss and the SGD step length
Stanisław Jastrzebski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amost Storkey · 2019
Earlier work this paper cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Cited alongside, same era.
The break-even point on optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Cited alongside, same era.
Learning rate annealing can provably help generalization, even for convex problems
Preetum Nakkiran · 2020
Cited alongside, same era.
Toward moderate overparameterization: global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Understanding the generalization benefit of normalization layers: sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Closest in time.
Beyond the quadratic approximation: the multiscale structure of neural network loss landscapes
Chao Ma, Daniel Kunin, Lei Wu, and Lexing Ying · 2022
Closest in time.
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry · 2022
Closest in time.
Convex analysis of the mean field Langevin dynamics
Atsushi Nitanda, Denny Wu, and Taiji Suzuki · 2022
Closest in time.
Trainability and accuracy of artificial neural networks: an interacting particle system approach
Grant M. Rotskoff and Eric Vanden-Eijnden · 2022
Closest in time.
Unveiling transformers with lego: a synthetic reasoning task
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Cited alongside, same era.
Catastrophic Fisher explosion: early phase Fisher matrix impacts generalization
Stanislaw Jastrzebski, Devansh Arpit, Oliver Astrand, Giancarlo B. Kerg, Huan Wang, Caiming Xiong, Richard Socher, Kyunghyun Cho, and Krzysztof J Geras · 2021
Cited alongside, same era.
Local signal adaptivity: provable feature learning in neural networks beyond kernels
Stefani Karp, Ezra Winston, Yuanzhi Li, and Aarti Singh · 2021
Cited alongside, same era.
The implicit bias of minima stability: a view from function space
Rotem Mulayoff, Tomer Michaeli, and Daniel Soudry · 2021
Cited alongside, same era.
Direction matters: on the implicit bias of stochastic gradient descent with moderate learning rate
Jingfeng Wu, Difan Zou, Vladimir Braverman, and Quanquan Gu · 2021
Cited alongside, same era.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Cited alongside, same era.
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Closest in time.
A mechanism for sample-efficient in-context learning for sparse retrieval tasks
Jacob Abernethy, Alekh Agarwal, Teodor V. Marinov, and Manfred K. Warmuth · 2023
Closest in time.
Transformers learn to implement preconditioned gradient descent for in-context learning
Kwangjun Ahn, Xiang Cheng, Hadi Daneshmand, and Suvrit Sra · 2023
Closest in time.
Physics of language models: part 1, context-free grammar
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Closest in time.
SGD with large step sizes learns sparse features
Maksym Andriushchenko, Aditya V. Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2023
Closest in time.
Beyond the edge of stability via two-step gradient updates
Lei Chen and Joan Bruna · 2023
Closest in time.
The crucial role of normalization in sharpness-aware minimization
Yan Dai, Kwangjun Ahn, and Suvrit Sra · 2023
Closest in time.
Self-stabilization: the implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2023
Closest in time.
How do transformers learn topic structure: towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Closest in time.
Trajectory alignment: understanding the edge of stability phenomenon via bifurcation theory
Minhak Song and Chulhee Yun · 2023
Closest in time.
Transformers learn in-context by gradient descent
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Closest in time.
Understanding edge-of-stability training dynamics with a minimalist example
Xingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou, and Rong Ge · 2023
Closest in time.