Fetching the paper…
Reading the bibliography…
Understanding the training dynamics of deep neural networks is challenging due to their high-dimensional nature and intricate loss landscapes.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A Roberts, and Ethan Dyer · 2018
Earlier work this paper cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham Kakade · 2018
Earlier work this paper cites.
Lectures on convex optimization , volume 137
Yurii Nesterov et al · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Earlier work this paper cites.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Earlier work this paper cites.
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Earlier work this paper cites.
On the relation between the sharpest directions of DNN loss and the SGD step length
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amost Storkey · 2019
Earlier work this paper cites.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Vardan Papyan · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2020
Earlier work this paper cites.
Momentum improves normalized sgd
Ashok Cutkosky and Harsh Mehta · 2020
Earlier work this paper cites.
Improving neural network training in low dimensional random bases
Frithjof Gressmann, Zach Eaton-Rosen, and Carlo Luschi · 2020
Earlier work this paper cites.
The break-even point on optimization trajectories of deep neural networks
Stanisław Jastrzębski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho*, and Krzysztof Geras* · 2020
Earlier work this paper cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2020
Cited alongside, same era.
Traces of class/cross-class structure pervade deep learning spectra
Vardan Papyan · 2020
Cited alongside, same era.
Why are adaptive methods good for attention models?
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Cited alongside, same era.
Label noise SGD provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D. Lee · 2021
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2023
Later among the works it cites.
When and why momentum accelerates sgd: An empirical study
Jingwen Fu, Bohan Wang, Huishuai Zhang, Zhizheng Zhang, Wei Chen, and Nanning Zheng · 2023
Later among the works it cites.
Gradient descent monotonically decreases the sharpness of gradient flow solutions in scalar networks and beyond
Itai Kreisler, Mor Shpigel Nacson, Daniel Soudry, and Yair Carmon · 2023
Later among the works it cites.
Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2023
Later among the works it cites.
A new characterization of the edge of stability based on a sharpness measure aware of batch gradient distribution
Sungyoon Lee and Cheongjae Jang · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2021
Cited alongside, same era.
Privately learning subspaces
Vikrant Singhal and Thomas Steinke · 2021
Cited alongside, same era.
Bypassing the ambient dimension: Private {sgd} with gradient subspace identification
Yingxue Zhou, Steven Wu, and Arindam Banerjee · 2021
Cited alongside, same era.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Cited alongside, same era.
Towards understanding sharpness-aware minimization
Maksym Andriushchenko and Nicolas Flammarion · 2022
Cited alongside, same era.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Cited alongside, same era.
Later among the works it cites.
Practical sharpness-aware minimization cannot converge all the way to optima
Dongkuk Si and Chulhee Yun · 2023
Later among the works it cites.
Trajectory alignment: Understanding the edge of stability phenomenon via bifurcation theory
Minhak Song and Chulhee Yun · 2023
Later among the works it cites.
How sharpness-aware minimization minimizes sharpness?
Kaiyue Wen, Tengyu Ma, and Zhiyuan Li · 2023
Later among the works it cites.
Implicit bias of gradient descent for logistic regression at the edge of stability
Jingfeng Wu, Vladimir Braverman, and Jason D. Lee · 2023
Later among the works it cites.
Understanding edge-of-stability training dynamics with a minimalist example
Xingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou, and Rong Ge · 2023
Later among the works it cites.
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
Atish Agarwala and Jeffrey Pennington · 2024
Closest in time.
Adam with model exponential moving average is effective for nonconvex optimization
Kwangjun Ahn and Ashok Cutkosky · 2024
Closest in time.
High-dimensional SGD aligns with emerging outlier eigenspaces
Gerard Ben Arous, Reza Gheissari, Jiaoyang Huang, and Aukosh Jagannath · 2024
Closest in time.
Heavy-tailed class imbalance and why adam outperforms gradient descent on language models
Frederik Kunstner, Robin Yadav, Alan Milligan, Mark Schmidt, and Alberto Bietti · 2024
Closest in time.
Sharpness-aware minimization and the edge of stability
Philip M. Long and Peter L. Bartlett · 2024
Closest in time.
Identifying policy gradient subspaces
Jan Schneider, Pierre Schumacher, Simon Guist, Le Chen, Daniel Haeufle, Bernhard Schölkopf, and Dieter Büchler · 2024
Closest in time.
The marginal value of momentum for small learning rate SGD
Runzhe Wang, Sadhika Malladi, Tianhao Wang, Kaifeng Lyu, and Zhiyuan Li · 2024
Closest in time.
Implicit bias of adamw: ℓ ∞ \ell_{\infty} norm constrained optimization
Shuo Xie and Zhiyuan Li · 2024
Closest in time.
Compressible dynamics in deep overparameterized low-rank learning & adaptation
Can Yaras, Peng Wang, Laura Balzano, and Qing Qu · 2024
Closest in time.
Why transformers need adam: A hessian perspective
Yushun Zhang, Congliang Chen, Tian Ding, Ziniu Li, Ruoyu Sun, and Zhi-Quan Luo · 2024
Closest in time.
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
Libin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, and Mikhail Belkin · 2024
Closest in time.