Fetching the paper…
Reading the bibliography…
Normalization layers (e.g., Batch Normalization, Layer Normalization) were introduced to help with optimization difficulties in very deep nets, but they clearly also help generalization, even in not-so-deep nets.
Positively scale-invariant flatness of ReLU neural networks
Mingyang Yi, Qi Meng, Wei Chen, Zhi-ming Ma, and Tie-Yan Liu · 1903
Earlier work this paper cites.
Differentiation of the Limit Mapping in a Dynamical System
K. J. Falconer · 1983
Earlier work this paper cites.
Shorter notes: Regularity of the distance function
Robert L. Foote · 1984
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Simplified PAC-Bayesian margin bounds
David McAllester · 2003
Earlier work this paper cites.
Neural networks for machine learning lecture 6a: Overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, J. P. Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Centered weight normalization in accelerating training of deep neural networks
Lei Huang, Xianglong Liu, Yang Liu, Bo Lang, and Dacheng Tao · 2017
Earlier work this paper cites.
Three factors influencing minima in SGD
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
First-order methods almost always avoid saddle points
Jason D Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Earlier work this paper cites.
Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions
Ioannis Panageas and Georgios Piliouras · 2017
Earlier work this paper cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization
Twan van Laarhoven · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Understanding batch normalization
Johan Bjorck, Carla Gomes, and Bart Selman · 2018
Earlier work this paper cites.
Norm matters: efficient and accurate normalization schemes in deep networks
Elad Hoffer, Ron Banner, Itay Golan, and Daniel Soudry · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
An alternative view: When does SGD escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Earlier work this paper cites.
Bayesian uncertainty estimation for batch normalized deep networks
Mattias Teye, Hossein Azizpour, and Kevin Smith · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Earlier work this paper cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Cited alongside, same era.
A quantitative analysis of the effect of batch normalization on gradient descent
Yongqiang Cai, Qianxiao Li, and Zuowei Shen · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Cited alongside, same era.
Nonconvex optimization meets low-rank matrix factorization: An overview
Yuejie Chi, Yue M. Lu, and Yuxin Chen · 2019
Cited alongside, same era.
Online normalization for training neural networks
Vitaliy Chiley, Ilya Sharapov, Atli Kosson, Urs Koster, Ryan Reece, Sofia Samaniego de la Fuente, Vishal Subbiah, and Michael James · 2019
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Later among the works it cites.
Normalized flat minima: Exploring scale invariant definition of flat minima for neural networks using PAC-Bayesian analysis
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama · 2020
Later among the works it cites.
Implicit regularization and convergence for weight normalization
Xiaoxia Wu, Edgar Dobriban, Tongzheng Ren, Shanshan Wu, Zhiyuan Li, Suriya Gunasekar, Rachel Ward, and Qiang Liu · 2020
Later among the works it cites.
Implicit gradient regularization
David Barrett and Benoit Dherin · 2021
Later among the works it cites.
Characterizing signal propagation to close the performance gap in unnormalized resnets
Andrew Brock, Soham De, and Samuel L Smith · 2021
Later among the works it cites.
How much over-parameterization is sufficient to learn deep ReLU networks?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Cited alongside, same era.
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Cited alongside, same era.
The normalization method for alleviating pathological sharpness in wide neural networks
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2019
Cited alongside, same era.
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2021
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
Label noise SGD provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
Later among the works it cites.
Batch normalization orthogonalizes representations in deep random networks
Hadi Daneshmand, Amir Joudaki, and Francis Bach · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Later among the works it cites.
Understanding deflation process in over-parametrized tensor decomposition
Rong Ge, Yunwei Ren, Xiang Wang, and Mo Zhou · 2021
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z. HaoChen, Colin Wei, Jason Lee, and Tengyu Ma · 2021
Later among the works it cites.
Exponential escape efficiency of SGD from sharp minima in non-stationary regime
Hikaru Ibayashi and Masaaki Imaizumi · 2021
Later among the works it cites.
Proxy-normalizing activations to match batch normalization while removing batch dependence
Antoine Labatie, Dominic Masters, Zach Eaton-Rosen, and Carlo Luschi · 2021
Later among the works it cites.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2021
Later among the works it cites.
Why spectral normalization stabilizes GANs: Analysis and improvements
Zinan Lin, Vyas Sekar, and Giulia Fanti · 2021
Later among the works it cites.
On the periodic behavior of neural network training with batch normalization and weight decay
Ekaterina Lobacheva, Maxim Kodryan, Nadezhda Chirkova, Andrey Malinin, and Dmitry P Vetrov · 2021
Later among the works it cites.
Beyond batchnorm: Towards a unified understanding of normalization in deep learning
Ekdeep S Lubana, Robert Dick, and Hidenori Tanaka · 2021
Later among the works it cites.
Gradient descent on two-layer nets: Margin maximization and simplicity bias
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang, and Sanjeev Arora · 2021
Later among the works it cites.
On linear stability of SGD and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Rotem Mulayoff, Tomer Michaeli, and Daniel Soudry · 2021
Later among the works it cites.
A scale invariant measure of flatness for deep network minima
Akshay Rangamani, Nam H. Nguyen, Abhishek Kumar, Dzung Phan, Sang Peter Chin, and Trac D. Tran · 2021
Later among the works it cites.
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
Dominik Stöger and Mahdi Soltanolkotabi · 2021
Later among the works it cites.
Noether’s learning dynamics: Role of symmetry breaking in neural networks
Hidenori Tanaka and Daniel Kunin · 2021
Later among the works it cites.
Spherical motion dynamics: Learning dynamics of normalized neural network using sgd and weight decay
Ruosi Wan, Zhanxing Zhu, Xiangyu Zhang, and Jian Sun · 2021
Later among the works it cites.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2021
Later among the works it cites.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Closest in time.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Closest in time.
On gradient descent convergence beyond the edge of stability
Lei Chen and Joan Bruna · 2022
Closest in time.
Flat minima generalize for low-rank matrix recovery
Lijun Ding, Dmitriy Drusvyatskiy, and Maryam Fazel · 2022
Closest in time.
A loss curvature perspective on training instabilities of deep learning models
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Edward Dahl, Zachary Nado, and Orhan Firat · 2022
Closest in time.
Batch normalization preconditioning for neural network training
Susanna Lange, Kyle Helfrich, and Qiang Ye · 2022
Closest in time.
A Riemannian mean field formulation for two-layer neural networks with batch normalization
Chao Ma and Lexing Ying · 2022
Closest in time.
Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks
Noam Razin, Asaf Maman, and Nadav Cohen · 2022
Closest in time.
Three-stage evolution and fast equilibrium for SGD with non-degerate critical points
Yi Wang and Zhiren Wang · 2022
Closest in time.
Large learning rate tames homogeneity: Convergence and balancing effect
Yuqing Wang, Minshuo Chen, Tuo Zhao, and Molei Tao · 2022
Closest in time.
Beyond the quadratic approximation: The multiscale structure of neural network loss landscapes
Chao Ma, Daniel Kunin, Lei Wu, and Lexing Ying · 2048
Closest in time.