Fetching the paper…
Reading the bibliography…
Normalization methods such as batch [Ioffe and Szegedy, 2015], weight [Salimansand Kingma, 2016], instance [Ulyanov et al., 2016], and layer normalization [Baet al., 2016] have been widely used in modern machine learning.
Theory and methods related to the singular-function expansion and landweber’s iteration for integral equations of the first kind
Otto Neall Strand · 1974
Earlier work this paper cites.
Generalization and parameter estimation in feedforward nets: Some experiments
Nelson Morgan and Hervé Bourlard · 1990
Earlier work this paper cites.
On gradient adaptation with unit-norm constraints
Scott C Douglas, Shun-ichi Amari, and S-Y Kung · 2000
Earlier work this paper cites.
Digital signal processing: principles algorithms and applications
John G Proakis · 2001
Earlier work this paper cites.
Least-mean-square adaptive filters
Simon Saher Haykin and Bernard Widrow · 2002
Earlier work this paper cites.
Adaptive filter theory
Simon S Haykin · 2005
Earlier work this paper cites.
Exact matrix completion via convex optimization
Emmanuel J Candès and Benjamin Recht · 2009
Earlier work this paper cites.
Statistical digital signal processing and modeling
Monson H Hayes · 2009
Earlier work this paper cites.
Variational analysis , volume 317
R Tyrrell Rockafellar and Roger J-B Wets · 2009
Earlier work this paper cites.
A randomized kaczmarz algorithm with exponential convergence
Thomas Strohmer and Roman Vershynin · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Optimization and dynamical systems
Uwe Helmke and John B Moore · 2012
Earlier work this paper cites.
Approximate computation and implicit regularization for very large-scale data analysis
Michael W Mahoney · 2012
Earlier work this paper cites.
The phase transition of matrix recovery from gaussian measurements matches the minimax mse of matrix denoising
David L Donoho, Matan Gavish, and Andrea Montanari · 2013
Earlier work this paper cites.
Dropout training as adaptive regularization
Stefan Wager, Sida Wang, and Percy S Liang · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Gradient descent converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P Kingma · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Closest in time.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2019
Closest in time.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2019
Closest in time.
A quantitative analysis of the effect of batch normalization on gradient descent
Yongqiang Cai, Qianxiao Li, and Zuowei Shen · 2019
Closest in time.
Invariance reduces variance: Understanding data augmentation in deep learning and beyond
Shuxiao Chen, Edgar Dobriban, and Jane H Lee · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J Zico Kolter, and Ryan J Tibshirani · 2018
Cited alongside, same era.
Theoretical analysis of auto rate-tuning by batch normalization
Sanjeev Arora, Zhiyuan Li, and Kaifeng Lyu · 2018
Cited alongside, same era.
High-dimensional asymptotics of prediction: Ridge regression and classification
Edgar Dobriban and Stefan Wager · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Closest in time.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Closest in time.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Closest in time.
Exponential convergence rates for batch normalization: The power of length-direction decoupling in non-convex optimization
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, Thomas Hofmann, Ming Zhou, and Klaus Neymeyr · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
Revisit batch normalization: New understanding and refinement via composition optimization
Xiangru Lian and Ji Liu · 2019
Closest in time.
Ridge regression: Structure, cross-validation, and sketching
Sifan Liu and Edgar Dobriban · 2019
Closest in time.
Towards understanding regularization in batch normalization
Ping Luo, Xinjiang Wang, Wenqi Shao, and Zhanglin Peng · 2019
Closest in time.
Traditional and heavy tailed self regularization in neural network models
Michael Mahoney and Charles Martin · 2019
Closest in time.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Closest in time.
Theoretical issues in deep networks: Approximation, optimization and generalization
Tomaso Poggio, Andrzej Banburski, and Qianli Liao · 2019
Closest in time.
Over-parameterization as a catalyst for better generalization of deep relu network
Yuandong Tian · 2019
Closest in time.
Luck matters: Understanding training dynamics of deep relu networks
Yuandong Tian, Tina Jiang, Qucheng Gong, and Ari Morcos · 2019
Closest in time.
Implicit regularization for optimal sparse recovery
Tomas Vaškevičius, Varun Kanade, and Patrick Rebeschini · 2019
Closest in time.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Closest in time.
Understanding and improving layer normalization
Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin · 2019
Closest in time.
Dropout: Explicit forms and capacity control
Raman Arora, Peter Bartlett, Poorya Mianjy, and Nathan Srebro · 2020
Closest in time.
Optimization theory for relu neural networks trained with normalization layers
Yonatan Dukler, Guido Montufar, and Quanquan Gu · 2020
Closest in time.