Fetching the paper…
Reading the bibliography…
When and why can a neural network be successfully trained? This article provides an overview of optimization algorithms and theory for training neural networks.
On connected sublevel sets in deep learning
Quynh Nguyen · 1901
Earlier work this paper cites.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 1901
Earlier work this paper cites.
Mean field limit of the learning dynamics of multilayer neural networks
Phan-Minh Nguyen · 1902
Earlier work this paper cites.
Training over-parameterized deep resnet is almost as easy as training a two-layer network
Huishuai Zhang, Da Yu, Wei Chen, and Tie-Yan Liu · 1903
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Scaling laws for the principled design, initialization and preconditioning of relu networks
Aaron Defazio and Léon Bottou · 1906
Earlier work this paper cites.
Restart procedures for the conjugate gradient method
Michael James David Powell · 1977
Earlier work this paper cites.
Two-point step size gradient methods
Jonathan Barzilai and Jonathan M Borwein · 1988
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Sue Becker, Yann Le Cun, et al · 1988
Earlier work this paper cites.
Parallel and distributed computation: numerical methods , volume 23
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
On the convergence of the lms algorithm with adaptive learning rate for linear feedforward networks
Zhi-Quan Luo · 1991
Earlier work this paper cites.
Error bounds and convergence analysis of feasible descent methods: a general approach
Zhi-Quan Luo and Paul Tseng · 1993
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Innovations-based MLSE for Rayleigh flat fading channels
X. Yu and S. Pasupathy · 1995
Earlier work this paper cites.
Nonlinear programming
Dimitri P Bertsekas · 1997
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Efficient backprop
Yann LeCun, Leon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Numerical optimization
Stephen Wright and Jorge Nocedal · 1999
Earlier work this paper cites.
Adaptive method of realizing natural gradient learning for multilayer perceptrons
Shun-Ichi Amari, Hyeyoung Park, and Kenji Fukumizu · 2000
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Local minima and convergence in low-rank semidefinite programming
Samuel Burer and Renato DC Monteiro · 2005
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Methods of information geometry , volume 191
Shun-ichi Amari and Hiroshi Nagaoka · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Earlier work this paper cites.
Step-sizes for the gradient method
Ya-xiang Yuan · 2008
Earlier work this paper cites.
Sgd-qn: Careful quasi-newton stochastic gradient descent
Antoine Bordes, Léon Bottou, and Patrick Gallinari · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Dumitru Erhan, Yoshua Bengio, Aaron Courville, Pierre-Antoine Manzagol, Pascal Vincent, and Samy Bengio · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Randomized methods for linear constraints: convergence rates and conditioning
Dennis Leventhal and Adrian S Lewis · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Efficiency of coordiate descent methods on huge-scale optimization problems
Y. Nesterov · 2012
Earlier work this paper cites.
Optimization for machine learning
Suvrit Sra, Sebastian Nowozin, and Stephen J Wright · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward–backward splitting, and regularized gauss–seidel methods
Hedy Attouch, Jérôme Bolte, and Benar Fux Svaiter · 2013
Earlier work this paper cites.
First-order methods with inexact oracle: the strongly convex case
Olivier Devolder, François Glineur, Yurii Nesterov, et al · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
A unified convergence analysis of block successive minimization methods for nonsmooth optimization
Meisam Razaviyayn, Mingyi Hong, and Zhi-Quan Luo · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Mark Schmidt and Nicolas Le Roux · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
First-order methods of smooth convex optimization with inexact oracle
Olivier Devolder, François Glineur, and Yurii Nesterov · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
James Martens · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
A universal catalyst for first-order optimization
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Dmytro Mishkin and Jiri Matas · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Russ R Salakhutdinov, and Nati Srebro · 2015
Earlier work this paper cites.
Adaptive restart for accelerated gradient schemes
Brendan O?donoghue and Emmanuel Candes · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Earlier work this paper cites.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
A descent lemma beyond lipschitz gradient continuity: first-order methods revisited and applications
Heinz H Bauschke, Jérôme Bolte, and Marc Teboulle · 2016
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2016
Earlier work this paper cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Cited alongside, same era.
Singularity of the hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Cited alongside, same era.
Guaranteed matrix completion via non-convex factorization
Ruoyu Sun and Zhi-Quan Luo · 2016
Cited alongside, same era.
How to start training: The effect of initialization and architecture
Boris Hanin and David Rolnick · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2018
Later among the works it cites.
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, et al · 2018
Later among the works it cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham Kakade · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Local minima in training of deep networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Cited alongside, same era.
Barzilai-borwein step size for stochastic gradient descent
Conghui Tan, Shiqian Ma, Yu-Hong Dai, and Yuqiu Qian · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Cited alongside, same era.
Extremely large minibatch sgd: training resnet-50 on imagenet in 15 minutes
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda · 2017
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James Brecht · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
Rethinking the value of network pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell · 2018
Later among the works it cites.
Relatively smooth convex optimization by first-order methods, and applications
Haihao Lu, Robert M Freund, and Yurii Nesterov · 2018
Later among the works it cites.
A mean field view of the landscape of two-layers neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
Massively distributed sgd: Imagenet/resnet-50 training in a flash
Hiroaki Mikami, Hisahiro Suganuma, Yoshiki Tanaka, Yuichi Kageyama, et al · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Later among the works it cites.
On the connection between learning two-layers neural networks and tensor decomposition
Marco Mondelli and Andrea Montanari · 2018
Later among the works it cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
Quynh Nguyen, Mahesh Chandra Mukkamala, and Matthias Hein · 2018
Later among the works it cites.
Learning deep models: Critical points and local openness
Maher Nouiehed and Meisam Razaviyayn · 2018
Later among the works it cites.
Second-order optimization method for large mini-batch: Training resnet-50 on imagenet in 35 epochs
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2018
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Later among the works it cites.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel S Schoenholz, and Surya Ganguli · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Grant M Rotskoff and Eric Vanden-Eijnden · 2018
Later among the works it cites.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Later among the works it cites.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Ohad Shamir · 2018
Later among the works it cites.
Mean field analysis of neural networks
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le · 2018
Later among the works it cites.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2018
Later among the works it cites.
Polynomial convergence of gradient descent for training one-hidden-layer neural networks
Santosh Vempala and John Wilmes · 2018
Later among the works it cites.
Learning relu networks on linearly separable data: Algorithm, optimality, and generalization
Gang Wang, Georgios B Giannakis, and Jie Chen · 2018
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2018
Later among the works it cites.
On the margin theory of feedforward neural networks
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2018
Later among the works it cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S Schoenholz, and Jeffrey Pennington · 2018
Later among the works it cites.
First-order stochastic algorithms for escaping from saddle points in almost linear time
Yi Xu, Jing Rong, and Tianbao Yang · 2018
Later among the works it cites.
Image classification at supercomputer scale
Chris Ying, Sameer Kumar, Dehao Chen, Tao Wang, and Youlong Cheng · 2018
Later among the works it cites.
Imagenet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Later among the works it cites.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Later among the works it cites.
Block coordinate descent for deep learning: Unified convergence guarantees
Jinshan Zeng, Tim Tsz-Kit Lau, Shaobo Lin, and Yuan Yao · 2018
Later among the works it cites.
Learning one-hidden-layer relu networks via gradient descent
Xiao Zhang, Yaodong Yu, Lingxiao Wang, and Quanquan Gu · 2018
Later among the works it cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Later among the works it cites.
Critical points of linear neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2018
Later among the works it cites.
On the convergence of adagrad with momentum for training deep neural networks
Fangyu Zou and Li Shen · 2018
Later among the works it cites.
A mean-field limit for certain deep neural networks, 2019
Dyego Araújo, Roberto I. Oliveira, and Daniel Yukimura · 2019
Closest in time.
Convergence analysis of a momentum algorithm with adaptive step size for non convex optimization
Anas Barakat and Pascal Bianchi · 2019
Closest in time.
Quasi-newton methods for deep learning: Forget the past, just sample
Albert S Berahas, Majid Jahani, and Martin Takáč · 2019
Closest in time.
A quantitative analysis of the effect of batch normalization on gradient descent
Yongqiang Cai, Qianxiao Li, and Zuowei Shen · 2019
Closest in time.
Nonconvex optimization meets low-rank matrix factorization: An overview
Yuejie Chi, Yue M Lu, and Yuxin Chen · 2019
Closest in time.
Spurious local minima exist for almost all over-parameterized neural networks
Tian Ding, Dawei Li, and Ruoyu Sun · 2019
Closest in time.
On the spherical convexity of quadratic functions
O. P. Ferreira and S. Z. Németh · 2019
Closest in time.
The lottery ticket hypothesis at scale
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin · 2019
Closest in time.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Closest in time.
Dynamical isometry and a mean field theory of lstms and grus
Dar Gilboa, Bo Chang, Minmin Chen, Greg Yang, Samuel S Schoenholz, Ed H Chi, and Jeffrey Pennington · 2019
Closest in time.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2019
Closest in time.
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Closest in time.
Generalization error in deep learning
Daniel Jakubovitz, Raja Giryes, and Miguel RD Rodrigues · 2019
Closest in time.
Elimination of all bad local minima in deep learning
Kenji Kawaguchi and Leslie Pack Kaelbling · 2019
Closest in time.
Exponential convergence rates for batch normalization: The power of length-direction decoupling in non-convex optimization
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, Thomas Hofmann, Ming Zhou, and Klaus Neymeyr · 2019
Closest in time.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Sanjeev Arora, and Rong Ge · 2019
Closest in time.
SNIP: SINGLE-SHOT NETWORK PRUNING BASED ON CONNECTION SENSITIVITY
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip Torr · 2019
Closest in time.
On random deep weight-tied autoencoders: Exact asymptotic analysis, phase transitions, and implications to training
Ping Li and Phan-Minh Nguyen · 2019
Closest in time.
Enhanced convolutional neural tangent kernels, 2019
Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S. Du, Wei Hu, Ruslan Salakhutdinov, and Sanjeev Arora · 2019
Closest in time.
A sensitive-eigenvector based global algorithm for quadratically constrained quadratic programming
Cheng Lu, Zhibin Deng, Jing Zhou, and Xiaoling Guo · 2019
Closest in time.
Switchable normalization for learning-to-normalize deep representation
Ping Luo, Ruimao Zhang, Jiamin Ren, Zhanglin Peng, and Jingyu Li · 2019
Closest in time.
Analysis of the gradient descent algorithm for a deep neural network model with skip-connections
Chao Ma, Lei Wu, et al · 2019
Closest in time.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Closest in time.
Ari S Morcos, Haonan Yu, Michela Paganini, and Yuandong Tian · 2019
Closest in time.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee, Greg Yang, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-dickstein · 2019
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.
The singular values of convolutional layers
Hanie Sedghi, Vineet Gupta, and Philip M. Long · 2019
Closest in time.
Effects of depth, width, and initialization: A convergence analysis of layer-wise training for deep linear neural networks, 2019
Yeonjong Shin · 2019
Closest in time.
Mean field analysis of deep neural networks, 2019
Justin Sirignano and Konstantinos Spiliopoulos · 2019
Closest in time.
On the tunability of optimizers in deep learning, 2019
Prabhu Teja Sivaprasad, Florian Mai, Thijs Vogels, Martin Jaggi, and François Fleuret · 2019
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2019
Closest in time.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V Le · 2019
Closest in time.
Dynamical isometry is achieved in residual networks in a universal way for any activation function
Wojciech Tarnowski, Piotr Warchoł, Stanisław Jastrz?bski, Jacek Tabor, and Maciej Nowak · 2019
Closest in time.
Gradient dynamics of shallow univariate relu networks, 2019
Francis Williams, Matthew Trager, Claudio Silva, Daniele Panozzo, Denis Zorin, and Joan Bruna · 2019
Closest in time.
Yet another accelerated sgd: Resnet-50 training on imagenet in 74.7 seconds
Masafumi Yamazaki, Akihiko Kasagi, Akihiro Tabuchi, Takumi Honda, Masahiro Miwa, Naoto Fukumoto, Tsuguchika Tabaru, Atsushi Ike, and Kohta Nakashima · 2019
Closest in time.
Greg Yang · 2019
Closest in time.
Positively scale-invariant flatness of relu neural networks
Mingyang Yi, Qi Meng, Wei Chen, Zhi-ming Ma, and Tie-Yan Liu · 2019
Closest in time.
Reducing bert pre-training time from 3 days to 76 minutes
Yang You, Jing Li, Jonathan Hseu, Xiaodan Song, James Demmel, and Cho-Jui Hsieh · 2019
Closest in time.
Network slimming by slimmable networks: Towards one-shot architecture search for channel numbers
Jiahui Yu and Thomas Huang · 2019
Closest in time.
Depth creates no more spurious local minima
Li Zhang · 2019
Closest in time.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Hattie Zhou, Janice Lan, Rosanne Liu, and Jason Yosinski · 2019
Closest in time.