Fetching the paper…
Reading the bibliography…
Proper regularization is critical for speeding up training, improving generalization performance, and learning compact models that are cost efficient.
Optimal brain damage
Yann Le Cun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G Stork · 1993
Earlier work this paper cites.
The generic chaining: upper and lower bounds of stochastic processes
M. Talagrand · 2006
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Benjamin Recht, Maryam Fazel, and Pablo A Parrilo · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2010
Earlier work this paper cites.
Tight oracle inequalities for low-rank matrix recovery from a minimal number of noisy random measurements
Emmanuel J Candes and Yaniv Plan · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
The convex geometry of linear inverse problems
V. Chandrasekaran, B. Recht, P. A. Parrilo, and A. S. Willsky · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Restricted strong convexity and weighted matrix completion: Optimal bounds with noise
Sahand Negahban and Martin J Wainwright · 2012
Earlier work this paper cites.
Tail bounds via generic chaining
S. Dirksen · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Hanson-wright inequality and sub-gaussian concentration
Mark Rudelson, Roman Vershynin, et al · 2013
Earlier work this paper cites.
Living on the edge: Phase transitions in convex programs with random data
Dennis Amelunxen, Martin Lotz, Michael B McCoy, and Joel A Tropp · 2014
Earlier work this paper cites.
Provable bounds for learning some deep representations
Sanjeev Arora, Aditya Bhaskara, Rong Ge, and Tengyu Ma · 2014
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
Emily L Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Learning without concentration
S. Mendelson · 2014
Earlier work this paper cites.
Randomized sketches of convex programs with sharp guarantees
M. Pilanci and M. J. Wainwright · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Gaussian processes and the generic chaining
Michel Talagrand · 2014
Cited alongside, same era.
Convex recovery of a structured signal from independent random linear measurements
J. A. Tropp · 2014
Cited alongside, same era.
A note on the hanson-wright inequality for random vectors with dependencies
Radoslaw Adamczak et al · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Cited alongside, same era.
Song Han, Huizi Mao, and William J Dally · 2015
Cited alongside, same era.
Net-trim: Convex pruning of deep neural networks with performance guarantee
Alireza Aghasi, Afshin Abdi, Nam Nguyen, and Justin Romberg · 2017
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Later among the works it cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Later among the works it cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Xin Dong, Shangyu Chen, and Sinno Pan · 2017
Later among the works it cites.
When is a convolutional filter easy to learn?
Simon S Du, Jason D Lee, and Yuandong Tian · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Cited alongside, same era.
Beating the perils of non-convexity: Guaranteed training of neural networks using tensor methods
Majid Janzamin, Hanie Sedghi, and Anima Anandkumar · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
Beyond sub-gaussian measurements: High-dimensional structured estimation with sub-exponential designs
Vidyashankar Sivakumar, Arindam Banerjee, and Pradeep K Ravikumar · 2015
Cited alongside, same era.
Complete dictionary recovery over the sphere
Ju Sun, Qing Qu, and John Wright · 2015
Cited alongside, same era.
Low-rank solutions of linear matrix equations via procrustes flow
Stephen Tu, Ross Boczar, Max Simchowitz, Mahdi Soltanolkotabi, and Benjamin Recht · 2015
Cited alongside, same era.
Simon S Du, Jason D Lee, Yuandong Tian, Barnabas Poczos, and Aarti Singh · 2017
Later among the works it cites.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Later among the works it cites.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Later among the works it cites.
Generalization in deep learning
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Later among the works it cites.
Pac-bayesian margin bounds for convolutional neural networks-technical report
Pitas Konstantinos, Mike Davies, and Pierre Vandergheynst · 2017
Later among the works it cites.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Later among the works it cites.
Sharp time–data tradeoffs for linear inverse problems
Samet Oymak, Benjamin Recht, and Mahdi Soltanolkotabi · 2017
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2017
Later among the works it cites.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Later among the works it cites.
Mahdi Soltanolkotabi · 2017
Later among the works it cites.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2017
Later among the works it cites.
Yuandong Tian · 2017
Later among the works it cites.
Learning non-overlapping convolutional neural networks with multiple kernels
Kai Zhong, Zhao Song, and Inderjit S Dhillon · 2017
Later among the works it cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Later among the works it cites.
End-to-end learning of a convolutional neural network via deep tensor decomposition
Samet Oymak and Mahdi Soltanolkotabi · 2018
Closest in time.
Learning the input layer of a deep convolutional neural network via centered gradient descent
Samet Oymak and Mahdi Soltanolkotabi · 2018
Closest in time.
Convergence results for neural networks via electrodynamics
Rina Panigrahy, Ali Rahimi, Sushant Sachdeva, and Qiuyi Zhang · 2018
Closest in time.