Fetching the paper…
Reading the bibliography…
Deep neural networks (DNNs) have demonstrated dominating performance in many fields; since AlexNet, networks used in practice are going wider and deeper.
Training a 3-node neural network is np-complete
Avrim L Blum and Ronald L Rivest · 1993
Earlier work this paper cites.
Introductory Lectures on Convex Programming Volume: A Basic course , volume I
Yurii Nesterov · 2004
Earlier work this paper cites.
A robust gradient sampling algorithm for nonsmooth, nonconvex optimization
James V Burke, Adrian S Lewis, and Michael L Overton · 2005
Earlier work this paper cites.
Cryptographic hardness for learning intersections of halfspaces
Adam R Klivans and Alexander A Sherstov · 2009
Earlier work this paper cites.
A variant of azuma’s inequality for martingales with subgaussian tails
Ohad Shamir · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Training very deep networks
Rupesh K Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Earlier work this paper cites.
Complexity theoretic limitations on learning halfspaces
Amit Daniely · 2016
Earlier work this paper cites.
Complexity theoretic limitations on learning dnf’s
Amit Daniely and Shai Shalev-Shwartz · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Follow the Compressed Leader: Faster Online Learning of Eigenvectors and Faster MMWU
Zeyuan Allen-Zhu and Yuanzhi Li · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Neon2: Finding Local Minima via First-Order Oracles
Zeyuan Allen-Zhu and Yuanzhi Li · 2018
Closest in time.
Gradient descent with identity initialization efficiently learns positive definite linear transformations
Peter Bartlett, Dave Helmbold, and Phil Long · 2018
Closest in time.
Safely learning to control the constrained linear quadratic regulator
Sarah Dean, Stephen Tu, Nikolai Matni, and Benjamin Recht · 2018
Closest in time.
Learning one convolutional layer with overlapping patches
Surbhi Goel, Adam Klivans, and Raghu Meka · 2018
Closest in time.
Gradient descent learns linear dynamical systems
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2017
Cited alongside, same era.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D. Lee, and Tengyu Ma · 2017
Cited alongside, same era.
Reliably learning the ReLU in polynomial time
Surbhi Goel, Varun Kanade, Adam Klivans, and Justin Thaler · 2017
Cited alongside, same era.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Cited alongside, same era.
Learning linear dynamical systems via spectral filtering
Elad Hazan, Karan Singh, and Cyril Zhang · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with ReLU activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Elad Hazan, Holden Lee, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Closest in time.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Closest in time.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Closest in time.
The computational complexity of training ReLU(s)
Pasin Manurangsi and Daniel Reichman · 2018
Closest in time.
Robust spectral filtering and anomaly detection
Jakub Marecek and Tigran Tchrakian · 2018
Closest in time.
Learning compact neural networks with regularization
Samet Oymak · 2018
Closest in time.
Non-asymptotic identification of LTI systems from a single trajectory
Samet Oymak and Necmiye Ozay · 2018
Closest in time.
Convergence results for neural networks via electrodynamics
Rina Panigrahy, Ali Rahimi, Sushant Sachdeva, and Qiuyi Zhang · 2018
Closest in time.
Spurious local minima are common in two-layer ReLU neural networks
Itay Safran and Ohad Shamir · 2018
Closest in time.
Learning without mixing: Towards a sharp analysis of linear system identification
Max Simchowitz, Horia Mania, Stephen Tu, Michael I Jordan, and Benjamin Recht · 2018
Closest in time.
Classification on CIFAR-10/100 and ImageNet with PyTorch, 2018
Wei Yang · 2018
Closest in time.
Learning long term dependencies via Fourier recurrent units
Jiong Zhang, Yibo Lin, Zhao Song, and Inderjit S Dhillon · 2018
Closest in time.
What Can ResNet Learn Efficiently, Going Beyond Kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Closest in time.