Fetching the paper…
Reading the bibliography…
Deep neural networks' remarkable ability to correctly fit training data when optimized by gradient-based algorithms is yet to be fully understood.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Complexity of linear regions in deep networks
Boris Hanin and David Rolnick · 1901
Earlier work this paper cites.
Training over-parameterized deep resnet is almost as easy as training a two-layer network
Huishuai Zhang, Da Yu, Wei Chen, and Tie-Yan Liu · 1903
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Avrim Blum and Ronald L Rivest · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, Halbert White, et al · 1989
Earlier work this paper cites.
A cooperative coevolutionary approach to function optimization
Mitchell A Potter and Kenneth A De Jong · 1994
Earlier work this paper cites.
The volume of convex bodies and Banach space geometry , volume 94
Gilles Pisier · 1999
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Sparse nonnegative matrix approximation: new formulations and algorithms
Rashish Tandon and Suvrit Sra · 2010
Earlier work this paper cites.
On the expressive power of deep architectures
Yoshua Bengio and Olivier Delalleau · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Cltune: A generic auto-tuner for opencl kernels
Cedric Nugteren and Valeriu Codreanu · 2015
Earlier work this paper cites.
Training very deep networks
Rupesh K Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Representation benefits of deep feedforward networks
Matus Telgarsky · 2015
Earlier work this paper cites.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li · 2015
Earlier work this paper cites.
On the expressive power of deep learning: A tensor analysis
Nadav Cohen, Or Sharir, and Amnon Shashua · 2016
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Binarized neural networks
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Why deep neural networks for function approximation?
Shiyu Liang and Rayadurgam Srikant · 2016
Earlier work this paper cites.
Local minima in training of deep networks
Grzegorz Swirszcz, Wojciech Marian Czarnecki, and Razvan Pascanu · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Improving training of deep neural networks via singular value bounding
Kui Jia, Dacheng Tao, Shenghua Gao, and Xiangmin Xu · 2017
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2017
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Later among the works it cites.
How much over-parameterization is sufficient to learn deep relu networks?
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Later among the works it cites.
Gradient descent finds global minima for generalizable deep neural networks of practical sizes
Kenji Kawaguchi and Jiaoyang Huang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2017
Cited alongside, same era.
Diracnets: Training very deep neural networks without skip-connections
Sergey Zagoruyko and Nikos Komodakis · 2017
Cited alongside, same era.
Learning non-overlapping convolutional neural networks with multiple kernels
Kai Zhong, Zhao Song, and Inderjit S Dhillon · 2017
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
Peter Bartlett, Dave Helmbold, and Philip Long · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
Simon S Du and Jason D Lee · 2018
Cited alongside, same era.
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby · 2019
Later among the works it cites.
Faster least squares optimization
Jonathan Lacotte and Mert Pilanci · 2019
Later among the works it cites.
Enhanced convolutional neural tangent kernels
Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S Du, Wei Hu, Ruslan Salakhutdinov, and Sanjeev Arora · 2019
Later among the works it cites.
Dying relu and initialization: Theory and numerical examples
Lu Lu, Yeonjong Shin, Yanhui Su, and George Em Karniadakis · 2019
Later among the works it cites.
Xnas: Neural architecture search with expert advice
Niv Nayman, Asaf Noy, Tal Ridnik, Itamar Friedman, Rong Jin, and Lihi Zelnik · 2019
Later among the works it cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Later among the works it cites.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Xiaoxia Wu, Simon S Du, and Rachel Ward · 2019
Later among the works it cites.
Disentangling trainability and generalization in deep learning
Lechao Xiao, Jeffrey Pennington, and Sam Schoenholz · 2019
Later among the works it cites.
Small relu networks are powerful memorizers: a tight analysis of memorization capacity
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2019
Later among the works it cites.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Later among the works it cites.
Rezero is all you need: Fast convergence at large depth
Thomas Bachlechner, Bodhisattwa Prasad Majumder, Huanru Henry Mao, Garrison W Cottrell, and Julian McAuley · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le · 2020
Later among the works it cites.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2020
Later among the works it cites.
Kaixuan Huang, Yuqing Wang, Molei Tao, and Tuo Zhao · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel S Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
Just interpolate: Kernel “ridgeless” regression can generalize
Tengyuan Liang, Alexander Rakhlin, et al · 2020
Later among the works it cites.
Asap: Architecture search, anneal and prune
Asaf Noy, Niv Nayman, Tal Ridnik, Nadav Zamir, Sivan Doveh, Itamar Friedman, Raja Giryes, and Lihi Zelnik · 2020
Later among the works it cites.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Tresnet: High performance gpu-dedicated architecture
Tal Ridnik, Hussam Lawen, Asaf Noy, and Itamar Friedman · 2020
Later among the works it cites.