Elementary Principles in Statistical Mechanics
Josiah Willard Gibbs · 1902
Earlier work this paper cites.
Contributions to the Theory of Games
A.W. Tucker and R.D. Luce · 1959
Earlier work this paper cites.
Visual feature extraction by a multilayered network of analog threshold elements
Kunihiko Fukushima · 1969
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Kunihiko Fukushima · 1980
Earlier work this paper cites.
Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition. neurocomputing: Algorithms, architectures and applications
John S Bridle, Soulié F.F., and Hérault J · 1989
Earlier work this paper cites.
Permitted and forbidden sets in symmetric threshold-linear networks
Richard Hahnloser and H Sebastian Seung · 2000
Earlier work this paper cites.
Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit
Richard HR Hahnloser, Rahul Sarpeshkar, Misha A Mahowald, Rodney J Douglas, and H Sebastian Seung · 2000
Earlier work this paper cites.
Efficient backprop
Yann LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2002
Earlier work this paper cites.
Pattern recognition and machine learning
C. M. Bishop · 2006
Earlier work this paper cites.
Learning deep architectures for ai
Y. Bengio, P. Simard, and P. Frasconi · 2006
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Earlier work this paper cites.
Multiplying matrices faster than coppersmith-winograd
Virginia Wassilevska Williams · 2012
Earlier work this paper cites.
Powers of tensors and fast matrix multiplication
Francois Le Gall · 2014
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
A near-optimal algorithm for approximating the john ellipsoid
Michael B Cohen, Ben Cousins, Yin Tat Lee, and Xin Yang · 2019
Earlier work this paper cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Earlier work this paper cites.
Gram-gauss-newton method: Learning overparameterized neural networks for regression problems
Original
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Earlier work this paper cites.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Earlier work this paper cites.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Matrix theory: optimization, concentration, and algorithms
Zhao Song · 2019
Earlier work this paper cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Original
Zhao Song and Xin Yang · 2019
Earlier work this paper cites.