Fetching the paper…
Reading the bibliography…
The study of deep neural networks (DNNs) in the infinite-width limit, via the so-called neural tangent kernel (NTK) approach, has provided new insights into the dynamics of learning, generalization, and the impact of initialization.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua. Bengio, Patrice. Simard, and Paolo. Frasconi · 1994
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M Neal · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chih-Chung Chang and Chih-Jen Lin · 2011
Earlier work this paper cites.
An iterative construction of solutions of the tap equations for the sherrington–kirkpatrick model
Erwin Bolthausen · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Avoiding pathologies in very deep networks
David Duvenaud, Oren Rippel, Ryan Adams, and Zoubin Ghahramani · 2014
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Earlier work this paper cites.
Recurrent orthogonal networks and long-memory tasks
Mikael Henaff, Arthur Szlam, and Yann LeCun · 2016
Earlier work this paper cites.
Tunable efficient unitary neural networks (eunn) and their application to rnns
Li Jing, Yichen Shen, Tena Dubcek, John Peurifoy, Scott Skirlo, Yann LeCun, Max Tegmark, and Marin Soljačić · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Jascha Sohl-dickstein, Jeffrey Pennington, Roman Novak, Sam Schoenholz, and Yasaman Bahri · 2018
Cited alongside, same era.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
A max-affine spline perspective of recurrent neural networks
Zichao Wang, Randall Balestriero, and Richard Baraniuk · 2018
Cited alongside, same era.
Stochastic gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Cited alongside, same era.
An extensive experimental survey of regression methods
M. Fernández-Delgado, M.S. Sirsat, E. Cernadas, S. Alawadi, S. Barro, and M. Febrero-Bande · 2019
Later among the works it cites.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Later among the works it cites.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Yasaman Bahri, Jaehoon Lee, Greg Yang, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-dickstein · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Cited alongside, same era.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Cited alongside, same era.
The UCR time series classification archive, 2019
Hoang Anh Dau, Eamonn Keogh, Kaveh Kamgar, Chin-Chia Michael Yeh, Yan Zhu, Shaghayegh Gharghabi, Chotirat Ann Ratanamahatana, Yanping Chen, Bing Hu, Nurjahan Begum, Anthony Bagnall, Abdullah Mueen, Gustavo Batista, and ML Hexagon · 2019
Cited alongside, same era.
Graph neural tangent kernel: Fusing graph neural networks with graph kernels
Simon S. Du, Kangcheng Hou, Russ R Salakhutdinov, Barnabas Poczos, Ruosong Wang, and Keyulu Xu · 2019
Cited alongside, same era.
Greg Yang · 2019
Later among the works it cites.
Greg Yang · 2019
Later among the works it cites.
Harnessing the power of infinitely wide deep nets on small-data tasks
Sanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu · 2020
Closest in time.
Simple and effective regularization methods for training on noisily labeled data with generalization guarantee
Wei Hu, Zhiyuan Li, and Dingli Yu · 2020
Closest in time.
Kaixuan Huang, Yuqing Wang, Molei Tao, and Tuo Zhao · 2020
Closest in time.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel S Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Closest in time.
Tensor programs II: Neural tangent kernel for any architecture
Greg Yang · 2020
Closest in time.
Tensor programs III: Neural matrix laws
Greg Yang · 2020
Closest in time.