Fetching the paper…
Reading the bibliography…
Recent analyses of neural networks with shaped activations (i.e.
Bayesian learning for neural networks , volume 118
Radford M Neal · 1995
Earlier work this paper cites.
Multidimensional diffusion processes , volume 233
Daniel W Stroock and SR Srinivasa Varadhan · 1997
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
Markov processes: characterization and convergence
Stewart N Ethier and Thomas G Kurtz · 2009
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sympy: symbolic computing in python
Aaron Meurer, Christopher P. Smith, Mateusz Paprocki, Ondřej Čertík, Sergey B. Kirpichev, Matthew Rocklin, AMiT Kumar, Sergiu Ivanov, Jason K. Moore, Sartaj Singh, Thilina Rathnayake, Sean Vig, Brian E. Granger, Richard P. Muller, Francesco Bonazzi, Harsh Gupta, Shivam Vats, Fredrik Johansson, Fabian Pedregosa, Matthew J. Curry, Andy R. Terrel, Štěpán Roučka, Ashutosh Saboo, Isuru Fernando, Sumith Kulal, Robert Cimrman, and Anthony Scopatz · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
Mean field analysis of neural networks: A law of large numbers, 2018
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Tensor programs i: Wide feedforward or recurrent neural networks of any architecture are gaussian processes, 2019
Greg Yang · 2019
Cited alongside, same era.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Cited alongside, same era.
Deep learning: a statistical viewpoint
Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin · 2021
Cited alongside, same era.
Stable ResNet
Soufiane Hayou, Eugenio Clerico, Bobby He, George Deligiannidis, Arnaud Doucet, and Judith Rousseau · 2021
Cited alongside, same era.
Foundations of Modern Probability
O. Kallenberg · 2021
Cited alongside, same era.
The neural covariance sde: Shaped infinite depth-and-width networks at initialization
Mufan Bill Li, Mihai Nica, and Daniel M Roy · 2022
Later among the works it cites.
Deep learning without shortcuts: Shaping the kernel with tailored rectifiers
Guodong Zhang, Aleksandar Botev, and James Martens · 2022
Later among the works it cites.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit, 2023
Blake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin, and Cengiz Pehlevan · 2023
Closest in time.
Neural signature kernels as infinite-width-depth-limits of controlled resnets
Nicola Muca Cirone, Maud Lemercier, and Cristopher Salvi · 2023
Closest in time.
Optimal signal propagation in resnets through residual scaling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The future is log-gaussian: Resnets and their infinite-depth-and-width limit at initialization
Mufan Li, Mihai Nica, and Dan Roy · 2021
Cited alongside, same era.
James Martens, Andy Ballard, Guillaume Desjardins, Grzegorz Swirszcz, Valentin Dalibard, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2021
Cited alongside, same era.
Precise characterization of the prior predictive distribution of deep relu networks
Lorenzo Noci, Gregor Bachmann, Kevin Roth, Sebastian Nowozin, and Thomas Hofmann · 2021
Cited alongside, same era.
Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2021
Cited alongside, same era.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2022
Cited alongside, same era.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Cited alongside, same era.
Kirsten Fischer, David Dahmen, and Moritz Helias · 2023
Closest in time.
Commutative width and depth scaling in deep neural networks, 2023
Soufiane Hayou · 2023
Closest in time.
Width and depth limits commute in residual networks
Soufiane Hayou and Greg Yang · 2023
Closest in time.
Cameron Jakub and Mihai Nica · 2023
Closest in time.
Towards training without depth limits: Batch normalization without gradient explosion, 2023
Alexandru Meterez, Amir Joudaki, Francesco Orabona, Alexander Immer, Gunnar Rätsch, and Hadi Daneshmand · 2023
Closest in time.
The shaped transformer: Attention models in the infinite depth-and-width limit
Lorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He, Thomas Hofmann, Chris Maddison, and Daniel M Roy · 2023
Closest in time.
Tensor programs vi: Feature learning in infinite-depth neural networks, 2023
Greg Yang, Dingli Yu, Chen Zhu, and Soufiane Hayou · 2023
Closest in time.
Geometric dyson brownian motion and the free log-normal for minor of products of random matrices
Mufan Li, Jaume de Dios Pont, Mihai Nica, and Daniel M. Roy · 2024
Closest in time.