Fetching the paper…
Reading the bibliography…
The cost of hyperparameter tuning in deep learning has been rising with model sizes, prompting practitioners to find new tuning methods using a proxy of smaller networks.
109. stochastic integral
Kiyosi Itô · 1944
Earlier work this paper cites.
Principles of mathematical analysis
Walter Rudin · 1953
Earlier work this paper cites.
Statistical dynamics of classical systems
Paul Cecil Martin, ED Siggia, and HA Rose · 1973
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence Saul · 2009
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Radford M Neal · 2012
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, JP Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Deep residual networks and weight initialization
Masato Taki · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Earlier work this paper cites.
Finite size corrections for neural network gaussian processes
Joseph M Antognini · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Earlier work this paper cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Dynamical isometry is achieved in residual networks in a universal way for any activation function
Wojciech Tarnowski, Piotr Warchoł, Stanisław Jastrzebski, Jacek Tabor, and Maciej Nowak · 2019
Earlier work this paper cites.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 2019
Earlier work this paper cites.
On the distance between two neural networks and the stability of learning
Jeremy Bernstein, Arash Vahdat, Yisong Yue, and Ming-Yu Liu · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Products of many large random matrices and gradients in deep neural networks
Boris Hanin and Mihai Nica · 2020
Cited alongside, same era.
Improving transformer optimization through better initialization
Xiao Shi Huang, Felipe Perez, Jimmy Ba, and Maksims Volkovs · 2020
Cited alongside, same era.
Correlation functions in random fully connected neural networks at finite width
Boris Hanin · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
The neural covariance sde: Shaped infinite depth-and-width networks at initialization
Mufan Li, Mihai Nica, and Dan Roy · 2022
Later among the works it cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi · 2022
Later among the works it cites.
Unified field theoretical approach to deep and recurrent neuronal networks
Kai Segadlo, Bastian Epping, Alexander van Meegen, David Dahmen, Michael Krämer, and Moritz Helias · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification
Francesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborová · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu · 2020
Cited alongside, same era.
Non-gaussian processes and neural networks at finite widths
Sho Yaida · 2020
Cited alongside, same era.
Tensor programs ii: Neural tangent kernel for any architecture
Greg Yang · 2020
Cited alongside, same era.
Rezero is all you need: Fast convergence at large depth
Thomas Bachlechner, Bodhisattwa Prasad Majumder, Henry Mao, Gary Cottrell, and Julian McAuley · 2021
Cited alongside, same era.
Batch normalization orthogonalizes representations in deep random networks
Hadi Daneshmand, Amir Joudaki, and Francis Bach · 2021
Cited alongside, same era.
Later among the works it cites.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward J Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2022
Later among the works it cites.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Later among the works it cites.
Automatic gradient descent: Deep learning without hyperparameters
Jeremy Bernstein, Chris Mingard, Kevin Huang, Navid Azizan, and Yisong Yue · 2023
Closest in time.
Dynamics of finite width kernel and prediction fluctuations in mean field neural networks
Blake Bordelon and Cengiz Pehlevan · 2023
Closest in time.
Neural signature kernels as infinite-width-depth-limits of controlled resnets
Nicola Muca Cirone, Maud Lemercier, and Cristopher Salvi · 2023
Closest in time.
Optimal signal propagation in resnets through residual scaling
Kirsten Fischer, David Dahmen, and Moritz Helias · 2023
Closest in time.
Bayesian interpolation with deep linear networks
Boris Hanin and Alexander Zlokapa · 2023
Closest in time.
On the infinite-depth limit of finite-width neural networks
Soufiane Hayou · 2023
Closest in time.
Width and depth limits commute in residual networks
Soufiane Hayou and Greg Yang · 2023
Closest in time.
Depth dependence of μ \mu p learning rates in relu mlps
Samy Jelassi, Boris Hanin, Ziwei Ji, Sashank J Reddi, Srinadh Bhojanapalli, and Sanjiv Kumar · 2023
Closest in time.
On the impact of activation and normalization in obtaining isometric embeddings at initialization
Amir Joudaki, Hadi Daneshmand, and Francis Bach · 2023
Closest in time.
Scaling laws for deep learning based image reconstruction
Tobit Klug and Reinhard Heckel · 2023
Closest in time.
The shaped transformer: Attention models in the infinite depth-and-width limit
Lorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He, Thomas Hofmann, Chris Maddison, and Daniel M Roy · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Llama: open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Feature-learning networks are consistent across widths at realistic scales, 2023
Nikhil Vyas, Alexander Atanasov, Blake Bordelon, Depen Morwani, Sabarish Sainathan, and Cengiz Pehlevan · 2023
Closest in time.