Fetching the paper…
Reading the bibliography…
Underpinning the past decades of work on the design, initialization, and optimization of neural networks is a seemingly innocuous assumption: that the network is trained on a \textit{stationary} data distribution.
On the representational efficiency of restricted boltzmann machines
James Martens, Arkadev Chattopadhya, Toni Pitassi, and Richard Zemel · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
How far can we go without convolution: Improving fully-connected networks
Zhouhan Lin, Roland Memisevic, and Kishore Konda · 2016
Earlier work this paper cites.
Path-normalized optimization of recurrent neural networks with relu activations
Behnam Neyshabur, Yuhuai Wu, Russ R Salakhutdinov, and Nati Srebro · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Earlier work this paper cites.
Learning values across many orders of magnitude
Hado P van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih, and David Silver · 2016
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, JP Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Deep information propagation
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Improving regression performance with distributional losses
Ehsan Imani and Martha White · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Cited alongside, same era.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Cited alongside, same era.
On the impact of the activation function on deep neural networks training
Soufiane Hayou, Arnaud Doucet, and Judith Rousseau · 2019
Cited alongside, same era.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz, and Noah A Smith · 2021
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
Mohammad Pezeshki, Oumar Kaba, Yoshua Bengio, Aaron C Courville, Doina Precup, and Guillaume Lajoie · 2021
Later among the works it cites.
Decoupling value and policy for generalization in reinforcement learning
Roberta Raileanu and Rob Fergus · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David GT Barrett, and Soham De · 2021
Later among the works it cites.
Deep learning without shortcuts: Shaping the kernel with tailored rectifiers
Guodong Zhang, Aleksandar Botev, and James Martens · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding and improving layer normalization
Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin · 2019
Cited alongside, same era.
Wide feedforward or recurrent neural networks of any architecture are gaussian processes
Greg Yang · 2019
Cited alongside, same era.
Scalable second order optimization for deep learning
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2020
Cited alongside, same era.
On warm-starting neural network training
Jordan Ash and Ryan P Adams · 2020
Cited alongside, same era.
Implicit gradient regularization
David GT Barrett and Benoit Dherin · 2020
Cited alongside, same era.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Aviral Kumar, Rishabh Agarwal, Dibya Ghosh, and Sergey Levine · 2020
Cited alongside, same era.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Cited alongside, same era.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2022
Later among the works it cites.
Learning dynamics and generalization in deep reinforcement learning
Clare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska, and Yarin Gal · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville · 2022
Later among the works it cites.
Loss of plasticity in continual deep reinforcement learning
Zaheer Abbas, Rosie Zhao, Joseph Modayil, Adam White, and Marlos C Machado · 2023
Later among the works it cites.
Resetting the optimizer in deep rl: An empirical study
Kavosh Asadi, Rasool Fakoor, and Shoham Sabach · 2023
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Peter Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al · 2023
Later among the works it cites.
Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation
Bobby He, James Martens, Guodong Zhang, Aleksandar Botev, Andrew Brock, Samuel L Smith, and Yee Whye Teh · 2023
Later among the works it cites.
Plastic: Improving input and label plasticity for sample efficient reinforcement learning
Hojoon Lee, Hanseul Cho, Hyunseung Kim, Daehoon Gwak, Joonkee Kim, Jaegul Choo, Se-Young Yun, and Chulhee Yun · 2023
Later among the works it cites.
Curvature explains loss of plasticity
Alex Lewandowski, Haruto Tanaka, Dale Schuurmans, and Marlos C Machado · 2023
Later among the works it cites.
Understanding plasticity in neural networks
Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, and Will Dabney · 2023
Later among the works it cites.
Deep reinforcement learning with plasticity injection
Evgenii Nikishin, Junhyuk Oh, Georg Ostrovski, Clare Lyle, Razvan Pascanu, Will Dabney, and André Barreto · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro, and Utku Evci · 2023
Later among the works it cites.
Regression as classification: Influence of task formulation on neural network features
Lawrence Stewart, Francis Bach, Quentin Berthet, and Jean-Philippe Vert · 2023
Later among the works it cites.
Small-scale proxies for large-scale transformer training instabilities
Mitchell Wortsman, Peter J Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D Co-Reyes, Izzeddin Gur, Abhishek Kumar, Roman Novak, et al · 2023
Later among the works it cites.