Fetching the paper…
Reading the bibliography…
Deep ResNets are recognized for achieving state-of-the-art results in complex machine learning tasks.
Convergence theory of learning over-parameterized ResNet: A full characterization
H. Zhang, D. Yu, M. Yi, W. Chen, and T.-Y. Liu · 1903
Earlier work this paper cites.
Numerical Solution of Stochastic Differential Equations
P.E. Kloeden and E. Platen · 1992
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
The MNIST database of handwritten digit images for machine learning research
L. Deng · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
A. Maas, A. Hannun, and A. Ng · 2013
Earlier work this paper cites.
Concentration in unbounded metric spaces and algorithmic stability
A. Kontorovich · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D.P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Empirical evaluation of rectified activations in convolutional network
B. Xu, N. Wang, T. Chen, and M. Li · 2015
Earlier work this paper cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Probability in High Dimension
R. van Handel · 2016
Earlier work this paper cites.
Three factors influencing minima in SGD
S. Jastrzkebski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Mean field residual networks: On the edge of chaos
G. Yang and S. Schoenholz · 2017
Cited alongside, same era.
Neural ordinary differential equations
R.T.Q. Chen, Y. Rubanova, J. Bettencourt, and D.K. Duvenaud · 2018
Cited alongside, same era.
How to start training: The effect of initialization and architecture
B. Hanin and D. Rolnick · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Z. Allen-Zhu, Y. Li, and Z. Song · 2019
Is normalization indispensable for training deep neural network?
J. Shao, K. Hu, C. Wang, X. Xue, and B. Raj · 2020
Later among the works it cites.
ReZero is all you need: Fast convergence at large depth
T. Bachlechner, B.P. Majumder, H. Mao, G. Cottrell, and J. Auley · 2021
Later among the works it cites.
Characterizing signal propagation to close the performance gap in unnormalized ResNets
A. Brock, S. De, and S.L. Smith · 2021
Later among the works it cites.
Scaling properties of deep residual networks
A.-S. Cohen, R. Cont, A. Rossier, and R. Xu · 2021
Later among the works it cites.
Framing RNN as a kernel method: A neural ODE approach
A. Fermanian, P. Marion, J.-P. Vert, and G. Biau · 2021
Later among the works it cites.
Rate of convergence to the circular law via smoothing inequalities for log-potentials
F. Götze and J. Jalowy · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How to initialize your network? Robust initialization for WeightNorm & ResNets
D. Arpit, V. Campos, and Y. Bengio · 2019
Cited alongside, same era.
AntisymmetricRNN: A dynamical system view on recurrent neural networks
B. Chang, M. Chen, E. Haber, and E.H. Chi · 2019
Cited alongside, same era.
A mean-field optimal control formulation of deep learning
W. E, J. Han, and Q. Li · 2019
Cited alongside, same era.
FFJORD: Free-form continuous dynamics for scalable reversible generative models
W. Grathwohl, R.T.Q. Chen, J. Bettencourt, I. Sutskever, and D. Duvenaud · 2019
Cited alongside, same era.
Uniform error bounds for Gaussian process regression with application to safe control
A. Lederer, J. Umlauft, and S. Hirche · 2019
Cited alongside, same era.
Batch normalization biases residual blocks towards the identity function in deep networks
S. De and S. Smith · 2020
Cited alongside, same era.
Later among the works it cites.
Stable ResNet
S. Hayou, E. Clerico, B. He, G. Deligiannidis, A. Doucet, and J. Rousseau · 2021
Later among the works it cites.
Efficient and accurate gradients for neural SDEs
P. Kidger, J. Foster, X. Li, and T. Lyons · 2021
Later among the works it cites.
Asymptotic analysis of deep residual networks
R. Cont, A. Rossier, and R. Xu · 2022
Closest in time.
Do residual neural networks discretize neural ordinary differential equations?
M.E. Sander, P. Ablin, and G. Peyré · 2022
Closest in time.
Stabilize deep resnet with a sharp scaling factor τ \tau
Huishuai Zhang, Da Yu, Mingyang Yi, Wei Chen, and Tie-Yan Liu · 2022
Closest in time.
Stability of deep neural networks via discrete rough paths
Christian Bayer, Peter K Friz, and Nikolas Tapia · 2023
Closest in time.
Deep limits of residual neural networks
Matthew Thorpe and Yves van Gennip · 2023
Closest in time.
Implicit regularization of deep residual networks towards neural ODEs
P. Marion, Y.-H. Wu, M.E. Sander, and G. Biau · 2024
Closest in time.
Deepnet: Scaling transformers to 1,000 layers
Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, and Furu Wei · 2024
Closest in time.