Fetching the paper…
Reading the bibliography…
Residual networks have significantly better trainability and thus performance than feed-forward networks at large depth.
R. Price, A useful theorem for nonlinear devices having gaussian inputs, IRE Trans. Inf. Theory 4
1958
Earlier work this paper cites.
J. Hertz, A. Krogh, and R. G. Palmer, Introduction to the Theory of Neural Computation (Perseus Books, 1991)
1991
Earlier work this paper cites.
J. Zinn-Justin, Quantum field theory and critical phenomena (Clarendon Press, Oxford, 1996)
1996
Earlier work this paper cites.
A. Papoulis and S. U. Pillai, Probability, Random Variables, and Stochastic Processes , 4th ed. (McGraw-Hill, Boston, 2002)
2002
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, Learning multiple layers of features from tiny images , Tech. Rep. (2009)
2009
Earlier work this paper cites.
M. Welling and Y. W. Teh, Bayesian learning via stochastic gradient langevin dynamics, in Proceedings of the 28th International Conference on International Conference on Machine Learning (2011) pp. 681–688
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, Imagenet classification with deep convolutional neural networks, in Adv. Neural Inf. Process. Syst. , Vol. 25, edited by F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger (Curran Associates, Inc., 2012) pp. 1097–1105
2012
Earlier work this paper cites.
S. Ioffe and C. Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift, in Proceedings of the 32nd International Conference on Machine Learning , Proc. Mach. Learn. Res., Vol. 37, edited by F. Bach and D. Blei (PMLR, Lille, France, 2015) pp. 448–456
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al. , Mastering the game of go with deep neural networks and tree search, Nature 529
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, Deep residual learning for image recognition, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, Identity mappings in deep residual networks, in European Conference on Computer Vision , Vol. 9908 (2016) pp. 630–645
2016
Earlier work this paper cites.
S. Li, J. Jiao, Y. Han, and T. Weissman, Demystifying resnet, CoRR abs/1611.01186
2016
Earlier work this paper cites.
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli, Exponential expressivity in deep neural networks through transient chaos, in Advances in Neural Information Processing Systems 29 (2016)
2016
Earlier work this paper cites.
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, Inception-v4, inception-resnet and the impact of residual connections on learning, in Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence , AAAI’17 (AAAI Press, 2017) pp. 4278–4284
2017
Earlier work this paper cites.
S. S. Schoenholz, J. Gilmer, S. Ganguli, and J. Sohl-Dickstein, Deep information propagation, 5th International Conference on Learning Representations, ICLR 2017 - Conference Track Proceedings (2017)
2017
Earlier work this paper cites.
G. Yang and S. Schoenholz, Mean field residual networks: On the edge of chaos, in Adv. Neural Inf. Process. Syst. , Vol. 30, edited by I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Curran Associates, Inc., 2017)
2017
Earlier work this paper cites.
S. M, t, M. D. Hoffman, and D. M. Blei, Stochastic gradient descent as approximate bayesian inference, Journal of Machine Learning Research 18
2017
Earlier work this paper cites.
D. Balduzzi, M. Frean, L. Leary, J. P. Lewis, K. W.-D. Ma, and B. McWilliams, The shattered gradients problem: If resnets are the answer, then what is the question?, in Proceedings of the 34th International Conference on Machine Learning , Proc. Mach. Learn. Res., Vol. 70, edited by D. Precup and Y. W. Teh (PMLR, 2017) pp. 342–350
2017
Earlier work this paper cites.
A. Jacot, F. Gabriel, and C. Hongler, Neural tangent kernel: Convergence and generalization in neural networks, in Advances in Neural Information Processing Systems 31 (2018) pp. 8580–8589
2018
Earlier work this paper cites.
B. Hanin, Which neural net architectures give rise to exploding and vanishing gradients?, in Advances in Neural Information Processing Systems , Vol. 31, edited by S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., 2018)
2018
Earlier work this paper cites.
R. T. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud, Neural ordinary differential equations, in Advances in neural information processing systems (2018) pp. 6571–6583
2018
Cited alongside, same era.
J. Zhang, B. Han, L. Wynter, B. K. H. Low, and M. Kankanhalli, Towards robust resnet: A small step but a giant leap, in Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19 (International Joint Conferences on Artificial Intelligence Organization, 2019) pp. 4285–4291
2019
Cited alongside, same era.
D. Arpit, V. Campos, and Y. Bengio, How to initialize your network? robust initialization for weightnorm & resnets, in Advances in Neural Information Processing Systems , Vol. 32, edited by H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Curran Associates, Inc., 2019)
2019
Cited alongside, same era.
L. Chizat, E. Oyallon, and F. Bach, On lazy training in differentiable programming, in Adv. Neural Inf. Process. Syst. , Vol. 32 (2019)
J. A. Zavatone-Veth, W. L. Tong, and C. Pehlevan, Contrasting random and learned features in deep bayesian linear regression, Phys. Rev. E 105
2022
Later among the works it cites.
R. Barboni, G. Peyré, and F.-X. Vialard, On global convergence of resnets: From finite to infinite width using linear parameterization, in Advances in Neural Information Processing Systems , Vol. 35, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc., 2022) pp. 16385–16397
2022
Later among the works it cites.
D. Barzilai, A. Geifman, M. Galun, and R. Basri, A kernel perspective of skip connections in convolutional networks, in The Eleventh International Conference on Learning Representations (2023)
2023
Closest in time.
I. Seroussi, G. Naveh, and Z. Ringel, Separation of scales and a thermodynamic description of feature learning in some cnns, Nat. Commun. 14
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Y. Furusho and K. Ikeda, Resnet and batch-normalization improve data separability, in Proceedings of The Eleventh Asian Conference on Machine Learning , Proc. Mach. Learn. Res., Vol. 101, edited by W. S. Lee and T. Suzuki (PMLR, 2019) pp. 94–108
2019
Cited alongside, same era.
Z. Ling and R. C. Qiu, Spectrum concentration in deep residual learning: A free probability approach, 7
2019
Cited alongside, same era.
K. Huang, Y. Wang, M. Tao, and T. Zhao, Why do deep residual networks generalize better than deep feedforward networks? — a neural tangent kernel perspective, in Adv. Neural Inf. Process. Syst. , Vol. 33, edited by H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin (Curran Associates, Inc., 2020) pp. 2698–2709
2020
Cited alongside, same era.
T. Bachlechner, B. P. Majumder, H. Mao, G. Cottrell, and J. McAuley, Rezero is all you need: fast convergence at large depth, in Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence , Proceedings of Machine Learning Research, Vol. 161, edited by C. de Campos and M. H. Maathuis (PMLR, 2021) pp. 1352–1361
2021
Cited alongside, same era.
S. Hayou, E. Clerico, B. He, G. Deligiannidis, A. Doucet, and J. Rousseau, Stable resnet, in Proceedings of The 24th International Conference on Artificial Intelligence and Statistics , Proc. Mach. Learn. Res., Vol. 130, edited by A. Banerjee and K. Fukumizu (PMLR, 2021) pp. 1324–1332
2021
Cited alongside, same era.
S. Hayou, J.-F. Ton, A. Doucet, and Y. W. Teh, Robust pruning at initialization, in International Conference on Learning Representations (2021)
2021
Cited alongside, same era.
G. Yang, E. J. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao, Tuning large neural networks via zero-shot hyperparameter transfer, in Adv. Neural Inf. Process. Syst. , edited by A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (2021)
2021
Cited alongside, same era.
J. Halverson, A. Maiti, and K. Stoner, Neural networks and quantum field theory, Machine Learning: Science and Technology 2
2021
Cited alongside, same era.
A. X. Yang, M. Robeyns, E. Milsom, B. Anson, N. Schoots, and L. Aitchison, A theory of representation learning gives a deep generalisation of kernel methods, in Proceedings of the 40th International Conference on Machine Learning , Proc. Mach. Learn. Res., Vol. 202, edited by A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (PMLR, 2023) pp. 39380–39415
2023
Closest in time.
2023
Closest in time.
B. Bordelon and C. Pehlevan, Self-consistent dynamical field theory of kernel evolution in wide neural networks*, J. Stat. Mech. Theory Exp. 2023
2023
Closest in time.
2023
Closest in time.
D. Doshi, T. He, and A. Gromov, Critical initialization of wide and deep neural networks using partial jacobians: General theory and applications, in Thirty-seventh Conference on Neural Information Processing Systems (2023)
2023
Closest in time.
B. Bordelon, L. Noci, M. B. Li, B. Hanin, and C. Pehlevan, Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit, in The Twelfth International Conference on Learning Representations (2024)
2024
Closest in time.
K. Fischer, J. Lindner, D. Dahmen, Z. Ringel, M. Krämer, and M. Helias, Critical feature learning in deep neural networks, in Proceedings of the 41st International Conference on Machine Learning , Proceedings of Machine Learning Research, Vol. 235, edited by R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (PMLR, 2024) pp. 13660–13690
2024
Closest in time.
N. Rubin, I. Seroussi, and Z. Ringel, Grokking as a first order phase transition in two layer networks, in The Twelfth International Conference on Learning Representations (2024)
2024
Closest in time.
2024
Closest in time.
Y. Chen, F. Liu, Y. Lu, G. Chrysos, and V. Cevher, Generalization of scaled deep resnets in the mean-field regime, in The Twelfth International Conference on Learning Representations (2024)
2024
Closest in time.
M. B. Li and M. Nica, Differential equation scaling limits of shaped and unshaped neural networks, Transactions on Machine Learning Research (2024) , expert Certification
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. Rubin, K. Fischer, J. Lindner, I. Seroussi, Z. Ringel, M. Krämer, and M. Helias, From kernels to features: A multi-scale adaptive theory of feature learning, in Forty-second International Conference on Machine Learning (2025)
2025
Closest in time.
A.-S. Cohen, R. Cont, A. Rossier, and R. Xu, Scaling properties of deep residual networks, in Proceedings of the 38th International Conference on Machine Learning , Proceedings of Machine Learning Research, Vol. 139, edited by M. Meila and T. Zhang (PMLR, 2021) pp. 2039–2048
2048
Closest in time.