Fetching the paper…
Reading the bibliography…
Recent studies revealed the mathematical connection between deep neural networks (DNNs) and dynamic systems.
A shortest augmenting path algorithm for dense and sparse linear assignment problems
Roy Jonker and Anton Volgenant · 1987
Earlier work this paper cites.
A numerical method for the optimal time-continuous mass transport problem and related problems
Jean-David Benamou and Yann Brenier · 1999
Earlier work this paper cites.
Robustness and generalization
Huan Xu and Shie Mannor · 2012
Earlier work this paper cites.
A user’s guide to optimal transport
Luigi Ambrosio and Nicola Gigli · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
Andreas Veit, Michael J Wilber, and Serge Belongie · 2016
Earlier work this paper cites.
Highway and residual networks learn unrolled iterative estimation
Klaus Greff, Rupesh K Srivastava, and Jürgen Schmidhuber · 2017
Earlier work this paper cites.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Earlier work this paper cites.
Convolutional neural networks analyzed via convolutional sparse coding
Vardan Papyan, Yaniv Romano, and Michael Elad · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A proposal on machine learning via dynamical systems
E Weinan · 2017
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Residual connections encourage iterative inference
Stanisław Jastrzebski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio · 2018
Cited alongside, same era.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2019
Prevalence of neural collapse during the terminal phase of deep learning training
Vardan Papyan, X. Y. Han, and David L. Donoho · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu · 2020
Later among the works it cites.
Towards understanding residual and dilated dense neural networks via convolutional sparse coding
Zhiyang Zhang and Shihua Zhang · 2020
Later among the works it cites.
Revisiting the importance of individual units in cnns via ablation
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba · 2020
Later among the works it cites.
Residual alignment: Uncovering the mechanisms of residual networks
Jianing Li and Vardan Papyan · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Resnets ensemble via the feynman-kac formalism to improve natural and robust accuracies
Bao Wang, Zuoqiang Shi, and Stanley Osher · 2019
Cited alongside, same era.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba · 2020
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo
Cited in the paper.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang
Cited in the paper.
Transformer block coupling and its correlation with generalization in llms
Murdock Aubry, Haoming Meng, Anton Sugolov, and Vardan Papyan · 2025
Closest in time.
Tracing representation progression: Analyzing and enhancing layer-wise similarity
Jiachen Jiang, Jinxin Zhou, and Zhihui Zhu · 2025
Closest in time.
Progressive feedforward collapse of resnet training
Sicong Wang, Kuo Gai, and Shihua Zhang · 2025
Closest in time.
OTAD: An optimal transport-induced robust model for agnostic adversarial attack, 2024
Kuo Gai, Sicong Wang, and Shihua Zhang · 2026
Closest in time.
mhc: Manifold-constrained hyper-connections, 2026
Zhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao, Chengqi Deng, Jiashi Li, Damai Dai, Huazuo Gao, Jiang Chang, Kuai Yu, Liang Zhao, Shangyan Zhou, Zhean Xu, Zhengyan Zhang, Wangding Zeng, Shengding Hu, Yuqing Wang, Jingyang Yuan, Lean Wang, and Wenfeng Liang · 2026
Closest in time.