Fetching the paper…
Reading the bibliography…
In this paper, we provide an overview of a common phenomenon, condensation, observed during the nonlinear training of neural networks: During the nonlinear training of neural networks, neurons in the same layer tend to condense into groups with similar outputs.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Kenji Fukumizu and Shun-ichi Amari · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
The nature of statistical learning theory
Vladimir Vapnik · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Earlier work this paper cites.
Parameters as interacting particles: long time convergence and asymptotic error scaling of neural networks
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Earlier work this paper cites.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Earlier work this paper cites.
Why do larger models generalize better? a theoretical perspective via the xor problem
Alon Brutzkus and Amir Globerson · 2019
Earlier work this paper cites.
A phase shift deep neural network for high frequency wave equations in inhomogeneous media
Wei Cai, Xiaoguang Li, and Lizuo Liu · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Semi-flat minima and saddle points by embedding neural networks to overparameterization
Kenji Fukumizu, Shoichiro Yamaguchi, Yoh-ichi Mototake, and Mirai Tanaka · 2019
Earlier work this paper cites.
Gradient descent aligns the layers of deep linear networks
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Multi-scale deep neural network (mscalednn) for solving poisson-boltzmann equation in complex domains
Ziqi Liu, Wei Cai, and Zhi-Qin John Xu · 2020
Cited alongside, same era.
A multi-scale dnn algorithm for nonlinear elliptic equations with multiple scales
Xi-An Li, Zhi-Qin John Xu, and Lei Zhang · 2020
Cited alongside, same era.
Mean field analysis of neural networks: A central limit theorem
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Cited alongside, same era.
Learning a neuron by a shallow relu network: Dynamics and implicit bias for correlated inputs
Dmitry Chistikov, Matthias Englert, and Ranko Lazic · 2023
Later among the works it cites.
Phase diagram of initial condensation for two-layer neural networks
Zhengan Chen, Yuqing Li, Tao Luo, Zhangchen Zhou, and Zhi-Qin John Xu · 2023
Later among the works it cites.
Regression as classification: Influence of task formulation on neural network features
Lawrence Stewart, Francis Bach, Quentin Berthet, and Jean-Philippe Vert · 2023
Later among the works it cites.
Understanding the initial condensation of convolutional neural networks
Zhangchen Zhou, Hanxu Zhou, Yuqing Li, and Zhi-Qin John Xu · 2023
Later among the works it cites.
Optimistic estimate uncovers the potential of nonlinear models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fourier features let networks learn high frequency functions in low dimensional domains
Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng · 2020
Cited alongside, same era.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2021
Cited alongside, same era.
The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima
Yu Feng and Yuhai Tu · 2021
Cited alongside, same era.
Gradient descent on two-layer nets: Margin maximization and simplicity bias
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang, and Sanjeev Arora · 2021
Cited alongside, same era.
An upper limit of decaying rate with respect to frequency in deep neural network
Tao Luo, Zheng Ma, Zhiwei Wang, Zhi-Qin John Xu, and Yaoyu Zhang · 2021
Cited alongside, same era.
Phase diagram for two-layer relu neural networks at infinite-width limit
Tao Luo, Zhi-Qin John Xu, Zheng Ma, and Yaoyu Zhang · 2021
Cited alongside, same era.
Yaoyu Zhang, Zhongwang Zhang, Leyang Zhang, Zhiwei Bai, Tao Luo, and Zhi-Qin John Xu · 2023
Later among the works it cites.
Early alignment in two-layer networks training is a two-edged sword
Etienne Boursier and Nicolas Flammarion · 2024
Later among the works it cites.
On the dynamics of three-layer neural networks: initial condensation
Zheng-an Chen and Tao Luo · 2024
Later among the works it cites.
Analyzing multi-stage loss curve: Plateau and descent mechanisms in neural networks
Zheng-An Chen, Tao Luo, and GuiHong Wang · 2024
Later among the works it cites.
Efficient and flexible method for reducing moderate-size deep neural networks with condensation
Tianyi Chen and Zhi-Qin John Xu · 2024
Later among the works it cites.
Directional convergence near small initializations and saddles in two-homogeneous neural networks
Akshay Kumar and Jarvis Haupt · 2024
Later among the works it cites.
Early directional convergence in deep homogeneous neural networks for small initializations
Akshay Kumar and Jarvis Haupt · 2024
Later among the works it cites.
Early neuron alignment in two-layer relu networks with small initialization
Hancheng Min, Enrique Mallada, and Rene Vidal · 2024
Later among the works it cites.
Understanding multi-phase optimization dynamics and rich nonlinear behaviors of relu networks
Mingze Wang and Chao Ma · 2024
Later among the works it cites.
Overview frequency principle/spectral bias in deep learning
Zhi-Qin John Xu, Yaoyu Zhang, and Tao Luo · 2024
Later among the works it cites.
Stochastic modified equations and dynamics of dropout algorithm
Zhongwang Zhang, yuqing Li, Tao Luo, and Zhi-Qin John Xu · 2024
Later among the works it cites.
Initialization is critical to whether transformers fit composite functions by reasoning or memorizing
Zhongwang Zhang, Pengxiao Lin, Zhiwei Wang, Yaoyu Zhang, and Zhi-Qin John Xu · 2024
Later among the works it cites.
Implicit regularization of dropout
Zhongwang Zhang and Zhi-Qin John Xu · 2024
Later among the works it cites.
An analysis for reasoning bias of language models with small initialization
Junjie Yao, Zhongwang Zhang, and Zhi-Qin John Xu · 2025
Closest in time.
Complexity control facilitates reasoning-based compositional generalization in transformers
Zhongwang Zhang, Pengxiao Lin, Zhiwei Wang, Yaoyu Zhang, and Zhi-Qin John Xu · 2025
Closest in time.