Fetching the paper…
Reading the bibliography…
Due to common architecture designs, symmetries exist extensively in contemporary neural networks.
Pathological spectra of the fisher information metric and its variants in deep neural networks
Karakida, R., Akaho, S., and Amari, S.-i · 1910
Earlier work this paper cites.
Products of random matrices
Furstenberg, H. and Kesten, H · 1960
Earlier work this paper cites.
More is different: Broken symmetry and the nature of the hierarchical structure of science
Anderson, P. W · 1972
Earlier work this paper cites.
A regularity condition of the information matrix of a multilayer perceptron network
Fukumizu, K · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R · 1996
Earlier work this paper cites.
Markov chains
Norris, J. R · 1998
Earlier work this paper cites.
Probabilistic principal component analysis
Tipping, M. E. and Bishop, C. M · 1999
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Fukumizu, K. and Amari, S.-i · 2000
Earlier work this paper cites.
Non-equilibrium critical phenomena and phase transitions into absorbing states
Hinrichsen, H · 2000
Earlier work this paper cites.
Quasi-stationary distributions for stochastic processes with an absorbing state
Dickman, R. and Vidigal, R · 2002
Earlier work this paper cites.
Maximum-margin matrix factorization
Srebro, N., Rennie, J., and Jaakkola, T · 2004
Earlier work this paper cites.
Pattern recognition and machine learning , volume 4
Bishop, C. M. and Nasrabadi, N. M · 2006
Earlier work this paper cites.
Dynamics of learning in multilayer perceptrons near singularities
Cousseau, F., Ozeki, T., and Amari, S.-i · 2008
Earlier work this paper cites.
The group lasso for logistic regression
Meier, L., Van De Geer, S., and Bühlmann, P · 2008
Earlier work this paper cites.
Dynamics of learning near singularities in layered networks
Wei, H., Zhang, J., Cousseau, F., Ozeki, T., and Amari, S.-i · 2008
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction , volume 2
Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H · 2009
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Jaderberg, M., Vedaldi, A., and Zisserman, A · 2014
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Group equivariant convolutional networks
Cohen, T. and Welling, M · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Sagun, L., Bottou, L., and LeCun, Y · 2016
Vicreg: Variance-invariance-covariance regularization for self-supervised learning
Bardes, A., Ponce, J., and LeCun, Y · 2021
Later among the works it cites.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Dong, Y., Cordonnier, J.-B., and Loukas, A · 2021
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B · 2021
Later among the works it cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Simsek, B., Ged, F., Jacot, A., Spadaro, F., Hongler, C., Gerstner, W., and Brea, J · 2021
Later among the works it cites.
Equivalences between sparse models and neural networks
Tibshirani, R. J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sharp Minima Can Generalize For Deep Nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Papyan, V · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., et al · 2018
Cited alongside, same era.
Negative eigenvalues of the hessian in deep neural networks
Alain, G., Roux, N. L., and Manzagol, P.-A · 2019
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
Ghorbani, B., Krishnan, S., and Xiao, Y · 2019
Cited alongside, same era.
Sgd can converge to local maxima
Ziyin, L., Li, B., Simon, J. B., and Ueda, M · 2021
Later among the works it cites.
Deep contrastive learning is provably (almost) principal component analysis
Tian, Y · 2022
Later among the works it cites.
Posterior collapse of a linear latent variable model
Wang, Z. and Ziyin, L · 2022
Later among the works it cites.
Exact phase transitions in deep learning
Ziyin, L. and Ueda, M · 2022
Later among the works it cites.
Exact solutions of a deep linear network
Ziyin, L., Li, B., and Meng, X · 2022
Later among the works it cites.
Loss of plasticity in continual deep reinforcement learning
Abbas, Z., Zhao, R., Modayil, J., White, A., and Machado, M. C · 2023
Closest in time.
Investigating how relu-networks encode symmetries, 2023
Bökman, G. and Kahl, F · 2023
Closest in time.
Stochastic collapse: How gradient noise attracts sgd dynamics towards simpler subnetworks
Chen, F., Kunin, D., Yamamura, A., and Ganguli, S · 2023
Closest in time.
Maintaining plasticity in deep continual learning
Dohare, S., Hernandez-Garcia, J. F., Rahman, P., Sutton, R. S., and Mahmood, A. R · 2023
Closest in time.
Understanding plasticity in neural networks
Lyle, C., Zheng, Z., Nikishin, E., Pires, B. A., Pascanu, R., and Dabney, W · 2023
Closest in time.
spred: Solving L1 Penalty with SGD
Ziyin, L. and Wang, Z · 2023
Closest in time.
The implicit bias of gradient noise: A symmetry perspective, 2024
Ziyin, L., Wang, M., and Wu, L · 2024
Closest in time.