Fetching the paper…
Reading the bibliography…
Empirical studies of the loss landscape of deep networks have revealed that many local minima are connected through low-loss valleys.
Invariante variationsprobleme
Emmy Noether · 1918
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Statistics of critical points of gaussian fields on large-dimensional spaces
Alan J Bray and David S Dean · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Symmetry-invariant optimization in deep networks
Vijay Badrinarayanan, Bamdev Mishra, and Roberto Cipolla · 2015
Earlier work this paper cites.
Path-SGD: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Russ R Salakhutdinov, and Nati Srebro · 2015
Earlier work this paper cites.
On accelerated methods in optimization
Andre Wibisono and Ashia C Wilson · 2015
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
L2 regularization versus batch and weight normalization
Twan Van Laarhoven · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Earlier work this paper cites.
Improving optimization for models with continuous symmetry breaking
Robert Bamler and Stephan Mandt · 2018
Earlier work this paper cites.
The loss landscape of overparameterized neural networks
Yaim Cooper · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Cited alongside, same era.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Cited alongside, same era.
Using mode connectivity for loss landscape analysis
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
Noether: The more things change, the more stay the same
Grzegorz Głuch and Rüdiger Urbanke · 2021
Later among the works it cites.
Neural mechanics: Symmetry and broken conservation laws in deep learning dynamics
Daniel Kunin, Javier Sagastuy-Brena, Surya Ganguli, Daniel LK Yamins, and Hidenori Tanaka · 2021
Later among the works it cites.
Optimization and sampling under continuous symmetry: Examples and lie theory
Jonathan Leake and Nisheeth K Vishnoi · 2021
Later among the works it cites.
On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks
Hancheng Min, Salma Tarmoun, René Vidal, and Enrique Mallada · 2021
Later among the works it cites.
Relative flatness and generalization
Henning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu, and Mario Boley · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 2019
Cited alongside, same era.
𝒢 \mathcal{G} -SGD: Optimizing relu neural networks in its positively scale-invariant space
Qi Meng, Shuxin Zheng, Huishuai Zhang, Wei Chen, Zhi-Ming Ma, and Tie-Yan Liu · 2019
Cited alongside, same era.
Directional pruning of deep neural networks
Shih-Kang Chao, Zhanyu Wang, Yue Xing, and Guang Cheng · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Cited alongside, same era.
Curvature-corrected learning dynamics in deep neural networks
Dongsung Huh · 2020
Cited alongside, same era.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Berfin Şimşek, François Ged, Arthur Jacot, Francesco Spadaro, Clément Hongler, Wulfram Gerstner, and Johanni Brea · 2021
Later among the works it cites.
Noether’s learning dynamics: Role of symmetry breaking in neural networks
Hidenori Tanaka and Daniel Kunin · 2021
Later among the works it cites.
Understanding the dynamics of gradient flow in overparameterized linear models
Salma Tarmoun, Guilherme Franca, Benjamin D Haeffele, and Rene Vidal · 2021
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Rahim Entezari, Hanie Sedghi, Olga Saukh, and Behnam Neyshabur · 2022
Closest in time.
Universal approximation and model compression for radial neural networks
Iordan Ganev, Twan van Laarhoven, and Robin Walters · 2022
Closest in time.
Fisher sam: Information geometry and sharpness aware minimisation
Minyoung Kim, Da Li, Shell X Hu, and Timothy Hospedales · 2022
Closest in time.
Deep networks on toroids: Removing symmetries reveals the structure of flat regions in the landscape geometry
Fabrizio Pittorino, Antonio Ferraro, Gabriele Perugini, Christoph Feinauer, Carlo Baldassi, and Riccardo Zecchina · 2022
Closest in time.
Symmetry teleportation for accelerated optimization
Bo Zhao, Nima Dehmamy, Robin Walters, and Rose Yu · 2022
Closest in time.
Git re-basin: Merging models modulo permutation symmetries
Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa · 2023
Closest in time.
Neural teleportation
Marco Armenta, Thierry Judge, Nathan Painchaud, Youssef Skandarani, Carl Lemaire, Gabriel Gibeau Sanchez, Philippe Spino, and Pierre-Marc Jodoin · 2023
Closest in time.