Fetching the paper…
Reading the bibliography…
The energy landscape of high-dimensional non-convex optimization problems is crucial to understanding the effectiveness of modern deep neural network architectures.
Optimal transport: old and new , volume 338
C. Villani et al · 2009
Earlier work this paper cites.
Chernoff-hoeffding inequality and applications
J. M. Phillips · 2012
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2014
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C. D. Freeman and J. Bruna · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
F. Draxler, K. Veschgini, M. Salmhofer, and F. Hamprecht · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
Using mode connectivity for loss landscape analysis
A. Gotmare, N. S. Keskar, C. Xiong, and R. Socher · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
S. Mei, A. Montanari, and P.-M. Nguyen · 2018
Earlier work this paper cites.
Going beyond linear mode connectivity: The layerwise linear feature connectivity
Z. Zhou, Y. Yang, X. Yang, J. Yan, and W. Hu · 2018
Earlier work this paper cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
R. Kuditipudi, X. Wang, H. Lee, Y. Zhang, Z. Li, W. Hu, R. Ge, and S. Arora · 2019
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
S. Mei, T. Misiakiewicz, and A. Montanari · 2019
Cited alongside, same era.
Computational optimal transport: With applications to data science
G. Peyré, M. Cuturi, et al · 2019
Cited alongside, same era.
Spurious valleys in one-hidden-layer neural network optimization landscapes
L. Venturi, A. S. Bandeira, and J. Bruna · 2019
Cited alongside, same era.
Sharp asymptotic and finite-sample rates of convergence of empirical measures in wasserstein distance
J. Weed and F. Bach · 2019
Cited alongside, same era.
Bayesian nonparametric federated learning of neural networks
M. Yurochkin, M. Agarwal, S. Ghosh, K. Greenewald, N. Hoang, and Y. Khazaeni · 2019
Cited alongside, same era.
Linear mode connectivity and the lottery ticket hypothesis
Analyzing monotonic linear interpolation in neural network loss landscapes
J. Lucas, J. Bae, M. R. Zhang, S. Fort, R. Zemel, and R. Grosse · 2021
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
S. K. Ainsworth, J. Hayase, and S. Srinivasa · 2022
Later among the works it cites.
Wasserstein barycenter-based model fusion and linear mode connectivity of neural networks
A. K. Akash, S. Li, and N. G. Trillos · 2022
Later among the works it cites.
A general framework for proving the equivariant strong lottery ticket hypothesis
D. Ferbach, C. Tsirigotis, G. Gidel, and A. Bose · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Frankle, G. K. Dziugaite, D. Roy, and M. Carbin · 2020
Cited alongside, same era.
Proving the lottery ticket hypothesis: Pruning is all you need
E. Malach, G. Yehudai, S. Shalev-Schwartz, and O. Shamir · 2020
Cited alongside, same era.
Linear mode connectivity in multitask and continual learning
S. I. Mirzadeh, M. Farajtabar, D. Gorur, R. Pascanu, and H. Ghasemzadeh · 2020
Cited alongside, same era.
What is being transferred in transfer learning?
B. Neyshabur, H. Sedghi, and C. Zhang · 2020
Cited alongside, same era.
Optimal lottery tickets via subset sum: Logarithmic over-parameterization is sufficient
A. Pensia, S. Rajput, A. Nagle, H. Vishwakarma, and D. Papailiopoulos · 2020
Cited alongside, same era.
Landscape connectivity and dropout stability of sgd solutions for over-parameterized neural networks
A. Shevchenko and M. Mondelli · 2020
Cited alongside, same era.
Model fusion via optimal transport
S. P. Singh and M. Jaggi · 2020
Cited alongside, same era.
J. Juneja, R. Bansal, K. Cho, J. Sedoc, and N. Saphra · 2022
Later among the works it cites.
Deep networks on toroids: removing symmetries reveals the structure of flat regions in the landscape geometry
F. Pittorino, A. Ferraro, G. Perugini, C. Feinauer, C. Baldassi, and R. Zecchina · 2022
Later among the works it cites.
Exploring mode connectivity for pre-trained language models
Y. Qin, C. Qian, J. Yi, W. Chen, Y. Lin, X. Han, Z. Liu, M. Sun, and J. Zhou · 2022
Later among the works it cites.
Diverse weight averaging for out-of-distribution generalization
A. Rame, M. Kirchmeyer, T. Rahier, A. Rakotomamonjy, P. Gallinari, and M. Cord · 2022
Later among the works it cites.
What can linear interpolation of neural network loss landscapes tell us?
T. J. Vlaar and J. Frankle · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, et al · 2022
Later among the works it cites.
On convexity and linear mode connectivity in neural networks
D. Yunis, K. K. Patel, P. H. P. Savarese, G. Vardi, J. Frankle, M. Walter, K. Livescu, and M. Maire · 2022
Later among the works it cites.
Layerwise linear mode connectivity
L. Adilova, A. Fischer, and M. Jaggi · 2023
Closest in time.
Mechanistic mode connectivity
E. S. Lubana, E. J. Bigelow, R. P. Dick, D. Krueger, and H. Tanaka · 2023
Closest in time.
A rigorous framework for the mean field limit of multilayer neural networks
P.-M. Nguyen and H. T. Pham · 2023
Closest in time.
Packing, covering, and consequences on minimax risk
Y. Wu · 2023
Closest in time.