Fetching the paper…
Reading the bibliography…
We systematize the approach to the investigation of deep neural network landscapes by basing it on the geometry of the space of implemented functions rather than the space of parameters.
Brea, J., Simsek, B., Illing, B., and Gerstner, W · 1907
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Hinton, G. E. and Van Camp, D · 1993
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Nudged elastic band method for finding minimum energy paths of transitions
Jónsson, H., Mills, G., and Jacobsen, K. W · 1998
Earlier work this paper cites.
Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models
Rocks, J. W. and Mehta, P · 2010
Earlier work this paper cites.
Subdominant dense clusters allow for simple learning and high computational performance in neural networks with discrete synapses
Baldassi, C., Ingrosso, A., Lucibello, C., Saglietti, L., and Zecchina, R · 2015
Earlier work this paper cites.
Accelerating very deep convolutional networks for classification and detection
Zhang, X., Zou, J., He, K., and Sun, J · 2015
Earlier work this paper cites.
Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes
Baldassi, C., Borgs, C., Chayes, J. T., Ingrosso, A., Lucibello, C., Saglietti, L., and Zecchina, R · 2016
Earlier work this paper cites.
Binarized neural networks
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
Dziugaite, G. K. and Roy, D. M · 2017
Earlier work this paper cites.
Role of synaptic stochasticity in training low-precision neural networks
Baldassi, C., Gerace, F., Kappen, H. J., Lucibello, C., Saglietti, L., Tartaglione, E., and Zecchina, R · 2018
Earlier work this paper cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Chaudhari, P. and Soatto, S · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Cited alongside, same era.
Using mode connectivity for loss landscape analysis
Gotmare, A., Keskar, N. S., Xiong, C., and Socher, R · 2018
Cited alongside, same era.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2018
Cited alongside, same era.
Properties of the geometry of solutions and capacity of multilayer neural networks with rectified linear unit activations
Baldassi, C., Malatesta, E. M., and Zecchina, R · 2019
Cited alongside, same era.
Bad global minima exist and sgd can reach them
Liu, S., Papailiopoulos, D., and Achlioptas, D · 2020
Later among the works it cites.
Landscape connectivity and dropout stability of SGD solutions for over-parameterized neural networks
Shevchenko, A. and Mondelli, M · 2020
Later among the works it cites.
Model fusion via optimal transport
Singh, S. P. and Jaggi, M · 2020
Later among the works it cites.
Optimizing mode connectivity via neuron alignment
Tatro, N., Chen, P.-Y., Das, P., Melnyk, I., Sattigeri, P., and Lai, R · 2020
Later among the works it cites.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2019
Cited alongside, same era.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Kuditipudi, R., Wang, X., Lee, H., Zhang, Y., Li, Z., Hu, W., Arora, S., and Ge, R · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Mei, S. and Montanari, A · 2019
Cited alongside, same era.
A review of binarized neural networks
Simons, T. and Lee, D.-J · 2019
Cited alongside, same era.
Shaping the learning landscape in neural networks around wide flat minima
Baldassi, C., Pittorino, F., and Zecchina, R · 2020
Cited alongside, same era.
Generalisation error in learning with random features and the hidden manifold model
Gerace, F., Loureiro, B., Krzakala, F., Mézard, M., and Zdeborová, L · 2020
Cited alongside, same era.
Salr: Sharpness-aware learning rates for improved generalization
Yue, X., Nouiehed, M., and Kontar, R. A · 2020
Later among the works it cites.
Bridging mode connectivity in loss landscapes and adversarial robustness
Zhao, P., Chen, P.-Y., Das, P., Ramamurthy, K. N., and Lin, X · 2020
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B · 2021
Later among the works it cites.
The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima
Feng, Y. and Tu, Y · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Later among the works it cites.
Entropic gradient descent algorithms and wide flat minima
Pittorino, F., Lucibello, C., Feinauer, C., Perugini, G., Baldassi, C., Demyanenko, E., and Zecchina, R · 2021
Later among the works it cites.
Basins with tentacles
Zhang, Y. and Strogatz, S. H · 2021
Later among the works it cites.
A variance principle explains why dropout finds flatter minima
Zhang, Z., Zhou, H., and Xu, Z.-Q. J · 2021
Later among the works it cites.