Fetching the paper…
Reading the bibliography…
In this paper, we conjecture that if the permutation invariance of neural networks is taken into account, SGD solutions will likely have no barrier in the linear interpolation between them.
Brea, J., Simsek, B., Illing, B., and Gerstner, W. (2019a) · 1907
Earlier work this paper cites.
Implicit regularization and convergence for weight normalization
Wu, X., Dobriban, E., Ren, T., Wu, S., Li, Z., Gunasekar, S., Ward, R., and Liu, Q. (2019) · 1911
Earlier work this paper cites.
Deep ensembles: A loss landscape perspective
Fort, S., Hu, H., and Lakshminarayanan, B. (2019) · 1912
Earlier work this paper cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I. (2019) · 1912
Earlier work this paper cites.
Principles of neurodynamics. perceptrons and the theory of brain mechanisms
Rosenblatt, F. (1961) · 1961
Earlier work this paper cites.
Local minima and plateaus in hierarchical structures of multilayer perceptrons
Fukumizu, K. and Amari, S. (2000) · 2000
Earlier work this paper cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Liu, C., Zhu, L., and Belkin, M. (2020) · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Cifar-100 and cifar-10 (canadian institute for advanced research)
Krizhevsky, A., Nair, V., and Hinton, G. (2009) · 2009
Earlier work this paper cites.
Optimizing mode connectivity via neuron alignment
Tatro, N. J., Chen, P.-Y., Das, P., Melnyk, I., Sattigeri, P., and Lai, R. (2020) · 2009
Earlier work this paper cites.
MNIST handwritten digit database
LeCun, Y. and Cortes, C. (2010) · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Earlier work this paper cites.
Path-SGD: Path-normalized optimization in deep neural networks
Neyshabur, B., Salakhutdinov, R., and Srebro, N. (2015) · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2015) · 2015
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J. (2016) · 2016
Cited alongside, same era.
List-based simulated annealing algorithm for traveling salesman problem
Zhan, S.-h., Lin, J., Zhang, Z.-j., and Zhong, Y.-w. (2016) · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2017) · 2017
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. (2017) · 2017
Cited alongside, same era.
Large scale structure of neural network loss landscapes
Fort, S. and Jastrzebski, S. (2019) · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M. (2019) · 2019
Later among the works it cites.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
Geiger, M., Spigler, S., d’Ascoli, S., Sagun, L., Baity-Jesi, M., Biroli, G., and Wyart, M. (2019) · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
Nagarajan, V. and Kolter, J. Z. (2019) · 2019
Later among the works it cites.
Optimal convergence rates for convex distributed optimization in networks
Scaman, K., Bach, F., Bubeck, S., Lee, Y., and Massoulié, L. (2019) · 2019
Later among the works it cites.
Shaping the learning landscape in neural networks around wide flat minima
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N. (2017) · 2017
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. (2018) · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of DNNs
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D., and Wilson, A. G. (2018) · 2018
Cited alongside, same era.
Multi-task zipping via layer-wise neuron sharing
He, X., Zhou, Z., and Thiele, L. (2018) · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C. (2018) · 2018
Cited alongside, same era.
Over-parameterized deep neural networks have no strict local minima for any continuous activations
Li, D., Ding, T., and Sun, R. (2018) · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Mei, S., Montanari, A., and Nguyen, P.-M. (2018) · 2018
Cited alongside, same era.
Baldassi, C., Pittorino, F., and Zecchina, R. (2020) · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M. (2020) · 2020
Later among the works it cites.
Towards learning convolutions from scratch
Neyshabur, B. (2020) · 2020
Later among the works it cites.
What is being transferred in transfer learning?
Neyshabur, B., Sedghi, H., and Zhang, C. (2020) · 2020
Later among the works it cites.
Caliban: Docker-based job manager for reproducible workflows
Ritchie, S., Slone, A., and Ramasesh, V. (2020) · 2020
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2017) · 2020
Later among the works it cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Şimşek, B., Ged, F., Jacot, A., Spadaro, F., Hongler, C., Gerstner, W., and Brea, J. (2021) · 2021
Closest in time.
Learning neural network subspaces
Wortsman, M., Horton, M., Guestrin, C., Farhadi, A., and Rastegari, M. (2021) · 2021
Closest in time.