Fetching the paper…
Reading the bibliography…
In this paper we look into the conjecture of Entezari et al.
Brea, J., Simsek, B., Illing, B., and Gerstner, W. (2019) · 1907
Earlier work this paper cites.
Deep ensembles: A loss landscape perspective
Fort, S., Hu, H., and Lakshminarayanan, B. (2019) · 1912
Earlier work this paper cites.
The hungarian method for the assignment problem
Kuhn, H. W. (1955) · 1955
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
Rosenblatt, F. (1958) · 1958
Earlier work this paper cites.
Relations between two sets of variates
Hotelling, H. (1992) · 1992
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y. (1998) · 1998
Earlier work this paper cites.
Federated learning with matched averaging
Wang, H., Yurochkin, M., Sun, Y., Papailiopoulos, D., and Khazaeni, Y. (2020) · 2002
Earlier work this paper cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Liu, C., Zhu, L., and Belkin, M. (2020) · 2003
Earlier work this paper cites.
Entropic gradient descent algorithms and wide flat minima
Pittorino, F., Lucibello, C., Feinauer, C., Malatesta, E. M., Perugini, G., Baldassi, C., Negri, M., Demyanenko, E., and Zecchina, R. (2020) · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Optimizing mode connectivity via neuron alignment
Tatro, N. J., Chen, P.-Y., Das, P., Melnyk, I., Sattigeri, P., and Lai, R. (2020) · 2009
Earlier work this paper cites.
Barycenters in the wasserstein space
Agueh, M. and Carlier, G. (2011) · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Earlier work this paper cites.
A method for finding similarity between multi-layer perceptrons by forward bipartite alignment
Ashmore, S. and Gashler, M. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J. (2015) · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Cited alongside, same era.
Topology and geometry of half-rectified network optimization
Freeman, C. D. and Bruna, J. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
List-based simulated annealing algorithm for traveling salesman problem
Zhan, S.-h., Lin, J., Zhang, Z.-j., and Zhong, Y.-w. (2016) · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2017) · 2017
Cited alongside, same era.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
Geiger, M., Spigler, S., d’Ascoli, S., Sagun, L., Baity-Jesi, M., Biroli, G., and Wyart, M. (2019) · 2019
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G. (2019) · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
Nagarajan, V. and Kolter, J. Z. (2019) · 2019
Later among the works it cites.
Shaping the learning landscape in neural networks around wide flat minima
Baldassi, C., Pittorino, F., and Zecchina, R. (2020) · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M. (2020) · 2020
Later among the works it cites.
Model fusion via optimal transport
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. (2017) · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N. (2017) · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R. (2017) · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2017) · 2017
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. (2018) · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G. (2018) · 2018
Cited alongside, same era.
Multi-task zipping via layer-wise neuron sharing
He, X., Zhou, Z., and Thiele, L. (2018) · 2018
Cited alongside, same era.
Singh, S. P. and Jaggi, M. (2020) · 2020
Later among the works it cites.
Safe crossover of neural networks through neuron alignment
Uriot, T. and Izzo, D. (2020) · 2020
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B. (2021) · 2021
Later among the works it cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Şimşek, B., Ged, F., Jacot, A., Spadaro, F., Hongler, C., Gerstner, W., and Brea, J. (2021) · 2021
Later among the works it cites.
Fixup initialization: Residual learning without normalization
Zhang, H., Dauphin, Y. N., and Ma, T. (2019) · 2021
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S. K., Hayase, J., and Srinivasa, S. (2022) · 2022
Closest in time.
Stochastic weight averaging revisited
Guo, H., Jin, J., and Liu, B. (2022) · 2022
Closest in time.
Patching open-vocabulary models by interpolating weights
Ilharco, G., Wortsman, M., Gadre, S. Y., Song, S., Hajishirzi, H., Kornblith, S., Farhadi, A., and Schmidt, L. (2022) · 2022
Closest in time.
Linear connectivity reveals generalization strategies
Juneja, J., Bansal, R., Cho, K., Sedoc, J., and Saphra, N. (2022) · 2022
Closest in time.
Fusing batch normalization and convolution in runtime
Markuš, N. (2018) · 2022
Closest in time.
Pittorino, F., Ferraro, A., Perugini, G., Feinauer, C., Baldassi, C., and Zecchina, R. (2022) · 2022
Closest in time.