Fetching the paper…
Reading the bibliography…
Neural networks typically exhibit permutation symmetries which contribute to the non-convexity of the networks' loss landscapes, since linearly interpolating between two permuted versions of a trained network tends to encounter a high loss barrier.
Krizhevsky, A.: Learning multiple layers of features from tiny images (2009), https://www.cs.toronto.edu/~kriz/learning-features-2009-TR.pdf
2009
Earlier work this paper cites.
Tange, O.: Gnu parallel - the command-line power tool. ;login: The USENIX Magazine 36
2011
Earlier work this paper cites.
Goodfellow, I.J., Vinyals, O., Saxe, A.M.: Qualitatively characterizing neural network optimization problems (2015). https://doi.org/10.48550/arXiv.1412.6544
2015
Earlier work this paper cites.
Li, Y., Yosinski, J., Clune, J., Lipson, H., Hopcroft, J.: Convergent learning: Do different neural networks learn the same representations? In: Proceedings of the 1st International Workshop on Feature Extraction: Modern Questions and Challenges at NIPS 2015. vol. 44, pp. 196–212. PMLR (2015), https://proceedings.mlr.press/v44/li15convergent.html
2015
Earlier work this paper cites.
Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization (2016). https://doi.org/10.48550/arXiv.1607.06450
2016
Earlier work this paper cites.
Alain, G., Bengio, Y.: Understanding intermediate layers using linear classifier probes (2017), https://openreview.net/forum?id=ryF7rTqgl
2017
Earlier work this paper cites.
Raghu, M., Gilmer, J., Yosinski, J., Sohl-Dickstein, J.: SVCCA: Singular vector canonical correlation analysis for deep learning dynamics and interpretability. In: Advances in Neural Information Processing Systems. vol. 30, pp. 6076–6085 (2017), https://proceedings.neurips.cc/paper_files/paper/2017/hash/dc6a7e655d7e5840e66733e9ee67cc69-Abstract.html
2017
Earlier work this paper cites.
Morcos, A., Raghu, M., Bengio, S.: Insights on representational similarity in neural networks with canonical correlation. In: Advances in Neural Information Processing Systems. vol. 31, pp. 5727–5736 (2018), https://proceedings.neurips.cc/paper/2018/hash/a7a3d70c6d17a73140918996d03c014f-Abstract.html
2018
Earlier work this paper cites.
Frankle, J., Carbin, M.: The lottery ticket hypothesis: Finding sparse, trainable neural networks. In: International Conference on Learning Representations (2019), https://openreview.net/forum?id=rJl-b3RcF7
2019
Earlier work this paper cites.
Nagarajan, V., Kolter, J.Z.: Uniform convergence may be unable to explain generalization in deep learning. In: Advances in Neural Information Processing Systems. vol. 32, pp. 11615–11626 (2019), https://proceedings.neurips.cc/paper/2019/hash/05e97c207235d63ceb1db43c60db7bbb-Abstract.html
2019
Earlier work this paper cites.
Yurochkin, M., Agarwal, M., Ghosh, S., Greenewald, K., Hoang, N., Khazaeni, Y.: Bayesian nonparametric federated learning of neural networks. In: Proceedings of the 36th International Conference on Machine Learning. vol. 97, pp. 7252–7261. PMLR (2019), https://proceedings.mlr.press/v97/yurochkin19a.html
2019
Earlier work this paper cites.
Frankle, J., Dziugaite, G.K., Roy, D., Carbin, M.: Linear mode connectivity and the lottery ticket hypothesis. In: Proceedings of the 37th International Conference on Machine Learning. vol. 119, pp. 3259–3269. PMLR (2020), https://proceedings.mlr.press/v119/frankle20a.html
2020
Cited alongside, same era.
Renda, A., Frankle, J., Carbin, M.: Comparing rewinding and fine-tuning in neural network pruning. In: International Conference on Learning Representations (2020), https://openreview.net/forum?id=S1gSj0NKvB
2020
Cited alongside, same era.
Singh, S.P., Jaggi, M.: Model fusion via optimal transport. In: Advances in Neural Information Processing Systems. vol. 33, pp. 22045–22055 (2020), https://proceedings.neurips.cc/paper/2020/hash/fb2697869f56484404c8ceee2985b01d-Abstract.html
2020
Cited alongside, same era.
Tatro, N., Chen, P.Y., Das, P., Melnyk, I., Sattigeri, P., Lai, R.: Optimizing mode connectivity via neuron alignment. In: Advances in Neural Information Processing Systems. vol. 33, pp. 15300–15311 (2020), https://proceedings.neurips.cc/paper/2020/hash/aecad42329922dfc97eee948606e1f8e-Abstract.html
Akash, A.K., Li, S., Trillos, N.G.: Wasserstein barycenter-based model fusion and linear mode connectivity of neural networks (2022). https://doi.org/10.48550/arXiv.2210.06671
2022
Later among the works it cites.
Altschuler, J.M., Boix-Adserà, E.: Wasserstein barycenters are NP-hard to compute. SIAM Journal on Mathematics of Data Science 4
2022
Later among the works it cites.
Benzing, F., Schug, S., Meier, R., Oswald, J.V., Akram, Y., Zucchet, N., Aitchison, L., Steger, A.: Random initialisations performing above chance and how to find them. In: OPT 2022: Optimization for Machine Learning (NeurIPS 2022 Workshop) (2022), https://openreview.net/forum?id=HS5zuN_qFI
2022
Later among the works it cites.
Entezari, R., Sedghi, H., Saukh, O., Neyshabur, B.: The role of permutation invariance in linear mode connectivity of neural networks. In: International Conference on Learning Representations (2022), https://openreview.net/forum?id=dNigytemkL
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Wang, H., Yurochkin, M., Sun, Y., Papailiopoulos, D., Khazaeni, Y.: Federated learning with matched averaging. In: International Conference on Learning Representations (2020), https://openreview.net/forum?id=BkluqlSFDS
2020
Cited alongside, same era.
Baldock, R., Maennel, H., Neyshabur, B.: Deep learning through the lens of example difficulty. In: Advances in Neural Information Processing Systems. vol. 34, pp. 10876–10889 (2021), https://proceedings.neurips.cc/paper_files/paper/2021/file/5a4b25aaed25c2ee1b74de72dc03c14e-Paper.pdf
2021
Cited alongside, same era.
Lucas, J.R., Bae, J., Zhang, M.R., Fort, S., Zemel, R., Grosse, R.B.: On monotonic linear interpolation of neural network parameters. In: Proceedings of the 38th International Conference on Machine Learning. vol. 139, pp. 7168–7179. PMLR (2021), https://proceedings.mlr.press/v139/lucas21a.html
2021
Cited alongside, same era.
O’Neill, J., V. Steeg, G., Galstyan, A.: Layer-wise neural network compression via layer fusion. In: Proceedings of The 13th Asian Conference on Machine Learning. vol. 157, pp. 1381–1396. PMLR (2021), https://proceedings.mlr.press/v157/o-neill21a.html
2021
Cited alongside, same era.
Simsek, B., Ged, F., Jacot, A., Spadaro, F., Hongler, C., Gerstner, W., Brea, J.: Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances. In: Proceedings of the 38th International Conference on Machine Learning. vol. 139, pp. 9722–9732. PMLR (2021), https://proceedings.mlr.press/v139/simsek21a.html
2021
Cited alongside, same era.
Wortsman, M., Horton, M.C., Guestrin, C., Farhadi, A., Rastegari, M.: Learning neural network subspaces. In: Proceedings of the 38th International Conference on Machine Learning. vol. 139, pp. 11217–11227. PMLR (2021), https://proceedings.mlr.press/v139/wortsman21a.html
2021
Cited alongside, same era.
Paul, M., Larsen, B., Ganguli, S., Frankle, J., Dziugaite, G.K.: Lottery tickets on a data diet: Finding initializations with sparse trainable networks. In: Advances in Neural Information Processing Systems. vol. 35, pp. 18916–18928 (2022), https://proceedings.neurips.cc/paper_files/paper/2022/hash/77dd8e90fe833eba5fae86cf017d7a56-Abstract-Conference.html
2022
Later among the works it cites.
Vlaar, T.J., Frankle, J.: What can linear interpolation of neural network loss landscapes tell us? In: Proceedings of the 39th International Conference on Machine Learning. vol. 162, pp. 22325–22341. PMLR (2022), https://proceedings.mlr.press/v162/vlaar22a.html
2022
Later among the works it cites.
Ainsworth, S., Hayase, J., Srinivasa, S.: Git re-basin: Merging models modulo permutation symmetries. In: The Eleventh International Conference on Learning Representations (2023), https://openreview.net/forum?id=CQsmMYmlP5T
2023
Later among the works it cites.
Jordan, K., Sedghi, H., Saukh, O., Entezari, R., Neyshabur, B.: REPAIR: REnormalizing Permuted Activations for Interpolation Repair. In: The Eleventh International Conference on Learning Representations (2023), https://openreview.net/forum?id=gU5sJ6ZggcX
2023
Later among the works it cites.
Paul, M., Chen, F., Larsen, B.W., Frankle, J., Ganguli, S., Dziugaite, G.K.: Unmasking the lottery ticket hypothesis: What’s encoded in a winning ticket’s mask? In: The Eleventh International Conference on Learning Representations (2023), https://openreview.net/forum?id=xSsW2Am-ukZ
2023
Later among the works it cites.
Peña, F.A.G., Medeiros, H.R., Dubail, T., Aminbeidokhti, M., Granger, E., Pedersoli, M.: Re-basin via implicit Sinkhorn differentiation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 20237–20246 (June 2023), https://openaccess.thecvf.com/content/CVPR2023/html/Pena_Re-Basin_via_Implicit_Sinkhorn_Differentiation_CVPR_2023_paper.html
2023
Later among the works it cites.