Fetching the paper…
Reading the bibliography…
Averaging neural network parameters is an intuitive method for fusing the knowledge of two independent models.
Robustness and generalization
Huan Xu and Shie Mannor · 2012
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2018
Earlier work this paper cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Earlier work this paper cites.
Federated learning with personalization layers
Manoj Ghuhan Arivazhagan, Vinay Aggarwal, Aaditya Kumar Singh, and Sunav Choudhary · 2019
Earlier work this paper cites.
Deep ensembles: A loss landscape perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan · 2019
Earlier work this paper cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Earlier work this paper cites.
Moment matching for multi-source domain adaptation
Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang · 2019
Earlier work this paper cites.
The intriguing role of module criticality in the generalization of deep networks
Niladri S Chatterji, Behnam Neyshabur, and Hanie Sedghi · 2020
Earlier work this paper cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Earlier work this paper cites.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Earlier work this paper cites.
Think locally, act globally: Federated learning with local and global representations
Paul Pu Liang, Terrance Liu, Liu Ziyin, Nicholas B Allen, Randy P Auerbach, David Brent, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2020
Earlier work this paper cites.
Bad global minima exist and sgd can reach them
Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas · 2020
Earlier work this paper cites.
Model fusion via optimal transport
Sidak Pal Singh and Martin Jaggi · 2020
Cited alongside, same era.
Federated learning with matched averaging
Hongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris Papailiopoulos, and Yasaman Khazaeni · 2020
Cited alongside, same era.
Revisiting model stitching to compare neural representations
Yamini Bansal, Preetum Nakkiran, and Boaz Barak · 2021
Cited alongside, same era.
Fedbn: Federated learning on non-iid features via local batch normalization
Xiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp, and Qi Dou · 2021
Cited alongside, same era.
Fedbabu: Toward enhanced representation for federated image classification
Jaehoon Oh, SangMook Kim, and Se-Young Yun · 2021
Cited alongside, same era.
Relative flatness and generalization
Henning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu, and Mario Boley · 2021
Cited alongside, same era.
Federated learning with partial model personalization
Krishna Pillutla, Kshitiz Malik, Abdel-Rahman Mohamed, Mike Rabbat, Maziar Sanjabi, and Lin Xiao · 2022
Later among the works it cites.
What can linear interpolation of neural network loss landscapes tell us?
Tiffany J Vlaar and Jonathan Frankle · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Later among the works it cites.
On convexity and linear mode connectivity in neural networks
David Yunis, Kumar Kshitij Patel, Pedro Henrique Pamplona Savarese, Gal Vardi, Jonathan Frankle, Matthew Walter, Karen Livescu, and Michael Maire · 2022
Later among the works it cites.
Are all layers created equal?
Chiyuan Zhang, Samy Bengio, and Yoram Singer · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The effects of mild over-parameterization on the optimization landscape of shallow relu neural networks
Itay M Safran, Gilad Yehudai, and Ohad Shamir · 2021
Cited alongside, same era.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances
Berfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro, Clément Hongler, Wulfram Gerstner, and Johanni Brea · 2021
Cited alongside, same era.
Learning neural network subspaces
Mitchell Wortsman, Maxwell C Horton, Carlos Guestrin, Ali Farhadi, and Mohammad Rastegari · 2021
Cited alongside, same era.
Git re-basin: Merging models modulo permutation symmetries
Samuel Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa · 2022
Cited alongside, same era.
Xiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen, Shuyang Gu, Weiming Zhang, Lu Yuan, Dong Chen, Fang Wen, and Nenghai Yu · 2022
Cited alongside, same era.
The role of permutation invariance in linear mode connectivity of neural networks
Rahim Entezari, Hanie Sedghi, Olga Saukh, and Behnam Neyshabur · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Closest in time.
Bottleneck structure in learned features: Low-dimension vs regularity tradeoff
Arthur Jacot · 2023
Closest in time.
Mechanistic mode connectivity
Ekdeep Singh Lubana, Eric J Bigelow, Robert P Dick, David Krueger, and Hidenori Tanaka · 2023
Closest in time.
Can neural network memorization be localized?
Pratyush Maini, Michael C Mozer, Hanie Sedghi, Zachary C Lipton, J Zico Kolter, and Chiyuan Zhang · 2023
Closest in time.
Normalization layers are all that sharpness-aware minimization needs
Maximilian Mueller, Tiffany Vlaar, David Rolnick, and Matthias Hein · 2023
Closest in time.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard · 2023
Closest in time.
Revisiting adapters with adversarial training
Sylvestre-Alvise Rebuffi, Francesco Croce, and Sven Gowal · 2023
Closest in time.
Going beyond linear mode connectivity: The layerwise linear feature connectivity
Zhanpeng Zhou, Yongyi Yang, Xiaojiang Yang, Junchi Yan, and Wei Hu · 2023
Closest in time.
Decentralized sgd and average-direction sam are asymptotically equivalent
Tongtian Zhu, Fengxiang He, Kaixuan Chen, Mingli Song, and Dacheng Tao · 2023
Closest in time.
Stochastic collapse: How gradient noise attracts sgd dynamics towards simpler subnetworks
Feng Chen, Daniel Kunin, Atsushi Yamamura, and Surya Ganguli · 2024
Closest in time.