Fetching the paper…
Reading the bibliography…
Model fusion aims to integrate several deep neural network (DNN) models' knowledge into one by fusing parameters, and it has promising applications, such as improving the generalization of foundation models and parameter averaging in federated learning.
R. Hecht-Nielsen, “On the algebraic structure of feedforward network weight spaces,” in Advanced Neural Computers . Elsevier, 1990, pp. 129–135
1990
Earlier work this paper cites.
R. K. Pace and R. Barry, “Sparse spatial autoregressions,” Statistics & Probability Letters , vol. 33, no. 3, pp. 291–297, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
A. Krizhevsky et al. , “Learning multiple layers of features from tiny images,” Citeseer , 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” 2009 IEEE conference on computer vision and pattern recognition , pp. 248–255, 2009
2009
Earlier work this paper cites.
Y. LeCun and C. Cortes, “MNIST handwritten digit database,” arxiv , 2010. [Online]. Available: http://yann.lecun.com/exdb/mnist/
2010
Earlier work this paper cites.
A. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts, “Learning word vectors for sentiment analysis,” in Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies , 2011, pp. 142–150
2011
Earlier work this paper cites.
A. Graves and A. Graves, “Long short-term memory,” Supervised sequence labelling with recurrent neural networks , pp. 37–45, 2012
2012
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Artificial intelligence and statistics . PMLR, 2017, pp. 1273–1282
2017
Earlier work this paper cites.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” in Proceedings of the 5th International Conference on Learning Representations (ICLR) , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
F. Draxler, K. Veschgini, M. Salmhofer, and F. Hamprecht, “Essentially no barriers in neural network energy landscape,” in International conference on machine learning . PMLR, 2018, pp. 1309–1318
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson, “Loss surfaces, mode connectivity, and fast ensembling of dnns,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
Y. Lin, S. Han, H. Mao, Y. Wang, and B. Dally, “Deep gradient compression: Reducing the communication bandwidth for distributed training,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein, “Visualizing the loss landscape of neural nets,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
M. Yurochkin, M. Agarwal, S. Ghosh, K. Greenewald, N. Hoang, and Y. Khazaeni, “Bayesian nonparametric federated learning of neural networks,” in International conference on machine learning . PMLR, 2019, pp. 7252–7261
2019
Earlier work this paper cites.
T. Vogels, S. P. Karimireddy, and M. Jaggi, “Powersgd: Practical low-rank gradient compression for distributed optimization,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. von Oswald, C. Henning, B. F. Grewe, and J. Sacramento, “Continual learning with hypernetworks,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI Blog , vol. 1, no. 8, 2019
2019
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “Glue: A multi-task benchmark and analysis platform for natural language understanding,” in Proceedings of the International Conference on Learning Representations (ICLR) , 2019. [Online]. Available: https://openreview.net/forum?id=rJ4km2R5t7
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
J. Frankle, G. K. Dziugaite, D. Roy, and M. Carbin, “Linear mode connectivity and the lottery ticket hypothesis,” in International Conference on Machine Learning . PMLR, 2020, pp. 3259–3269
2020
Cited alongside, same era.
S. P. Singh and M. Jaggi, “Model fusion via optimal transport,” Advances in Neural Information Processing Systems , vol. 33, pp. 22 045–22 055, 2020
2020
Cited alongside, same era.
D. A. E. Acar, Y. Zhao, R. Matas, M. Mattina, P. Whatmough, and V. Saligrama, “Federated learning based on dynamic regularization,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
H.-Y. Chen and W.-L. Chao, “On bridging generic and personalized federated learning for image classification,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Don-Yehiya, E. Venezian, C. Raffel, N. Slonim, and L. Choshen, “Cold fusion: Collaborative descent for distributed multitask finetuning,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , A. Rogers, J. L. Boyd-Graber, and N. Okazaki, Eds. Association for Computational Linguistics, 2023, pp. 788–806. [Online]. Available: https://doi.org/10.18653/v1/2023.acl-long.46
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Zhao, P.-Y. Chen, P. Das, K. N. Ramamurthy, and X. Lin, “Bridging mode connectivity in loss landscapes and adversarial robustness,” in International Conference on Learning Representations (ICLR 2020) , 2020
2020
Cited alongside, same era.
H. Wang, M. Yurochkin, Y. Sun, D. Papailiopoulos, and Y. Khazaeni, “Federated learning with matched averaging,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
N. Tatro, P.-Y. Chen, P. Das, I. Melnyk, P. Sattigeri, and R. Lai, “Optimizing mode connectivity via neuron alignment,” Advances in Neural Information Processing Systems , vol. 33, pp. 15 300–15 311, 2020
2020
Cited alongside, same era.
T. Lin, L. Kong, S. U. Stich, and M. Jaggi, “Ensemble distillation for robust model fusion in federated learning,” Advances in Neural Information Processing Systems , vol. 33, pp. 2351–2363, 2020
2020
Cited alongside, same era.
J. Wang, Q. Liu, H. Liang, G. Joshi, and H. V. Poor, “Tackling the objective inconsistency problem in heterogeneous federated optimization,” Advances in neural information processing systems , vol. 33, pp. 7611–7623, 2020
2020
Cited alongside, same era.
S. P. Karimireddy, S. Kale, M. Mohri, S. Reddi, S. Stich, and A. T. Suresh, “Scaffold: Stochastic controlled averaging for federated learning,” in International Conference on Machine Learning . PMLR, 2020, pp. 5132–5143
2020
Cited alongside, same era.
T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,” Proceedings of Machine Learning and Systems , vol. 2, pp. 429–450, 2020
2020
Cited alongside, same era.
X. Li, M. JIANG, X. Zhang, M. Kamp, and Q. Dou, “Fedbn: Federated learning on non-iid features via local batch normalization,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
2023
Later among the works it cites.
Z. Li, T. Lin, X. Shang, and C. Wu, “Revisiting weighted aggregation in federated learning with neural networks,” in Proceedings of the 40th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 202. PMLR, 23–29 Jul 2023, pp. 19 767–19 788
2023
Later among the works it cites.
P. Yadav, D. Tam, L. Choshen, C. A. Raffel, and M. Bansal, “Ties-merging: Resolving interference when merging models,” Advances in Neural Information Processing Systems , vol. 36, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Yao, P. Wang, B. Tian, S. Cheng, Z. Li, S. Deng, H. Chen, and N. Zhang, “Editing large language models: Problems, methods, and opportunities,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 10 222–10 240
2023
Later among the works it cites.
Z. Li, X. Shang, R. He, T. Lin, and C. Wu, “No fear of classifier biases: Neural collapse inspired federated learning with synthetic and fixed classifier,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 5319–5329
2023
Later among the works it cites.
G. Ilharco, M. T. Ribeiro, M. Wortsman, L. Schmidt, H. Hajishirzi, and A. Farhadi, “Editing models with task arithmetic,” in The Eleventh International Conference on Learning Representations , 2023
2023
Later among the works it cites.
G. Ortiz-Jimenez, A. Favero, and P. Frossard, “Task arithmetic in the tangent space: Improved editing of pre-trained models,” Advances in Neural Information Processing Systems , vol. 36, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
E. S. Lubana, E. J. Bigelow, R. P. Dick, D. Krueger, and H. Tanaka, “Mechanistic mode connectivity,” in International Conference on Machine Learning . PMLR, 2023, pp. 22 965–23 004
2023
Later among the works it cites.
F. A. G. Peña, H. R. Medeiros, T. Dubail, M. Aminbeidokhti, E. Granger, and M. Pedersoli, “Re-basin via implicit sinkhorn differentiation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 20 237–20 246
2023
Later among the works it cites.
R. Ye, M. Xu, J. Wang, C. Xu, S. Chen, and Y. Wang, “Feddisco: Federated learning with discrepancy-aware collaboration,” in International Conference on Machine Learning . PMLR, 2023, pp. 39 879–39 902
2023
Later among the works it cites.
2023
Later among the works it cites.
“Tiny imagenet,” https://tiny-imagenet.herokuapp.com/ , Accessed: 2023, a subset of the ImageNet dataset
2023
Later among the works it cites.
Y. Dai, Z. Chen, J. Li, S. Heinecke, L. Sun, and R. Xu, “Tackling data heterogeneity in federated learning with class prototypes,” 2023
2023
Later among the works it cites.
OpenAI and the co authors, “Gpt-4 technical report,” 2024
2024
Closest in time.
OpenAI, “Creating video from text,” https://openai.com/sora , February 15 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Choi, H. Lee, G. Nam, and J. Lee, “Sparse weight averaging with multiple parti-cles for iterative magnitude pruning,” ICLR 2024
2024
Closest in time.