Fetching the paper…
Reading the bibliography…
Deep model fusion/merging is an emerging technique that merges the parameters or predictions of multiple deep learning models into a single one.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Neural network ensembles
Lars Kai Hansen and Peter Salamon · 1990
Earlier work this paper cites.
On the algebraic structure of feedforward network weight spaces
Robert Hecht-Nielsen · 1990
Earlier work this paper cites.
New stochastic approximation type procedures
Boris T Polyak · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Stacked generalization
David H Wolpert · 1992
Earlier work this paper cites.
On the geometry of feedforward neural network error surfaces
An Mei Chen, Haw-minn Lu, and Robert Hecht-Nielsen · 1993
Earlier work this paper cites.
Exponentially many local minima for single neurons
Peter Auer, Mark Herbster, and Manfred KK Warmuth · 1995
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
Weight averaging for neural networks and local resampling schemes
Joachim Utans · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A brief introduction to boosting
Robert E Schapire et al · 1999
Earlier work this paper cites.
Random forests
Leo Breiman · 2001
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Measuring thermodynamic length
Gavin E Crooks · 2007
Earlier work this paper cites.
A survey for the quadratic assignment problem
Eliane Maria Loiola, Nair Maria Maia De Abreu, Paulo Oswaldo Boaventura-Netto, Peter Hahn, and Tania Querido · 2007
Earlier work this paper cites.
Hierarchical beta processes and the indian buffet process
Romain Thibaux and Michael I Jordan · 2007
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2009
Earlier work this paper cites.
Multiparty differential privacy via aggregation of locally trained classifiers
Manas Pathak, Shantanu Rane, and Bhiksha Raj · 2010
Earlier work this paper cites.
Ensemble-based classifiers
Lior Rokach · 2010
Earlier work this paper cites.
The bernstein polynomial basis: A centennial retrospective
Rida T Farouki · 2012
Earlier work this paper cites.
Computer graphics: theory and practice
Jonas Gomes, Luiz Velho, and Mario Costa Sousa · 2012
Earlier work this paper cites.
Introduction to manifold learning
Alan Julian Izenman · 2012
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, Martin J Wainwright, and John C Duchi · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Ensemble deep learning for speech recognition
Li Deng and John Platt · 2014
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2014
Earlier work this paper cites.
Learning and transferring mid-level image representations using convolutional neural networks
Maxime Oquab, Leon Bottou, Ivan Laptev, and Josef Sivic · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Weight uncertainty in neural network
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Earlier work this paper cites.
Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser
Long Duong, Trevor Cohn, Steven Bird, and Paul Cook · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Deep neural decision forests
Peter Kontschieder, Madalina Fiterau, Antonio Criminisi, and Samuel Rota Bulo · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John Hopcroft · 2015
Earlier work this paper cites.
Multi-graph matching via affinity optimization with graduated consistency regularization
Junchi Yan, Minsu Cho, Hongyuan Zha, Xiaokang Yang, and Stephen M Chu · 2015
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C. D. Freeman and J. Bruna · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
An automated nudged elastic band method
Esben L Kolsbjerg, Michael N Groves, and Bjørk Hammer · 2016
Earlier work this paper cites.
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
A chinese question answering approach integrating count-based and embedding-based features
Benyou Wang, Jiabin Niu, Liqun Ma, Yuhua Zhang, Lipeng Zhang, Jingfei Li, Peng Zhang, and Dawei Song · 2016
Earlier work this paper cites.
Learnware: on the future of machine learning
Zhi-Hua Zhou · 2016
Earlier work this paper cites.
Checkpoint ensembles: Ensemble methods from a single training process
Hugh Chen, Scott Lundberg, and Su-In Lee · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Snapshot ensembles: Train 1, get m for free
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E Hopcroft, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Distributed consistent data association via permutation synchronization
Spyridon Leonardos, Xiaowei Zhou, and Kostas Daniilidis · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
An investigation of how neural networks learn from the experiences of peers through periodic weight averaging
Joshua Smith and Michael Gashler · 2017
Earlier work this paper cites.
Federated multi-task learning
Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, and Ameet S Talwalkar · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola · 2017
Earlier work this paper cites.
Ensemble of deep neural networks with probability-based fusion for facial expression recognition
Guihua Wen, Zhi Hou, Huihui Li, Danyang Li, Lijun Jiang, and Eryang Xun · 2017
Earlier work this paper cites.
Modal consistency based pre-trained multi-model reuse
Yang Yang De-Chuan Zhan Xiang and Yu Guo Yuan Jiang · 2017
Earlier work this paper cites.
Deep learning for fixed model reuse
Yang Yang, De-Chuan Zhan, Ying Fan, Yuan Jiang, and Zhi-Hua Zhou · 2017
Earlier work this paper cites.
Learning from multiple teacher networks
Shan You, Chang Xu, Chao Xu, and Dacheng Tao · 2017
Earlier work this paper cites.
Knowledge fusion in feedforward artificial neural networks
Milad I Akhlaghi and Sergey V Sukhov · 2018
Earlier work this paper cites.
Large scale distributed neural network training through online distillation
Rohan Anil, Gabriel Pereyra, Alexandre Passos, Robert Ormandi, George E Dahl, and Geoffrey E Hinton · 2018
Earlier work this paper cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Digital retina: revolutionizing camera systems for the smart city
Wen Gao, Yonghong Tian, and Jian Wang · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Earlier work this paper cites.
Using mode connectivity for loss landscape analysis
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
Prateek Jain, Sham Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Earlier work this paper cites.
Eunjeong Jeong, Seungeun Oh, Hyesung Kim, Jihong Park, Mehdi Bennis, and Seong-Lyun Kim · 2018
Earlier work this paper cites.
Bag of experts architectures for model reuse in conversational language understanding
Rahul Jha, Alex Marin, Suvamsh Shivaprasad, and Imed Zitouni · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Earlier work this paper cites.
Paraphrasing complex network: Network compression via factor transfer
Jangho Kim, SeongUk Park, and Nojun Kwak · 2018
Earlier work this paper cites.
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Earlier work this paper cites.
A comparable study on model averaging, ensembling and reranking in nmt
Yuchen Liu, Long Zhou, Yining Wang, Yang Zhao, Jiajun Zhang, and Chengqing Zong · 2018
Earlier work this paper cites.
Variance networks: When expectation does not meet your expectations
Kirill Neklyudov, Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov · 2018
Earlier work this paper cites.
Iterate averaging as regularization for stochastic gradient descent
Gergely Neu and Lorenzo Rosasco · 2018
Earlier work this paper cites.
On the loss landscape of a class of deep neural networks with no bad local valleys
Quynh Nguyen, Mahesh Chandra Mukkamala, and Matthias Hein · 2018
Earlier work this paper cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman · 2018
Earlier work this paper cites.
Ensemble learning: A survey
Omer Sagi and Lior Rokach · 2018
Earlier work this paper cites.
Identifying generalization properties in neural networks
Huan Wang, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Earlier work this paper cites.
Graph hypernetworks for neural architecture search
Chris Zhang, Mengye Ren, and Raquel Urtasun · 2018
Cited alongside, same era.
An overview of multi-task learning
Yu Zhang and Qiang Yang · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
Sungsoo Ahn, Shell Xu Hu, Andreas Damianou, Neil D Lawrence, and Zhenwen Dai · 2019
Cited alongside, same era.
Ensemble approach for natural language question answering problem
Anna Aniol, Marcin Pietron, and Jerzy Duda · 2019
Cited alongside, same era.
Johanni Brea, Berfin Simsek, Bernd Illing, and Wulfram Gerstner · 2019
Cited alongside, same era.
Sparse training via boosting pruning plasticity with neuroregeneration
Shiwei Liu, Tianlong Chen, Xiaohan Chen, Zahra Atashgahi, Lu Yin, Huanyu Kou, Li Shen, Mykola Pechenizkiy, Zhangyang Wang, and Decebal Constantin Mocanu · 2021
Later among the works it cites.
Fedct: Federated collaborative transfer for recommendation
Shuchang Liu, Shuyuan Xu, Wenhui Yu, Zuohui Fu, Yongfeng Zhang, and Amelie Marian · 2021
Later among the works it cites.
Fedbabu: Towards enhanced representation for federated image classification
Jaehoon Oh, Sangmook Kim, and Se-Young Yun · 2021
Later among the works it cites.
Adaptive federated optimization, 2021
Sashank Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and H. Brendan McMahan · 2021
Later among the works it cites.
Fedaux: Leveraging unlabeled auxiliary data in federated learning
Felix Sattler, Tim Korjakow, Roman Rischke, and Wojciech Samek · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wojciech Marian Czarnecki, Simon Osindero, Razvan Pascanu, and Max Jaderberg · 2019
Cited alongside, same era.
Diversity with cooperation: Ensemble methods for few-shot classification
Nikita Dvornik, Cordelia Schmid, and Julien Mairal · 2019
Cited alongside, same era.
Deep ensembles: A loss landscape perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan · 2019
Cited alongside, same era.
Large scale structure of neural network loss landscapes
Stanislav Fort and Stanislaw Jastrzebski · 2019
Cited alongside, same era.
The goldilocks zone: Towards better understanding of neural network loss landscapes
Stanislav Fort and Adam Scherlis · 2019
Cited alongside, same era.
Neel Guha, Ameet Talwalkar, and Virginia Smith · 2019
Cited alongside, same era.
Collective model fusion for multiple black-box experts
Minh Hoang, Nghia Hoang, Bryan Kian Hsiang Low, and Carleton Kingsford · 2019
Cited alongside, same era.
Zoo-tuning: Adaptive transfer from a zoo of models
Yang Shu, Zhi Kou, Zhangjie Cao, Jianmin Wang, and Mingsheng Long · 2021
Later among the works it cites.
Boost neural networks by checkpoints
Feng Wang, Guoyizhe Wei, Qiao Liu, Jinxiang Ou, Hairong Lv, et al · 2021
Later among the works it cites.
Learning neural network subspaces
Mitchell Wortsman, Maxwell C Horton, Carlos Guestrin, Ali Farhadi, and Mohammad Rastegari · 2021
Later among the works it cites.
Peer collaborative learning for online knowledge distillation
Guile Wu and Shaogang Gong · 2021
Later among the works it cites.
Model reuse with reduced kernel mean embedding specification
Xi-Zhu Wu, Wenkai Xu, Song Liu, and Zhi-Hua Zhou · 2021
Later among the works it cites.
Deep neural network compression through interpretability-based filter pruning
Kaixuan Yao, Feilong Cao, Yee Leung, and Jiye Liang · 2021
Later among the works it cites.
Data-free knowledge distillation for heterogeneous federated learning
Zhuangdi Zhu, Junyuan Hong, and Jiayu Zhou · 2021
Later among the works it cites.
Multi-feature, multi-modal, and multi-source social event detection: A comprehensive survey
Imad Afyouni, Zaher Al Aghbari, and Reshma Abdul Razack · 2022
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Samuel K Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa · 2022
Later among the works it cites.
Wasserstein barycenter-based model fusion and linear mode connectivity of neural networks
Aditya Kumar Akash, Sixu Li, and Nicolás García Trillos · 2022
Later among the works it cites.
Ensemble of averages: Improving model selection and boosting performance in domain generalization
Devansh Arpit, Huan Wang, Yingbo Zhou, and Caiming Xiong · 2022
Later among the works it cites.
Random initialisations performing above chance and how to find them
Frederik Benzing, Simon Schug, Robert Meier, Johannes Von Oswald, Yassir Akram, Nicolas Zucchet, Laurence Aitchison, and Angelika Steger · 2022
Later among the works it cites.
Revisiting parameter-efficient tuning: Are we really there yet?
Guanzheng Chen, Fangyu Liu, Zaiqiao Meng, and Shangsong Liang · 2022
Later among the works it cites.
Where to start? analyzing the potential value of intermediate models
Leshem Choshen, Elad Venezian, Shachar Don-Yehia, Noam Slonim, and Yoav Katz · 2022
Later among the works it cites.
Fusing finetuned models for better pretraining
Leshem Choshen, Elad Venezian, Noam Slonim, and Yoav Katz · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Cold fusion: Collaborative descent for distributed multitask finetuning, 2022
Shachar Don-Yehiya, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen · 2022
Later among the works it cites.
Stylegan-nada: Clip-guided domain adaptation of image generators
Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or · 2022
Later among the works it cites.
Revisiting checkpoint averaging for neural machine translation
Yingbo Gao, Christian Herold, Zijian Yang, and Hermann Ney · 2022
Later among the works it cites.
Learning a subspace of policies for online adaptation in reinforcement learning, 2022
Jean-Baptiste Gaya, Laure Soulier, and Ludovic Denoyer · 2022
Later among the works it cites.
On the symmetries of deep learning models and their internal representations
Charles Godfrey, Davis Brown, Tegan Emerson, and Henry Kvinge · 2022
Later among the works it cites.
Data-free one-shot federated learning under very high statistical heterogeneity
Clare Elizabeth Heinbaugh, Emilio Luz-Ricca, and Huajie Shao · 2022
Later among the works it cites.
Achieving personalized federated learning with sparse local models
Tiansheng Huang, Shiwei Liu, Li Shen, Fengxiang He, Weiwei Lin, and Dacheng Tao · 2022
Later among the works it cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2022
Later among the works it cites.
Patching open-vocabulary models by interpolating weights
Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song, Hannaneh Hajishirzi, Simon Kornblith, Ali Farhadi, and Ludwig Schmidt · 2022
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng · 2022
Later among the works it cites.
Repair: Renormalizing permuted activations for interpolation repair
Keller Jordan, Hanie Sedghi, Olga Saukh, Rahim Entezari, and Behnam Neyshabur · 2022
Later among the works it cites.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra · 2022
Later among the works it cites.
Stop wasting my time! saving days of imagenet and bert training with latest weight averaging
Jean Kaddour · 2022
Later among the works it cites.
Branch-train-merge: Embarrassingly parallel training of expert language models
Margaret Li, Suchin Gururangan, Tim Dettmers, Mike Lewis, Tim Althoff, Noah A Smith, and Luke Zettlemoyer · 2022
Later among the works it cites.
Trainable weight averaging: Efficient training by optimizing historical solutions
Tao Li, Zhehao Huang, Qinghua Tao, Yingwen Wu, and Xiaolin Huang · 2022
Later among the works it cites.
Low dimensional trajectory hypothesis is true: Dnns can be trained in tiny subspaces
Tao Li, Lei Tan, Zhehao Huang, Qinghua Tao, Yipeng Liu, and Xiaolin Huang · 2022
Later among the works it cites.
Deep neural network fusion via graph matching with applications to model ensemble and federated learning
Chang Liu, Chenfei Lou, Runzhong Wang, Alan Yuhan Xi, Li Shen, and Junchi Yan · 2022
Later among the works it cites.
Nonlinear multi-model reuse
Yong Luo, Ling-Yu Duan, Yan Bai, Tongliang Liu, Yihang Lou, and Yonggang Wen · 2022
Later among the works it cites.
Merging models with fisher-weighted averaging
Michael S Matena and Colin A Raffel · 2022
Later among the works it cites.
Improving ensemble distillation with weight averaging and diversifying perturbation
Giung Nam, Hyungi Lee, Byeongho Heo, and Juho Lee · 2022
Later among the works it cites.
Re-basin via implicit sinkhorn differentiation
Fidel A Guerrero Peña, Heitor Rapela Medeiros, Thomas Dubail, Masih Aminbeidokhti, Eric Granger, and Marco Pedersoli · 2022
Later among the works it cites.
Deep networks on toroids: removing symmetries reveals the structure of flat regions in the landscape geometry
Fabrizio Pittorino, Antonio Ferraro, Gabriele Perugini, Christoph Feinauer, Carlo Baldassi, and Riccardo Zecchina · 2022
Later among the works it cites.
Exploring mode connectivity for pre-trained language models
Yujia Qin, Cheng Qian, Jing Yi, Weize Chen, Yankai Lin, Xu Han, Zhiyuan Liu, Maosong Sun, and Jie Zhou · 2022
Later among the works it cites.
Diverse weight averaging for out-of-distribution generalization
Alexandre Rame, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy, Patrick Gallinari, and Matthieu Cord · 2022
Later among the works it cites.
Towards summary candidates fusion
Mathieu Ravaut, Shafiq Joty, and Nancy F Chen · 2022
Later among the works it cites.
Talking about large language models
Murray Shanahan · 2022
Later among the works it cites.
Meta-learning without data via wasserstein distributionally-robust model fusion
Zhenyi Wang, Xiaoyang Wang, Li Shen, Qiuling Suo, Kaiqiang Song, Dong Yu, Yan Shen, and Mingchen Gao · 2022
Later among the works it cites.
lo-fi: distributed fine-tuning without communication
Mitchell Wortsman, Suchin Gururangan, Shen Li, Ali Farhadi, Ludwig Schmidt, Michael Rabbat, and Ari S Morcos · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al · 2022
Later among the works it cites.
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, et al · 2022
Later among the works it cites.
Ranking and tuning pre-trained models: a new paradigm for exploiting model hubs
Kaichao You, Yong Liu, Ziyang Zhang, Jianmin Wang, Michael I Jordan, and Mingsheng Long · 2022
Later among the works it cites.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Later among the works it cites.
Dense: Data-free one-shot federated learning
Jie Zhang, Chen Chen, Bo Li, Lingjuan Lyu, Shuang Wu, Shouhong Ding, Chunhua Shen, and Chao Wu · 2022
Later among the works it cites.
Fine-tuning global model via data-free knowledge distillation for non-iid federated learning
Lin Zhang, Li Shen, Liang Ding, Dacheng Tao, and Ling-Yu Duan · 2022
Later among the works it cites.
Fedsoup: Improving generalization and personalization in federated learning via selective model interpolation, 2023
Minghui Chen, Meirui Jiang, Qi Dou, Zehua Wang, and Xiaoxiao Li · 2023
Closest in time.
Adaptersoup: Weight averaging to improve generalization of pretrained language models
Alexandra Chronopoulou, Matthew E Peters, Alexander Fraser, and Jesse Dodge · 2023
Closest in time.
Seasoning model soups for robustness to adversarial and natural distribution shifts
Francesco Croce, Sylvestre-Alvise Rebuffi, Evan Shelhamer, and Sven Gowal · 2023
Closest in time.
Elastic weight removal for faithful and abstractive dialogue generation
Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, and Edoardo M Ponti · 2023
Closest in time.
Pareto manifold learning: Tackling multiple tasks via ensembles of single-task models
Nikolaos Dimitriadis, Pascal Frossard, and François Fleuret · 2023
Closest in time.
Hierarchical weight averaging for deep neural networks
Xiaozhe Gu, Zixun Zhang, Yuncheng Jiang, Tao Luo, Ruimao Zhang, Shuguang Cui, and Zhen Li · 2023
Closest in time.
Stochastic weight averaging revisited
Hao Guo, Jiyong Jin, and Bin Liu · 2023
Closest in time.
Lorahub: Efficient cross-task generalization via dynamic lora composition, 2023
Chengsong Huang, Qian Liu, Bill Yuchen Lin, Tianyu Pang, Chao Du, and Min Lin · 2023
Closest in time.
Experts weights averaging: A new general training scheme for vision transformers
Yongqi Huang, Peng Ye, Xiaoshui Huang, Sheng Li, Tao Chen, and Wanli Ouyang · 2023
Closest in time.
Exploring the benefits of training expert language models over instruction tuning, 2023
Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo · 2023
Closest in time.
A survey on multi-modal summarization
Anubhav Jangra, Sourajit Mukherjee, Adam Jatowt, Sriparna Saha, and Mohammad Hasanuzzaman · 2023
Closest in time.
Fedexp: Speeding up federated averaging via extrapolation
Divyansh Jhunjhunwala, Shiqiang Wang, and Gauri Joshi · 2023
Closest in time.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin · 2023
Closest in time.
Population parameter averaging (papa)
Alexia Jolicoeur-Martineau, Emy Gervais, Kilian Fatras, Yan Zhang, and Simon Lacoste-Julien · 2023
Closest in time.
Trainable weight averaging: A general approach for subspace training, 2023
Tao Li, Zhehao Huang, Qinghua Tao, Yingwen Wu, and Xiaolin Huang · 2023
Closest in time.
Hierarchical prompt learning for multi-task learning
Yajing Liu, Yuning Lu, Hao Liu, Yaozu An, Zhuoran Xu, Zhuokun Yao, Baofeng Zhang, Zhiwei Xiong, and Chenguang Gui · 2023
Closest in time.
Mechanistic mode connectivity
Ekdeep Singh Lubana, Eric J Bigelow, Robert P Dick, David Krueger, and Hidenori Tanaka · 2023
Closest in time.
Parameter-efficient weight ensembling facilitates task-level knowledge transfer
Xingtai Lv, Ning Ding, Yujia Qin, Zhiyuan Liu, and Maosong Sun · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Guillermo Ortiz-Jimenez, Alessandro Favero, and Pascal Frossard · 2023
Closest in time.
Model ratatouille: Recycling diverse models for out-of-distribution generalization
Alexandre Rame, Kartik Ahuja, Jianyu Zhang, Matthieu Cord, Leon Bottou, and David Lopez-Paz · 2023
Closest in time.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards, 2023
Alexandre Rame, Guillaume Couairon, Mustafa Shukor, Corentin Dancette, Jean-Baptiste Gaya, Laure Soulier, and Matthieu Cord · 2023
Closest in time.
Zipit! merging models from different tasks without training
George Stoica, Daniel Bolya, Jakob Bjorner, Taylor Hearn, and Judy Hoffman · 2023
Closest in time.
Multitask pre-training of modular prompt for chinese few-shot learning
Tianxiang Sun, Zhengfu He, Qin Zhu, Xipeng Qiu, and Xuan-Jing Huang · 2023
Closest in time.
An empirical study of multimodal model merging
Yi-Lin Sung, Linjie Li, Kevin Lin, Zhe Gan, Mohit Bansal, and Lijuan Wang · 2023
Closest in time.
Geodesic mode connectivity
Charlie Tan, Theodore Long, Sarah Zhao, and Rudolf Laine · 2023
Closest in time.
Improving heterogeneous model reuse by density estimation
Anke Tang, Yong Luo, Han Hu, Fengxiang He, Kehua Su, Bo Du, Yixin Chen, and Dacheng Tao · 2023
Closest in time.
Exploring diversified adversarial robustness in neural networks via robust mode connectivity
Ren Wang, Yuxuan Li, and Sijia Liu · 2023
Closest in time.
Ntk-approximating mlp fusion for efficient language model fine-tuning
Tianxin Wei, Zeming Guo, Yifan Chen, and Jingrui He · 2023
Closest in time.
Optimizing mode connectivity for class incremental learning
Haitao Wen, Haoyang Cheng, Heqian Qiu, Lanxiao Wang, Lili Pan, and Hongliang Li · 2023
Closest in time.
Traversing between modes in function space for fast ensembling
EungGu Yun, Hyungi Lee, Giung Nam, and Juho Lee · 2023
Closest in time.
Zhijian: A unifying and rapidly deployable toolbox for pre-trained model reuse
Yi-Kai Zhang, Lu Ren, Chao Yi, Qi-Wei Wang, De-Chuan Zhan, and Han-Jia Ye · 2023
Closest in time.
A survey of large language models, 2023
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, and Xiaolei Wang · 2023
Closest in time.
Sparse model soups: A recipe for improved pruning via model averaging
Max Zimmer, Christoph Spiegel, and Sebastian Pokutta · 2023
Closest in time.