On the mathematical foundations of theoretical statistics
Ronald A Fisher. 1922 · 1922
Earlier work this paper cites.
Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts. In SIGKDD . ACM, 1930–1939
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. 2018 · 1939
Earlier work this paper cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel. 2022b · 1965
Earlier work this paper cites.
Animating rotation with quaternion curves. In Proceedings of the 12th annual conference on Computer graphics and interactive techniques . 245–254
Ken Shoemake. 1985 · 1985
Earlier work this paper cites.
Catastrophic forgetting, rehearsal and pseudorehearsal
Anthony Robins. 1995 · 1995
Earlier work this paper cites.
Weight averaging for neural networks and local resampling schemes. In AAAI Workshop . Citeseer, 133–138
Joachim Utans. 1996 · 1996
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Ensemble learning
Thomas G Dietterich et al · 2002
Earlier work this paper cites.
Torchvision the machine-vision package of torch. In ACM MM . 1485–1488
Sébastien Marcel and Yann Rodriguez. 2010 · 2010
Earlier work this paper cites.
ESC: Dataset for environmental sound classification. In ACM MM . 1015–1018
Karol J Piczak. 2015 · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent. In ICML . PMLR, 1225–1234
Moritz Hardt, Ben Recht, and Yoram Singer. 2016 · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data. In AISTATS . PMLR, 1273–1282
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017 · 2017
Earlier work this paper cites.
An investigation of how neural networks learn from the experiences of peers through periodic weight averaging. In 2017 16th IEEE ICML and Applications (ICMLA) . IEEE, 731–736
Joshua Smith and Michael Gashler. 2017 · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks. In ICML . PMLR, 794–803
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. 2018 · 2018
Earlier work this paper cites.
AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models. In EACL . 2009–2018
Alexandra Chronopoulou, Matthew E Peters, Alexander Fraser, and Jesse Dodge. 2023a · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape. In ICML . PMLR, 1309–1318
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht. 2018 · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson. 2018 · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization, In UAI
Original
P Izmailov, AG Wilson, D Podoprikhin, D Vetrov, and T Garipov. 2018 · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler. 2018 · 2018
Earlier work this paper cites.
Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
Prateek Jain, Sham M Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford. 2018 · 2018
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2018 · 2018
Earlier work this paper cites.
The california consumer privacy act: Towards a european-style privacy regime in the united states
Stuart L Pardau. 2018 · 2018
Earlier work this paper cites.
Ensemble learning: A survey
Omer Sagi and Lior Rokach. 2018 · 2018
Earlier work this paper cites.
Multi-Task Learning as Multi-Objective Optimization. In NeurIPS . 525–536
Ozan Sener and Vladlen Koltun. 2018 · 2018
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification. In WWW . 491–500
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019 · 2019
Earlier work this paper cites.
Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
Dan Hendrycks and Thomas Dietterich. 2019 · 2019
Earlier work this paper cites.
The European Union general data protection regulation: what it is and what it means
Chris Jay Hoofnagle, Bart Van Der Sloot, and Frederik Zuiderveen Borgesius. 2019 · 2019
Earlier work this paper cites.
Learning private neural language modeling with attentive aggregation. In IJCNN . IEEE, 1–8
Shaoxiong Ji, Shirui Pan, Guodong Long, Xue Li, Jing Jiang, and Zi Huang. 2019 · 2019
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Original
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
Explaining landscape connectivity of low-cost solutions for multilayer nets
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Rong Ge, and Sanjeev Arora. 2019 · 2019
Earlier work this paper cites.
Uniform convergence may be unable to explain generalization in deep learning
Vaishnavh Nagarajan and J Zico Kolter. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
PyTorch Image Models
Ross Wightman. 2019 · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Original
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Earlier work this paper cites.
SWALP: Stochastic weight averaging in low precision training. In ICML . PMLR, 7015–7024
Guandao Yang, Tianyi Zhang, Polina Kirichenko, Junwen Bai, Andrew Gordon Wilson, and Chris De Sa. 2019 · 2019
Earlier work this paper cites.
Bayesian nonparametric federated learning of neural networks. In ICML . PMLR, 7252–7261
Mikhail Yurochkin, Mayank Agarwal, Soumya Ghosh, Kristjan Greenewald, Nghia Hoang, and Yasaman Khazaeni. 2019 · 2019
Earlier work this paper cites.
Fine-Tuning Language Models from Human Preferences
Original
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 2019
Earlier work this paper cites.
Linear mode connectivity and the lottery ticket hypothesis. In ICML . PMLR, 3259–3269
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. 2020 · 2020
Earlier work this paper cites.
Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes Well. In ICLR . OpenReview.net
Vipul Gupta, Santiago Akle Serrano, and Dennis DeCoste. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Federated learning: Challenges, methods, and future directions
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith. 2020 · 2020
Earlier work this paper cites.
Model fusion via optimal transport
Sidak Pal Singh and Martin Jaggi. 2020 · 2020
Earlier work this paper cites.
Adashare: Learning what to share for efficient deep multi-task learning
Ximeng Sun, Rameswar Panda, Rogerio Feris, and Kate Saenko. 2020 · 2020
Earlier work this paper cites.
Optimizing mode connectivity via neuron alignment
Norman Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk, Prasanna Sattigeri, and Rongjie Lai. 2020 · 2020
Earlier work this paper cites.
Tackling the objective inconsistency problem in heterogeneous federated optimization
Jianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi, and H Vincent Poor. 2020a · 2020
Earlier work this paper cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. 2020 · 2020
Earlier work this paper cites.
Swa object detection
Original
Haoyang Zhang, Ying Wang, Feras Dayoub, and Niko Sünderhauf. 2020 · 2020
Earlier work this paper cites.
Loss surface simplexes for mode connecting volumes and fast ensembling. In ICML . PMLR, 769–779
Gregory Benton, Wesley Maddox, Sanae Lotfi, and Andrew Gordon Gordon Wilson. 2021 · 2021
Earlier work this paper cites.
Swad: Domain generalization by seeking flat minima
Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park. 2021 · 2021
Earlier work this paper cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. 2021 · 2021
Earlier work this paper cites.
GeDi: Generative Discriminator Guided Sequence Generation
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2021 · 2021
Earlier work this paper cites.
Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances. In ICML . PMLR, 9722–9732
Berfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro, Clément Hongler, Wulfram Gerstner, and Johanni Brea. 2021 · 2021
Earlier work this paper cites.
Multi-task learning for dense prediction tasks: A survey
Simon Vandenhende, Stamatios Georgoulis, Wouter Van Gansbeke, Marc Proesmans, Dengxin Dai, and Luc Van Gool. 2021 · 2021
Earlier work this paper cites.
Neural networks with late-phase weights. In ICLR
Johannes von Oswald, Seijin Kobayashi, Alexander Meulemans, Christian Henning, Benjamin F. Grewe, and João Sacramento. 2021 · 2021
Earlier work this paper cites.
CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLP. In EMNLP . 7163–7189
Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren. 2021 · 2021
Earlier work this paper cites.
Ensemble of averages: Improving model selection and boosting performance in domain generalization
Devansh Arpit, Huan Wang, Yingbo Zhou, and Caiming Xiong. 2022 · 2022
Earlier work this paper cites.
GAN Cocktail: mixing GANs without dataset access. In ECCV . Springer, 205–221
Omri Avrahami, Dani Lischinski, and Ohad Fried. 2022 · 2022
Earlier work this paper cites.
The role of permutation invariance in linear mode connectivity of neural networks
Rahim Entezari, Hanie Sedghi, Olga Saukh, and Behnam Neyshabur. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models. In ICLR
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Patching open-vocabulary models by interpolating weights
Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song, Hannaneh Hajishirzi, Simon Kornblith, Ali Farhadi, and Ludwig Schmidt. 2022 · 2022
Earlier work this paper cites.
Stop wasting my time! saving days of imagenet and bert training with latest weight averaging
Jean Kaddour. 2022 · 2022
Earlier work this paper cites.
When do flat minima optimizers work?
Jean Kaddour, Linqing Liu, Ricardo Silva, and Matt J Kusner. 2022 · 2022
Earlier work this paper cites.
Branch-train-merge: Embarrassingly parallel training of expert language models
Original
Margaret Li, Suchin Gururangan, Tim Dettmers, Mike Lewis, Tim Althoff, Noah A Smith, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
Merging models with fisher-weighted averaging
Michael S Matena and Colin A Raffel. 2022 · 2022
Earlier work this paper cites.
Cross-Task Generalization via Natural Language Crowdsourcing Instructions. In ACL . Association for Computational Linguistics (ACL), 3470–3487
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022 · 2022
Earlier work this paper cites.
Diverse weight averaging for out-of-distribution generalization
Alexandre Rame, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy, Patrick Gallinari, and Matthieu Cord. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models. In CVPR . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Earlier work this paper cites.
Unrolling sgd: Understanding factors influencing machine unlearning. In EuroS&P . IEEE, 303–319
Anvith Thudi, Gabriel Deza, Varun Chandrasekaran, and Nicolas Papernot. 2022 · 2022
Earlier work this paper cites.
Plateau in Monotonic Linear Interpolation—A" Biased" View of Loss Landscape for Deep Networks. In ICLR
Xiang Wang, Annie N Wang, Mo Zhou, and Rong Ge. 2022 · 2022
Earlier work this paper cites.
BackdoorBench: A Comprehensive Benchmark of Backdoor Learning. In NeurIPS Datasets and Benchmarks Track
Baoyuan Wu, Hongrui Chen, Mingda Zhang, Zihao Zhu, Shaokui Wei, Danni Yuan, and Chao Shen. 2022 · 2022
Earlier work this paper cites.
Gpt-4 technical report
Original
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Git Re-Basin: Merging Models modulo Permutation Symmetries. In ICLR
Samuel Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa. 2023 · 2023
Earlier work this paper cites.
Robust weight signatures: gaining robustness as easy as patching weights?. In ICML . PMLR, 3495–3506
Ruisi Cai, Zhenyu Zhang, and Zhangyang Wang. 2023 · 2023
Earlier work this paper cites.
Task Arithmetic with LoRA for Continual Learning
Original
Rajas Chitale, Ankit Vaidya, Aditya Kane, and Archana Ghotkar. 2023 · 2023
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Earlier work this paper cites.
Language and Task Arithmetic with Parameter-Efficient Layers for Zero-Shot Summarization
Original
Alexandra Chronopoulou, Jonas Pfeiffer, Joshua Maynez, Xinyi Wang, Sebastian Ruder, and Priyanka Agrawal. 2023b · 2023
Earlier work this paper cites.
Seasoning model soups for robustness to adversarial and natural distribution shifts. In CVPR . 12313–12323
Francesco Croce, Sylvestre-Alvise Rebuffi, Evan Shelhamer, and Sven Gowal. 2023 · 2023
Earlier work this paper cites.
Model breadcrumbs: Scaling multi-task model merging with sparse masks
Original
MohammadReza Davari and Eugene Belilovsky. 2023 · 2023
Earlier work this paper cites.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ning Ding, Yujia Qin, Guang Yang, Fuchao Wei, Zonghan Yang, Yusheng Su, Shengding Hu, Yulin Chen, Chi-Min Chan, Weize Chen, et al · 2023
Earlier work this paper cites.
MerA: Merging pretrained adapters for few-shot learning
Original
Shwai He, Run-Ze Fan, Liang Ding, Li Shen, Tianyi Zhou, and Dacheng Tao. 2023a · 2023
Earlier work this paper cites.
Editing models with task arithmetic. In ICLR
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023 · 2023
Earlier work this paper cites.
Dart: Diversify-aggregate-repeat training improves generalization of neural networks. In CVPR . 16048–16059
Samyak Jain, Sravanti Addepalli, Pawan Kumar Sahu, Priyam Dey, and R Venkatesh Babu. 2023 · 2023
Earlier work this paper cites.
Personalized soups: Personalized large language model alignment via post-hoc parameter merging
Original
Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu. 2023a · 2023
Earlier work this paper cites.
ForkMerge: Mitigating Negative Transfer in Auxiliary-Task Learning
Junguang Jiang, Baixu Chen, Junwei Pan, Ximei Wang, Dapeng Liu, Jie Jiang, and Mingsheng Long. 2023 · 2023
Earlier work this paper cites.
Dataless Knowledge Fusion by Merging Weights of Language Models. In ICLR
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng. 2023 · 2023
Earlier work this paper cites.
Repair: Renormalizing permuted activations for interpolation repair
Keller Jordan, Hanie Sedghi, Olga Saukh, Rahim Entezari, and Behnam Neyshabur. 2023 · 2023
Earlier work this paper cites.