Fetching the paper…
Reading the bibliography…
Large language models (LLM) have become a critical component in many applications of machine learning.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2019
Earlier work this paper cites.
Lookahead optimizer: k steps forward, 1 step back
Michael R. Zhang, James Lucas, Geoffrey Hinton, and Jimmy Ba · 2019
Earlier work this paper cites.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin · 2020
Earlier work this paper cites.
Faster on-device training using new federated momentum algorithm
Zhouyuan Huo, Qian Yang, Bin Gu, and Lawrence Carin. Heng Huang · 2020
Earlier work this paper cites.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U. Stich, Kumar Kshitij Patel, and Martin Jaggi · 2020
Earlier work this paper cites.
Swarm training, 2020
Shawn Presser · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Slowmo: Improving communication-efficient distributed sgd with slow momentum
Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Michael Rabbat · 2020
Cited alongside, same era.
Distributed deep learning in open collaborations
Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin, Lucile Saulnier, Quentin Lhoest, Anton Sinitsin, Dmitry Popov, Dmitry Pyrkin, Maxim Kashirin, Alexander Borzunov, Albert Villanova del Moral, Denis Mazur, Ilia Kobelev, Yacine Jernite, Thomas Wolf, and Gennady Pekhimenko · 2021
Cited alongside, same era.
Trade-offs of local sgd at scale: An empirical study
Jose Javier Gonzalez Ortiz, Jonathan Frankle, Mike Rabbat, Ari Morcos, and Nicolas Ballas · 2021
Cited alongside, same era.
Adaptive federated optimization
Sashank Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and H. Brendan McMahan · 2021
Cited alongside, same era.
Branch-train-merge: Embarrassingly parallel training of expert language models
Margaret Li, Suchin Gururangan, Tim Dettmers, Mike Lewis, Tim Althoff, Noah A. Smith, and Luke Zettlemoyer · 2022
Later among the works it cites.
Revisiting adapters with adversarial training
Sylvestre-Alvise Rebuffi, Francesco Croce, and Sven Gowal · 2022
Later among the works it cites.
Why (and when) does local sgd generalize better than sgd?
Xinran Gu, Kaifeng Lyu, Longbo Huang, and Sanjeev Arora · 2023
Closest in time.
Continual pre-training of large language models: How to (re)warm your model?
Kshitij Gupta, Benjamin Thérien, Adam Ibrahim, Mats L. Richter, Quentin Anthony, Eugene Belilovsky, Irina Rish, and Timothée Lesort · 2023
Closest in time.
Scaling expert language models with unsupervised domain discovery
Suchin Gururangan, Margaret Li, Mike Lewis, Weijia Shi, Tim Althoff, Noah A. Smith, and Luke Zettlemoyer · 2023
Closest in time.
Dataless knowledge fusion by merging weights of language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Moshpit sgd: Communication-efficient decentralized training on heterogeneous unreliable devices
Max Ryabinin, Eduard Gorbunov, Vsevolod Plokhotnyuk, and Gennady Pekhimenko · 2021
Cited alongside, same era.
Learning neural network subspaces
Mitchell Wortsman, Maxwell Horton, Carlos Guestrin, Ali Farhadi, and Mohammad Rastegari · 2021
Cited alongside, same era.
Petals: Collaborative inference and fine-tuning of large models
Alexander Borzunov, Dmitry Baranchuk, Tim Dettmers, Max Ryabinin, Younes Belkada, Artem Chumachenko, Pavel Samygin, and Colin Raffel · 2022
Cited alongside, same era.
A survey on heterogeneous federated learning
Dashan Gao, Xin Yao, and Qiang Yang · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Cited alongside, same era.
Patching open-vocabulary models by interpolating weights
Gabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song, Hannaneh Hajishirzi, Simon Kornblith, Ali Farhadi, and Ludwig Schmidt · 2022
Cited alongside, same era.
Stop wasting my time! saving days of imagenet and bert training with latest weight averaging
Jean Kaddour · 2022
Cited alongside, same era.
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng · 2023
Closest in time.
Population parameter averaging (papa)
Alexia Jolicoeur-Martineau, Emy Gervais, Kilian Fatras, Yan Zhang, and Simon Lacoste-Julien · 2023
Closest in time.
Repair: Renormalizing permuted activations for interpolation repair
Keller Jordan, Hanie Sedghi, Olga Saukh, Rahim Entezari, and Behnam Neyshabur · 2023
Closest in time.
Git-theta: A git extension for collaborative development of machine learning models
Nikhil Kandpal, Brian Lester, Mohammed Muqeeth, Anisha Mascarenhas, Monty Evans, Vishal Baskaran, Tenghao Huang, Haokun Liu, and Colin Raffel · 2023
Closest in time.
Zipit! merging models from different tasks without training
George Stoica, Daniel Bolya, Jakob Bjorner, Taylor Hearn, and Judy Hoffman · 2023
Closest in time.
Communication-efficient distributed deep learning: A comprehensive survey
Zhenheng Tang, Shaohuai Shi, Wei Wang, Bo Li, and Xiaowen Chu · 2023
Closest in time.
Resolving interference when merging models
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal · 2023
Closest in time.