Fetching the paper…
Reading the bibliography…
Language model training in distributed settings is limited by the communication cost of gradient exchanges.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
James C Spall · 1992
Earlier work this paper cites.
Optimization over discrete sets via spsa
László Gerencsér, Stacy D Hill, and Zsuzsanna Vago · 1999
Earlier work this paper cites.
Efficient global optimization using spsa
John L Maryak and Daniel C Chin · 1999
Earlier work this paper cites.
Climateprediction. net: a global community for research in climate physics
David A Stainforth, Myles R Allen, David Frame, Jamie Kettleborough, Carl Christensen, Tolu Aina, and Matthew Collins · 2004
Earlier work this paper cites.
Parallelized-over-parts computation of absolute binding free energy with docking and molecular dynamics
Guha Jayachandran, Michael R Shirts, Sanghyun Park, and Vijay S Pande · 2006
Earlier work this paper cites.
A model free automatic tuning method for a restricted structured controller by using simultaneous perturbation stochastic approximation (spsa)
QingHui Yuan · 2008
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Communication with imperfectly shared randomness
Clément Louis Canonne, Venkatesan Guruswami, Raghu Meka, and Madhu Sudan · 2015
Earlier work this paper cites.
Psr j1906+ 0722: an elusive gamma-ray pulsar
CJ Clark, HJ Pletsch, J Wu, Lucas Guillemot, Markus Ackermann, B Allen, Annalisa de Angelis, C Aulbert, Luca Baldini, Jean Ballet, et al · 2015
Earlier work this paper cites.
Amyloid β \beta protein and alzheimer’s disease: When computer simulations complement experimental studies
Jessica Nasica-Labouze, Phuong H Nguyen, Fabio Sterpone, Olivia Berthoumieu, Nicolae-Viorel Buchete, Sebastien Cote, Alfonso De Simone, Andrew J Doig, Peter Faller, Angel Garcia, et al · 2015
Earlier work this paper cites.
c-spsa: Cluster-wise simultaneous perturbation stochastic approximation algorithm and its application to dynamic origin–destination matrix estimation
Athina Tympakianaki, Haris N Koutsopoulos, and Erik Jenelius · 2015
Earlier work this paper cites.
Analysis of gradient descent methods with nondiminishing bounded errors
Arunselvan Ramaswamy and Shalabh Bhatnagar · 2017
Earlier work this paper cites.
Robust spsa algorithms for dynamic od matrix estimation
Athina Tympakianaki, Haris N Koutsopoulos, and Erik Jenelius · 2018
Earlier work this paper cites.
Spsa for layer-wise training of deep networks
Benjamin Wulff, Jannis Schuecker, and Christian Bauckhage · 2018
Earlier work this paper cites.
Communication-constrained inference and the role of shared randomness
Jayadev Acharya, Clément Canonne, and Himanshu Tyagi · 2019
Earlier work this paper cites.
Ieee standard for floating-point arithmetic
IEEE 754 Working Group et al · 2019
Earlier work this paper cites.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Anastasia Koloskova, Sebastian Stich, and Martin Jaggi · 2019
Cited alongside, same era.
Hop: Heterogeneity-aware decentralized training
Qinyi Luo, Jinkun Lin, Youwei Zhuo, and Xuehai Qian · 2019
Cited alongside, same era.
Efficient sparse secure aggregation for federated learning
Constance Beguier, Mathieu Andreux, and Eric W Tramel · 2020
Cited alongside, same era.
Privacy-preserving learning via deep net pruning
Yangsibo Huang, Yushan Su, Sachin Ravi, Zhao Song, Sanjeev Arora, and Kai Li · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Basil: A fast and byzantine-resilient approach for decentralized training
Ahmed Roushdy Elkordy, Saurav Prakash, and Salman Avestimehr · 2022
Later among the works it cites.
Sparse random networks for communication-efficient federated learning
Berivan Isik, Francesco Pase, Deniz Gunduz, Tsachy Weissman, and Michele Zorzi · 2022
Later among the works it cites.
Branch-train-merge: Embarrassingly parallel training of expert language models
Margaret Li, Suchin Gururangan, Tim Dettmers, Mike Lewis, Tim Althoff, Noah A Smith, and Luke Zettlemoyer · 2022
Later among the works it cites.
Intrinsic gradient compression for scalable and efficient federated learning
Luke Melas-Kyriazi and Franklyn Wang · 2022
Later among the works it cites.
Lossy compression with gaussian diffusion
Lucas Theis, Tim Salimans, Matthew D Hoffman, and Fabian Mentzer · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dspg: Decentralized simultaneous perturbations gradient descent scheme
Arunselvan Ramaswamy · 2020
Cited alongside, same era.
Simultaneous perturbation stochastic approximation of the quantum fisher information
Julien Gacon, Christa Zoufal, Giuseppe Carleo, and Stefan Woerner · 2021
Cited alongside, same era.
Covariant quantum kernels for data with group structure
Jennifer R Glick, Tanvi P Gujarati, Antonio D Corcoles, Youngseok Kim, Abhinav Kandala, Jay M Gambetta, and Kristan Temme · 2021
Cited alongside, same era.
Model averaging in distributed machine learning: a case study with apache spark
Yunyan Guo, Zhipeng Zhang, Jiawei Jiang, Wentao Wu, Ce Zhang, Bin Cui, and Jianzhong Li · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Coordination through shared randomness
Gowtham R Kurri, Vinod M Prabhakaran, and Anand D Sarwate · 2021
Cited alongside, same era.
Convergence analysis of weighted spsa-based consensus algorithm in distributed parameter estimation problem
Anna Sergeenko, Victoria Erofeeva, Oleg Granichin, Olga Granichina, and Anton Proskurnikov · 2021
Cited alongside, same era.
Later among the works it cites.
Decentralized training of foundation models in heterogeneous environments
Binhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang, Tri Dao, Beidi Chen, Percy S Liang, Christopher Re, and Ce Zhang · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Training differentially private graph neural networks with random walk sampling
Morgane Ayle, Jan Schuchardt, Lukas Gosch, Daniel Zügner, and Stephan Günnemann · 2023
Closest in time.
Randomness should be consistent across devices with use_deterministic_algorithms
Antoine Bouthors · 2023
Closest in time.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2023
Closest in time.
Google fi mobile broadband consumer disclosure, 2023
Google Fi · 2023
Closest in time.
On the geometric convergence of byzantine-resilient distributed optimization algorithms
Kananart Kuwaranancharoen and Shreyas Sundaram · 2023
Closest in time.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma · 2023
Closest in time.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora · 2023
Closest in time.
Releasing gpt-jt powered by open-source ai
Together · 2023
Closest in time.
An empirical comparison of optimizers for quantum machine learning with spsa-based gradients
Marco Wiedmann, Marc Hölle, Maniraman Periyasamy, Nico Meyer, Christian Ufrecht, Daniel D Scherer, Axel Plinge, and Christopher Mutschler · 2023
Closest in time.
Anthropic’s $5b, 4-year plan to take on openai, April 2023
Kyle Wiggers, Devin Coldewey, and Manish Singh · 2023
Closest in time.
A modified second-order spsa optimization algorithm for finite samples
Xun Zhu and James C Spall · 2023
Closest in time.