Fetching the paper…
Reading the bibliography…
Generative pre-trained large language models (LLMs) have demonstrated impressive performance over a wide range of tasks, thanks to the unprecedented amount of data they have been trained on.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 1909
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Christopher Hesse, Rafal Józefowicz, Scott Gray, Catherine Olsson, Jakub Pachocki, Michael Petrov, Henrique Pondé de Oliveira Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang · 1912
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Faster on-device training using new federated momentum algorithm
Zhouyuan Huo, Qian Yang, Bin Gu, Lawrence Carin, and Heng Huang · 2002
Earlier work this paper cites.
Low-resource languages: A review of past work and future challenges
Alexandre Magueresse, Vincent Carles, and Evan Heetderks · 2006
Earlier work this paper cites.
Flower: A friendly federated learning research framework
Daniel J. Beutel, Taner Topal, Akhil Mathur, Xinchi Qiu, Javier Fernandez-Marques, Yan Gao, Lorenzo Sani, Kwing Hei Li, Titouan Parcollet, Pedro Porto Buarque de Gusmão, and Nicholas D. Lane · 2007
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon · 2011
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
Monthly energy consumption forecast: A deep learning approach
R. F. Berriel, A. T. Lopes, A. Rodrigues, F. M. Varejão, and T. Oliveira-Santos · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
A systematic literature review on microservices
Hulya Vural, Murat Koyuncu, and Sinem Guney · 2017
Earlier work this paper cites.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, and Blake A. Hechtman · 2018
Earlier work this paper cites.
Learning differentially private recurrent language models
H. Brendan McMahan, Daniel Ramage, Kunal Talwar, and Li Zhang · 2018
Earlier work this paper cites.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the AI2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? A new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2019
Earlier work this paper cites.
Towards federated learning at scale: System design
Kallista A. Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloé Kiddon, Jakub Konečný, Stefano Mazzocchi, Brendan McMahan, Timon Van Overveldt, David Petrou, Daniel Ramage, and Jason Roselander · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
Hydra - a framework for elegantly configuring complex applications
Omry Yadan · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
Don’t use large mini-batches, use local SGD
Tao Lin, Sebastian U. Stich, Kumar Kshitij Patel, and Martin Jaggi · 2020
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Zero: memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Earlier work this paper cites.
Federated optimization in heterogeneous networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith · 2020
Earlier work this paper cites.
Model compression and hardware acceleration for neural networks: A comprehensive survey
Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Fair resource allocation in federated learning
Tian Li, Maziar Sanjabi, Ahmad Beirami, and Virginia Smith · 2020
Earlier work this paper cites.
Personalized federated learning with theoretical guarantees: A model-agnostic meta-learning approach
Alireza Fallah, Aryan Mokhtari, and Asuman E. Ozdaglar · 2020
Earlier work this paper cites.
PIQA: reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2020
Cited alongside, same era.
Zero-offload: Democratizing billion-scale model training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He · 2021
Cited alongside, same era.
Trade-offs of local SGD at scale: An empirical study
Jose Javier Gonzalez Ortiz, Jonathan Frankle, Mike Rabbat, Ari S. Morcos, and Nicolas Ballas · 2021
Cited alongside, same era.
Scaling federated learning for fine-tuning of large language models
Agrin Hilmkil, Sebastian Callh, Matteo Barbieri, Leon René Sütfeld, Edvin Listo Zec, and Olof Mogren · 2021
Cited alongside, same era.
https://lajavaness.medium.com/llm-large-language-model-cost-analysis-d5022bb43e9e , 2023
La Javaness · 2023
Later among the works it cites.
Distributed inference and fine-tuning of large language models over the internet, 2023
Alexander Borzunov, Max Ryabinin, Artem Chumachenko, Dmitry Baranchuk, Tim Dettmers, Younes Belkada, Pavel Samygin, and Colin Raffel · 2023
Later among the works it cites.
Performance analysis of federated learning algorithms for multilingual protest news detection using pre-trained distilbert and BERT
Pascal Riedel, Manfred Reichert, Reinhold von Schwerin, Alexander Hafner, Daniel Schaudt, and Gaurav Singh · 2023
Later among the works it cites.
Can public large language models help private cross-device federated learning?
Boxin Wang, Jacky Yibo Zhang, Yuan Cao, Bo Li, H. Brendan McMahan, Sewoong Oh, Zheng Xu, and Manzil Zaheer · 2023
Later among the works it cites.
Towards building the federated GPT: federated instruction tuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Cited alongside, same era.
Adaptive federated optimization
Sashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and Hugh Brendan McMahan · 2021
Cited alongside, same era.
Differentially private learning with adaptive clipping
Galen Andrew, Om Thakkar, Brendan McMahan, and Swaroop Ramaswamy · 2021
Cited alongside, same era.
Tilted empirical risk minimization
Tian Li, Ahmad Beirami, Maziar Sanjabi, and Virginia Smith · 2021
Cited alongside, same era.
On large-cohort training for federated learning
Zachary Charles, Zachary Garrett, Zhouyuan Huo, Sergei Shmulyian, and Virginia Smith · 2021
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2021
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2021
Cited alongside, same era.
Jianyi Zhang, Saeed Vahidian, Martin Kuo, Chunyuan Li, Ruiyi Zhang, Guoyin Wang, and Yiran Chen · 2023
Later among the works it cites.
FATE-LLM: A industrial grade federated learning framework for large language models
Tao Fan, Yan Kang, Guoqiang Ma, Weijing Chen, Wenbin Wei, Lixin Fan, and Qiang Yang · 2023
Later among the works it cites.
Weirui Kuang, Bingchen Qian, Zitao Li, Daoyuan Chen, Dawei Gao, Xuchen Pan, Yuexiang Xie, Yaliang Li, Bolin Ding, and Jingren Zhou · 2023
Later among the works it cites.
Low-parameter federated learning with large language models
Jingang Jiang, Xiangyang Liu, and Chenyou Fan · 2023
Later among the works it cites.
Reducing communication overhead in federated learning for pre-trained language models using parameter-efficient finetuning
Shubham Malaviya, Manish Shukla, and Sachin Lodha · 2023
Later among the works it cites.
Training large-vocabulary neural language models by private federated learning for resource-constrained devices
Mingbin Xu, Congzheng Song, Ye Tian, Neha Agrawal, Filip Granqvist, Rogier C. van Dalen, Xiao Zhang, Arturo Argueta, Shiyi Han, Yaqiao Deng, Leo Liu, Anmol Walia, and Alex Jin · 2023
Later among the works it cites.
Slora: Federated parameter efficient fine-tuning of language models
Sara Babakniya, Ahmed Roushdy Elkordy, Yahya H. Ezzeldin, Qingfeng Liu, Kee-Bong Song, Mostafa El-Khamy, and Salman Avestimehr · 2023
Later among the works it cites.
Client-customized adaptation for parameter-efficient federated learning
Yeachan Kim, Junho Kim, Wing-Lam Mok, Jun-Hyung Park, and SangKeun Lee · 2023
Later among the works it cites.
Fedprompt: Communication-efficient and privacy-preserving prompt tuning in federated learning
Haodong Zhao, Wei Du, Fangqi Li, Peixuan Li, and Gongshen Liu · 2023
Later among the works it cites.
Federated learning of large language models with parameter-efficient prompt tuning and adaptive optimization
Tianshi Che, Ji Liu, Yang Zhou, Jiaxiang Ren, Jiwen Zhou, Victor S. Sheng, Huaiyu Dai, and Dejing Dou · 2023
Later among the works it cites.
Neural machine translation for low-resource languages: A survey
Surangika Ranathunga, En-Shiun Annie Lee, Marjana Prifti Skenduli, Ravi Shekhar, Mehreen Alam, and Rishemjit Kaur · 2023
Later among the works it cites.
Federated learning for computationally constrained heterogeneous devices: A survey
Kilian Pfeiffer, Martin Rapp, Ramin Khalili, and Jörg Henkel · 2023
Later among the works it cites.
Towards federated foundation models: Scalable dataset pipelines for group-structured learning
Zachary Charles, Nicole Mitchell, Krishna Pillutla, Michael Reneer, and Zachary Garrett · 2023
Later among the works it cites.
Introducing MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs, 2023
MosaicML NLP Team · 2023
Later among the works it cites.
Fedl2p: Federated learning to personalize
Royson Lee, Minyoung Kim, Da Li, Xinchi Qiu, Timothy M. Hospedales, Ferenc Huszar, and Nicholas D. Lane · 2023
Later among the works it cites.
Generalised winograd schema and its contextuality
Kin Ian Lo, Mehrnoosh Sadrzadeh, and Shane Mansfield · 2023
Later among the works it cites.
Introducing FlowerLLM, 2024
Nicholas D. Lane, Lorenzo Sani, and Alex Iacob · 2024
Closest in time.
An archival perspective on pretraining data
Meera A Desai, Irene V Pasquetto, Abigail Z Jacobs, and Dallas Card · 2024
Closest in time.
Securing large language models: Threats, vulnerabilities and responsible practices
Sara Abdali, Richard Anarfi, C. J. Barberan, and Jia He · 2024
Closest in time.
Asynchronous local-sgd training for language modeling
Bo Liu, Rachita Chhaparia, Arthur Douillard, Satyen Kale, Andrei A. Rusu, Jiajun Shen, Arthur Szlam, and Marc’Aurelio Ranzato · 2024
Closest in time.
Dipaco: Distributed path composition
Arthur Douillard, Qixuan Feng, Andrei A. Rusu, Adhiguna Kuncoro, Yani Donchev, Rachita Chhaparia, Ionel Gog, Marc’Aurelio Ranzato, Jiajun Shen, and Arthur Szlam · 2024
Closest in time.
The era of 1-bit llms: All large language models are in 1.58 bits
Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei · 2024
Closest in time.
Fwdllm: Efficient fedllm using forward gradient, 2024
Mengwei Xu, Dongqi Cai, Yaozong Wu, Xiang Li, and Shangguang Wang · 2024
Closest in time.
Breaking physical and linguistic borders: Multilingual federated prompt tuning for low-resource languages
Wanru Zhao, Royson Lee, Yihong Chen, Xinchi Qiu, Yan Gao, Hongxiang Fan, and Nicholas Donald Lane · 2024
Closest in time.
OpenAI offers publishers as little as $1 million a year — the information, Jan 2024
Sahil Patel and Stephanie Palazzolo · 2024
Closest in time.
Understanding and improving model averaging in federated learning on heterogeneous data
Tailin Zhou, Zehong Lin, Jun Zhang, and Danny H.K. Tsang · 2024
Closest in time.
The Object Store for AI Data Infrastructure, 2024
Inc. MinIO · 2024
Closest in time.
Amazon S3, 2024
Amazon · 2024
Closest in time.
Boto3 - The AWS SDK for Python, 2024
the boto project · 2024
Closest in time.
mosaic research, 2024
Databricks · 2024
Closest in time.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws
Nikhil Sardana, Jacob Portes, Sasha Doubov, and Jonathan Frankle · 2024
Closest in time.
Exploring scaling laws for local sgd in large language model training, 2024
Qiaozhi He, Xiaomin Zhuang, and Zhihua Wu · 2024
Closest in time.