Fetching the paper…
Reading the bibliography…
The confluence of Federated Learning (FL) and Large Language Models (LLMs) is ushering in a new era in privacy-preserving natural language processing.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of convex optimization
Alekh Agarwal, Martin J Wainwright, Peter Bartlett, and Pradeep Ravikumar. 2009 · 2009
Earlier work this paper cites.
Query complexity of derivative-free optimization
Kevin G Jamieson, Robert Nowak, and Ben Recht. 2012 · 2012
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining . 661–670
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola. 2014 · 2014
Earlier work this paper cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
John C Duchi, Michael I Jordan, Martin J Wainwright, and Andre Wibisono. 2015 · 2015
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data. In AISTATS . PMLR, 1273–1282
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2017 · 2017
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny. 2017 · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou. 2017 · 2017
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal. 2018 · 2018
Earlier work this paper cites.
On the information-adaptive variants of the ADMM: an iteration complexity perspective
Xiang Gao, Bo Jiang, and Shuzhong Zhang. 2018 · 2018
Earlier work this paper cites.
Measuring the Intrinsic Dimension of Objective Landscapes. In International Conference on Learning Representations
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski. 2018 · 2018
Earlier work this paper cites.
An investigation into neural net optimization via hessian eigenvalue density. In International Conference on Machine Learning . PMLR, 2232–2241
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao. 2019 · 2019
Earlier work this paper cites.
ZONE: Zeroth-order nonconvex multiagent optimization over networks
Davood Hajinezhad, Mingyi Hong, and Alfredo Garcia. 2019 · 2019
Earlier work this paper cites.
On the Convergence of FedAvg on Non-IID Data. In ICLR
Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. 2019 · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Personalized federated learning: A meta-learning approach. In NeurIPS 2020
Alireza Fallah, Aryan Mokhtari, and Asuman Ozdaglar. 2020 · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2020
Earlier work this paper cites.
Federated optimization in heterogeneous networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. 2020 · 2020
Earlier work this paper cites.
A primer on zeroth-order optimization in signal processing and machine learning: Principals, recent advances, and applications
Sijia Liu, Pin-Yu Chen, Bhavya Kailkhura, Gaoyuan Zhang, Alfred O Hero III, and Pramod K Varshney. 2020 · 2020
Earlier work this paper cites.
Throughput-Optimal Topology Design for Cross-Silo Federated Learning. In NeurIPS , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 19478–19487
Othmane Marfoq, Chuan Xu, Giovanni Neglia, and Richard Vidal. 2020 · 2020
Earlier work this paper cites.
Fedfast: Going beyond average for faster training of federated recommender systems. In KDD . 1234–1242
Khalil Muhammad, Qinqin Wang, Diarmuid O’Reilly-Morgan, Elias Tragos, Barry Smyth, Neil Hurley, James Geraci, and Aonghus Lawlor. 2020 · 2020
Earlier work this paper cites.
Clustered Federated Learning: Model-Agnostic Distributed Multitask Optimization Under Privacy Constraints
Felix Sattler, Klaus-Robert Müller, and Wojciech Samek. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations . 38–45
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2020
Cited alongside, same era.
Pyhessian: Neural networks through the lens of the hessian. In 2020 IEEE international conference on big data (Big data) . IEEE, 581–590
Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael W Mahoney. 2020 · 2020
Cited alongside, same era.
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 7319–7328
Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer. 2021 · 2021
Zeroth-order algorithms for stochastic distributed nonconvex optimization
Xinlei Yi, Shengjun Zhang, Tao Yang, and Karl H Johansson. 2022 · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Code alpaca: An instruction-following llama model for code generation
Sahil Chaudhary. 2023 · 2023
Later among the works it cites.
Federated large language model: A position paper
Chaochao Chen, Xiaohua Feng, Jun Zhou, Jianwei Yin, and Xiaolin Zheng. 2023a · 2023
Later among the works it cites.
FATE-LLM: A Industrial Grade Federated Learning Framework for Large Language Models
Tao Fan, Yan Kang, Guoqiang Ma, Weijing Chen, Wenbin Wei, Lixin Fan, and Qiang Yang. 2023a · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2021 · 2021
Cited alongside, same era.
Efficient on-chip learning for optical neural networks through power-aware sparse zeroth-order optimization. In Proceedings of the AAAI conference on artificial intelligence , Vol. 35. 7583–7591
Jiaqi Gu, Chenghao Feng, Zheng Zhao, Zhoufeng Ying, Ray T Chen, and David Z Pan. 2021 · 2021
Cited alongside, same era.
Federated Adversarial Debiasing for Fair and Transferable Representations. In KDD
Junyuan Hong, Zhuangdi Zhu, Shuyang Yu, Zhangyang Wang, Hiroko Dodge, and Jiayu Zhou. 2021 · 2021
Cited alongside, same era.
Advances and Open Problems in Federated Learning
Peter Kairouz, H. Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, Rafael G. L. D’Oliveira, Hubert Eichner, Salim El Rouayheb, David Evans, Josh Gardner, Zachary Garrett, Adrià Gascón, Badih Ghazi, Phillip B. Gibbons, Marco Gruteser, Zaid Harchaoui, Chaoyang He, Lie He, Zhouyuan Huo, Ben Hutchinson, Justin Hsu, Martin Jaggi, Tara Javidi, Gauri Joshi, Mikhail Khodak, Jakub Konecný, Aleksandra Korolova, Farinaz Koushanfar, Sanmi Koyejo, Tancrède Lepoint, Yang Liu, Prateek Mittal, Mehryar Mohri, Richard Nock, Ayfer Özgür, Rasmus Pagh, Hang Qi, Daniel Ramage, Ramesh Raskar, Mariana Raykova, Dawn Song, Weikang Song, Sebastian U. Stich, Ziteng Sun, Ananda Theertha Suresh, Florian Tramèr, Praneeth Vepakomma, Jianyu Wang, Li Xiong, Zheng Xu, Qiang Yang, Felix X. Yu, Han Yu, and Sen Zhao. 2021 · 2021
Cited alongside, same era.
Communication-memory-efficient decentralized learning for audio representation. In 2021 International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–8
LeiLai Li, Jianzong Wang, Xiaoyang Qu, and Jing Xiao. 2021 · 2021
Cited alongside, same era.
Cross-node federated graph neural network for spatio-temporal data modeling. In KDD . 1202–1211
Chuizheng Meng, Sirisha Rambhatla, and Yan Liu. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Fate-llm: A industrial grade federated learning framework for large language models
Tao Fan, Yan Kang, Guoqiang Ma, Weijing Chen, Wenbin Wei, Lixin Fan, and Qiang Yang. 2023b · 2023
Later among the works it cites.
Does Federated Learning Really Need Backpropagation?
Haozhe Feng, Tianyu Pang, Chao Du, Wei Chen, Shuicheng Yan, and Min Lin. 2023 · 2023
Later among the works it cites.
Weirui Kuang, Bingchen Qian, Zitao Li, Daoyuan Chen, Dawei Gao, Xuchen Pan, Yuexiang Xie, Yaliang Li, Bolin Ding, and Jingren Zhou. 2023 · 2023
Later among the works it cites.
Fine-Tuning Language Models with Just Forward Passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora. 2023 · 2023
Later among the works it cites.
Federated Zeroth-Order Optimization using Trajectory-Informed Surrogate Gradients
Yao Shu, Xiaoqiang Lin, Zhongxiang Dai, and Bryan Kian Hsiang Low. 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
FederatedScope: A Flexible Federated Learning Platform for Heterogeneity
Yuexiang Xie, Zhen Wang, Dawei Gao, Daoyuan Chen, Liuyi Yao, Weirui Kuang, Yaliang Li, Bolin Ding, and Jingren Zhou. 2023 · 2023
Later among the works it cites.
Towards Building the Federated GPT: Federated Instruction Tuning
Jianyi Zhang, Saeed Vahidian, Martin Kuo, Chunyuan Li, Ruiyi Zhang, Guoyin Wang, and Yiran Chen. 2023a · 2023
Later among the works it cites.
Towards Building the FederatedGPT: Federated Instruction Tuning. In International Workshop on Federated Learning in the Age of Foundation Models in Conjunction with NeurIPS 2023
Jianyi Zhang, Saeed Vahidian, Martin Kuo, Chunyuan Li, Ruiyi Zhang, Tong Yu, Guoyin Wang, and Yiran Chen. 2023b · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Jiamu Bai, Daoyuan Chen, Bingchen Qian, Liuyi Yao, and Yaliang Li. 2024 · 2024
Closest in time.
Data-Juicer: A One-Stop Data Processing System for Large Language Models. In International Conference on Management of Data
Daoyuan Chen, Yilun Huang, Zhijian Ma, Hesen Chen, Xuchen Pan, Ce Ge, Dawei Gao, Yuexiang Xie, Zhaoyang Liu, Jinyang Gao, Yaliang Li, Bolin Ding, and Jingren Zhou. 2024 · 2024
Closest in time.
Zhen Qin, Daoyuan Chen, Bingchen Qian, Bolin Ding, Yaliang Li, and Shuiguang Deng. 2024 · 2024
Closest in time.