Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are transforming the landscape of mobile intelligence.
Estimation of the mean of a multivariate normal distribution
Charles M Stein · 1981
Earlier work this paper cites.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1985
Earlier work this paper cites.
Gradient following without back-propagation in layered networks
Andrew G Barto and Michael I Jordan · 1987
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
Online convex optimization in the bandit setting: gradient descent without a gradient
Abraham D Flaxman, Adam Tauman Kalai, and H Brendan McMahan · 2004
Earlier work this paper cites.
Differential privacy: A survey of results
Cynthia Dwork · 2008
Earlier work this paper cites.
Observed universality of phase transitions in high-dimensional geometry, with implications for modern data analysis and signal processing
David Donoho and Jared Tanner · 2009
Earlier work this paper cites.
Improved analysis of the subsampled randomized hadamard transform
Joel A Tropp · 2011
Earlier work this paper cites.
Advanced data assimilation for geosciences: Lecture notes of the LES Houches School of Physics: Special issue, June 2012
Éric Blayo, Marc Bocquet, Emmanuel Cosme, and Leticia F Cugliandolo · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Federated learning of deep networks using model averaging
H Brendan McMahan, Eider Moore, Daniel Ramage, and Blaise Agüera y Arcas · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
https://ai.googleblog.com/2017/04/federated-learning-collaborative.html , 2017
Federated learning: Collaborative machine learning without centralized training data · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Optimizing low memory killers for mobile devices using reinforcement learning
Cong Li, Jia Bao, and Haitao Wang · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Earlier work this paper cites.
Error compensated quantized sgd and its applications to large-scale distributed optimization
Jiaxiang Wu, Weidong Huang, Junzhou Huang, and Tong Zhang · 2018
Earlier work this paper cites.
Deeptype: On-device deep learning for input personalization service with minimal privacy concern
Mengwei Xu, Feng Qian, Qiaozhu Mei, Kang Huang, and Xuanzhe Liu · 2018
Earlier work this paper cites.
Deeptype: On-device deep learning for input personalization service with minimal privacy concern
Mengwei Xu, Feng Qian, Qiaozhu Mei, Kang Huang, and Xuanzhe Liu · 2018
Earlier work this paper cites.
Towards federated learning at scale: System design
Keith Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloe Kiddon, Jakub Konečnỳ, Stefano Mazzocchi, Brendan McMahan, et al · 2019
Earlier work this paper cites.
Deqa: On-device question answering
Qingqing Cao, Noah Weber, Niranjan Balasubramanian, and Aruna Balasubramanian · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Client selection for federated learning with heterogeneous resources in mobile edge
Takayuki Nishio and Ryo Yonetani · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Earlier work this paper cites.
Low-memory neural network training: A technical report
Nimit S Sohoni, Christopher R Aberger, Megan Leszczynski, Jian Zhang, and Christopher Ré · 2019
Earlier work this paper cites.
Adaptive federated learning in resource constrained edge computing systems
Shiqiang Wang, Tiffany Tuor, Theodoros Salonidis, Kin K Leung, Christian Makaya, Ting He, and Kevin Chan · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Earlier work this paper cites.
Deep leakage from gradients
Ligeng Zhu, Zhijian Liu, and Song Han · 2019
Earlier work this paper cites.
Tinytl: Reduce activations, not trainable parameters for efficient on-device learning
Han Cai, Chuang Gan, Ligeng Zhu, and Song Han · 2020
Earlier work this paper cites.
Mathematics for machine learning
Marc Peter Deisenroth, A Aldo Faisal, and Cheng Soon Ong · 2020
Earlier work this paper cites.
Dynabert: Dynamic bert with adaptive width and depth
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu · 2020
Cited alongside, same era.
Oort: Informed participant selection for scalable federated learning
Fan Lai, Xiangfeng Zhu, Harsha V Madhyastha, and Mosharaf Chowdhury · 2020
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Cited alongside, same era.
End the senseless killing: Improving memory management for mobile operating systems
Niel Lebeck, Arvind Krishnamurthy, Henry M Levy, and Irene Zhang · 2020
Cited alongside, same era.
Fast and scalable in-memory deep multitask learning via neural weight virtualization
Seulki Lee and Shahriar Nirjon · 2020
Cited alongside, same era.
Fedadapter: Efficient federated learning for modern nlp
Dongqi Cai, Yaozong Wu, Shangguang Wang, Felix Xiaozhu Lin, and Mengwei Xu · 2022
Later among the works it cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2022
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig · 2022
Later among the works it cites.
The forward-forward algorithm: Some preliminary investigations
Geoffrey Hinton · 2022
Later among the works it cites.
Band: coordinated multi-dnn inference on heterogeneous mobile processors
Joo Seong Jeong, Jingyu Lee, Donghyun Kim, Changmin Jeon, Changjin Jeong, Youngki Lee, and Byung-Gon Chun · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Federated optimization in heterogeneous networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith · 2020
Cited alongside, same era.
A primer on zeroth-order optimization in signal processing and machine learning: Principals, recent advances, and applications
Sijia Liu, Pin-Yu Chen, Bhavya Kailkhura, Gaoyuan Zhang, Alfred O Hero III, and Pramod K Varshney · 2020
Cited alongside, same era.
Fastbert: a self-distilling bert with adaptive inference time
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Haotang Deng, and Qi Ju · 2020
Cited alongside, same era.
The hsic bottleneck: Deep learning without back-propagation
Wan-Duo Kurt Ma, JP Lewis, and W Bastiaan Kleijn · 2020
Cited alongside, same era.
Capuchin: Tensor-based gpu memory management for deep learning
Xuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin, Weiliang Ma, Qian Xiong, Fan Yang, and Xuehai Qian · 2020
Cited alongside, same era.
Adapterhub: A framework for adapting transformers
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych · 2020
Cited alongside, same era.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou · 2020
Cited alongside, same era.
Later among the works it cites.
Fedscale: Benchmarking model and system performance of federated learning at scale
Fan Lai, Yinwei Dai, Sanjay Singapuram, Jiachen Liu, Xiangfeng Zhu, Harsha Madhyastha, and Mosharaf Chowdhury · 2022
Later among the works it cites.
Pyramidfl: A fine-grained client selection framework for efficient federated learning
Chenning Li, Xiao Zeng, Mi Zhang, and Zhichao Cao · 2022
Later among the works it cites.
Cutting down on prompts and parameters: Simple few-shot learning with language models
Robert Logan IV, Ivana Balažević, Eric Wallace, Fabio Petroni, Sameer Singh, and Sebastian Riedel · 2022
Later among the works it cites.
Scaling forward gradient with local losses
Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton · 2022
Later among the works it cites.
Fedbalancer: data and pace control for efficient federated learning on heterogeneous clients
Jaemin Shin, Yuanchun Li, Yunxin Liu, and Sung-Ju Lee · 2022
Later among the works it cites.
Bbtv2: towards a gradient-free future with large language models
Tianxiang Sun, Zhengfu He, Hong Qian, Yunhua Zhou, Xuan-Jing Huang, and Xipeng Qiu · 2022
Later among the works it cites.
Lst: Ladder side-tuning for parameter and memory efficient transfer learning, 2022
Yi-Lin Sung, Jaemin Cho, and Mohit Bansal · 2022
Later among the works it cites.
Fedbert: When federated learning meets pre-training
Yuanyishu Tian, Yao Wan, Lingjuan Lyu, Dezhong Yao, Hai Jin, and Lichao Sun · 2022
Later among the works it cites.
Melon: Breaking the memory wall for resource-efficient on-device machine learning
Qipeng Wang, Mengwei Xu, Chao Jin, Xinran Dong, Jinliang Yuan, Xin Jin, Gang Huang, Yunxin Liu, and Xuanzhe Liu · 2022
Later among the works it cites.
Mandheling: mixed-precision on-device dnn training with dsp offloading
Daliang Xu, Mengwei Xu, Qipeng Wang, Shangguang Wang, Yun Ma, Kang Huang, Gang Huang, Xin Jin, and Xuanzhe Liu · 2022
Later among the works it cites.
Fednlp: Benchmarking federated learning methods for natural language processing tasks
Bill Yuchen Lin, Chaoyang He, Zihang Zeng, Hulin Wang, Yufen Huang, Christophe Dupuy, Rahul Gupta, Mahdi Soltanolkotabi, Xiang Ren, and Salman Avestimehr · 2022
Later among the works it cites.
Towards practical few-shot federated nlp
Dongqi Cai, Yaozong Wu, Haitao Yuan, Shangguang Wang, Felix Xiaozhu Lin, and Mengwei Xu · 2023
Closest in time.
Federated large language model: A position paper
Chaochao Chen, Xiaohua Feng, Jun Zhou, Jianwei Yin, and Xiaolin Zheng · 2023
Closest in time.
Does federated learning really need backpropagation?
Haozhe Feng, Tianyu Pang, Chao Du, Wei Chen, Shuicheng Yan, and Min Lin · 2023
Closest in time.
Emsassist: An end-to-end mobile voice assistant at the edge for emergency medical services
Liuyi Jin, Tian Liu, Amran Haroon, Radu Stoleru, Michael Middleton, Ziwei Zhu, and Theodora Chaspari · 2023
Closest in time.
Client-customized adaptation for parameter-efficient federated learning
Yeachan Kim, Junho Kim, Wing-Lam Mok, Jun-Hyung Park, and SangKeun Lee · 2023
Closest in time.
Dynamic and efficient inference for text generation via bert family
Xiaobo Liang, Juntao Li, Lijun Wu, Ziqiang Cao, and Min Zhang · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2023
Closest in time.
Efficientvit: Memory efficient vision transformer with cascaded group attention
Xinyu Liu, Houwen Peng, Ningxin Zheng, Yuqing Yang, Han Hu, and Yixuan Yuan · 2023
Closest in time.
Gradma: A gradient-memory-based accelerated federated learning with alleviated catastrophic forgetting
Kangyang Luo, Xiang Li, Yunshi Lan, and Ming Gao · 2023
Closest in time.
Fine-tuning language models with just forward passes
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora · 2023
Closest in time.
Fedzen: Towards superlinear zeroth-order federated learning via incremental hessian estimation, 2023
Alessio Maritan, Subhrakanti Dey, and Luca Schenato · 2023
Closest in time.
Fedfwd: Federated learning without backpropagation
Seonghwan Park, Dahun Shin, Jinseok Chung, and Namhoon Lee · 2023
Closest in time.
Zhen Qin, Daoyuan Chen, Bingchen Qian, Bolin Ding, Yaliang Li, and Shuiguang Deng · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
An optical xnor-bitcount based accelerator for efficient inference of binary neural networks
Sairam Sri Vatsavai, Venkata Sai Praneeth Karempudi, and Ishan Thakkar · 2023
Closest in time.
Can public large language models help private cross-device federated learning?
Boxin Wang, Yibo Jacky Zhang, Yuan Cao, Bo Li, H Brendan McMahan, Sewoong Oh, Zheng Xu, and Manzil Zaheer · 2023
Closest in time.
Towards building the federated gpt: Federated instruction tuning
Jianyi Zhang, Saeed Vahidian, Martin Kuo, Chunyuan Li, Ruiyi Zhang, Guoyin Wang, and Yiran Chen · 2023
Closest in time.
Fedprompt: Communication-efficient and privacy-preserving prompt tuning in federated learning
Haodong Zhao, Wei Du, Fangqi Li, Peixuan Li, and Gongshen Liu · 2023
Closest in time.
When foundation model meets federated learning: Motivations, challenges, and future directions
Weiming Zhuang, Chen Chen, and Lingjuan Lyu · 2023
Closest in time.
A survey of resource-efficient llm and multimodal foundation models
Mengwei Xu, Wangsong Yin, Dongqi Cai, Rongjie Yi, Daliang Xu, Qipeng Wang, Bingyang Wu, Yihao Zhao, Chen Yang, Shihe Wang, et al · 2024
Closest in time.