Fetching the paper…
Reading the bibliography…
The conventional cloud-based large model learning framework is increasingly constrained by latency, cost, personalization, and privacy concerns.
Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models. In Proceedings of Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) . ACL, Mexico City, Mexico, 1964–1974
Keming Lu, Hongyi Yuan, Runji Lin, Junyang Lin, Zheng Yuan, Chang Zhou, and Jingren Zhou. 2024 · 1974
Earlier work this paper cites.
MAUI: making smartphones last longer with code offload. In Proceedings of International Conference on Mobile Systems, Applications, and Services (MobiSys) . ACM, San Francisco, CA, USA, 49–62
Eduardo Cuervo, Aruna Balasubramanian, Dae-ki Cho, Alec Wolman, Stefan Saroiu, Ranveer Chandra, and Paramvir Bahl. 2010 · 2010
Earlier work this paper cites.
CloneCloud: elastic execution between mobile device and cloud. In Proceedings of European conference on Computer systems (EuroSys) . ACM, Salzburg, Austria, 301–314
Byung-Gon Chun, Sunghwan Ihm, Petros Maniatis, Mayur Naik, and Ashwin Patti. 2011 · 2011
Earlier work this paper cites.
COMET: Code Offload by Migrating Execution Transparently. In Proceedings of USENIX Symposium on Operating Systems Design and Implementation (OSDI) . USENIX Association, Hollywood, CA, USA, 93–106
Mark S. Gordon, Davoud Anoushe Jamshidi, Scott A. Mahlke, Zhuoqing Morley Mao, and Xu Chen. 2012 · 2012
Earlier work this paper cites.
Accelerating Mobile Applications through Flip-Flop Replication. In Proceedings of Annual International Conference on Mobile Systems, Applications, and Services (MobiSys) . ACM, Florence, Italy, 137–150
Mark S. Gordon, David Ke Hong, Peter M. Chen, Jason Flinn, Scott A. Mahlke, and Zhuoqing Morley Mao. 2015 · 2015
Earlier work this paper cites.
Deep Learning with Limited Numerical Precision. In Proceedings of the 32nd International Conference on Machine Learning (ICML) . JMLR.org, Lille, France, 1737–1746
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015 · 2015
Earlier work this paper cites.
Learning both Weights and Connections for Efficient Neural Networks. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Montreal, Quebec, Canada, 1135–1143
Song Han, Jeff Pool, John Tran, and William J. Dally. 2015 · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
FitNets: Hints for Thin Deep Nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
TensorFlow: A System for Large-Scale Machine Learning. In Proceedings of USENIX Symposium on Operating Systems Design and Implementation (OSDI) . USENIX, Savannah, GA, USA, 265–283
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016 · 2016
Earlier work this paper cites.
Sequence-Level Knowledge Distillation. In Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP) . The Association for Computational Linguistics, Austin, Texas, USA, 1317–1327
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks. In Proceedings of European Conference on Computer Vision (ECCV) . Springer, Amsterdam, The Netherlands, 525–542
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
BranchyNet: Fast inference via early exiting from deep neural networks. In Proceedings of International Conference on Pattern Recognition (ICPR) . IEEE, Cancún, Mexico, 2464–2469
Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. 2016 · 2016
Earlier work this paper cites.
DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou. 2016 · 2016
Earlier work this paper cites.
TensorFlow Lite or LiteRT
Google. 2017 · 2017
Earlier work this paper cites.
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017 · 2017
Earlier work this paper cites.
Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge. In Proceedings of International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) . ACM, Xi’an, China, 615–629
Yiping Kang, Johann Hauswald, Cao Gao, Austin Rovinski, Trevor N. Mudge, Jason Mars, and Lingjia Tang. 2017 · 2017
Earlier work this paper cites.
DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset. In Proceedings of International Joint Conference on Natural Language Processing (IJCNLP) . Asian Federation of Natural Language Processing, Taipei, Taiwan, 986–995
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Earlier work this paper cites.
Learning Efficient Convolutional Networks through Network Slimming. In Proceedings of IEEE International Conference on Computer Vision (ICCV) . IEEE Computer Society, Venice, Italy, 2755–2763
Zhuang Liu, Jianguo Li, Zhiqiang Shen, Gao Huang, Shoumeng Yan, and Changshui Zhang. 2017 · 2017
Earlier work this paper cites.
Communication-Efficient Learning of Deep Networks from Decentralized Data. In Proceedings of International Conference on Artificial Intelligence and Statistics (AISTATS) . PMLR, Fort Lauderdale, FL, USA, 1273–1282
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas. 2017 · 2017
Earlier work this paper cites.
Federated Multi-Task Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Long Beach, CA, USA, 4424–4434
Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, and Ameet Talwalkar. 2017 · 2017
Earlier work this paper cites.
Distributed Deep Neural Networks Over the Cloud, the Edge and End Devices. In Proceedings of IEEE International Conference on Distributed Computing Systems (ICDCS) . IEEE, Atlanta, GA, USA, 328–339
Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. 2017 · 2017
Earlier work this paper cites.
A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE Computer Society, Honolulu, HI, USA, 7130–7138
Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim. 2017 · 2017
Earlier work this paper cites.
Optimized Cost per Click in Taobao Display Advertising. In Proceedings of ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD) . ACM, Halifax, NS, Canada, 2191–2200
Han Zhu, Junqi Jin, Chang Tan, Fei Pan, Yifan Zeng, Han Li, and Kun Gai. 2017 · 2017
Earlier work this paper cites.
Neural Architecture Search with Reinforcement Learning. In Proceedings of International Conference on Learning Representations, (ICLR) . OpenReview.net, Toulon, France, 16 pages
Barret Zoph and Quoc V. Le. 2017 · 2017
Earlier work this paper cites.
Ali_Display_Ad_Click
Alimama. 2018 · 2018
Earlier work this paper cites.
LEAF: A Benchmark for Federated Settings
Sebastian Caldas, Peter Wu, Tian Li, Jakub Konečný, H. Brendan McMahan, Virginia Smith, and Ameet Talwalkar. 2018 · 2018
Earlier work this paper cites.
TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of USENIX Symposium on Operating Systems Design and Implementation (OSDI) . USENIX, Carlsbad, CA, USA, 578–594
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018 · 2018
Earlier work this paper cites.
Distributed learning of deep neural network over multiple agents
Otkrist Gupta and Ramesh Raskar. 2018 · 2018
Earlier work this paper cites.
Transformers
Hugging Face. 2018 · 2018
Earlier work this paper cites.
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE Computer Society, Salt Lake City, UT, 2704–2713
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018 · 2018
Earlier work this paper cites.
MobileNetV2: Inverted Residuals and Linear Bottlenecks. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE Computer Society, Salt Lake City, UT, USA, 4510–4520
Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018 · 2018
Earlier work this paper cites.
Towards Federated Learning at Scale: System Design. In Proceedings of Conference on Machine Learning and Systems (SysML) . mlsys.org, Stanford, California, 15 pages
Kallista Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloé Kiddon, Jakub Konečný, Stefano Mazzocchi, Brendan McMahan, Timon Van Overveldt, David Petrou, Daniel Ramage, and Jason Roselander. 2019 · 2019
Earlier work this paper cites.
Scaling Video Analytics on Constrained Edge Nodes. In Proceedings of Machine Learning and Systems (MLSys) . mlsys.org, Stanford, California, 12 pages
Christopher Canel, Thomas Kim, Giulio Zhou, Conglong Li, Hyeontaek Lim, David G. Andersen, Michael Kaminsky, and Subramanya Dulloor. 2019 · 2019
Earlier work this paper cites.
Federated Meta-Learning with Fast Convergence and Efficient Communication
Fei Chen, Mi Luo, Zhenhua Dong, Zhenguo Li, and Xiuqiang He. 2019 · 2019
Earlier work this paper cites.
HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision. In Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Seoul, Korea (South), 293–302
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2019 · 2019
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, New Orleans, LA, USA, 42 pages
Jonathan Frankle and Michael Carbin. 2019 · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Long Beach, California, USA, 2790–2799
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Searching for MobileNetV3. In Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Seoul, South Korea, 1314–1324
Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le, Mark Sandler, Bo Chen, Weijun Wang, Liang-Chieh Chen, Mingxing Tan, Grace Chu, Vijay Vasudevan, and Yukun Zhu. 2019 · 2019
Earlier work this paper cites.
LinkShare: device-centric control for concurrent and continuous mobile-cloud interactions. In Proceedings of ACM/IEEE Symposium on Edge Computing (SEC) . ACM, Arlington, Virginia, USA, 15–29
Bo Hu and Wenjun Hu. 2019 · 2019
Earlier work this paper cites.
Collaborative learning between cloud and end devices: an empirical study on location prediction. In Proceedings of the ACM/IEEE Symposium on Edge Computing (SEC) . ACM/IEEE, Washington DC, USA, 139–151
Yan Lu, Yuanchao Shu, Xu Tan, Yunxin Liu, Mengyu Zhou, Qi Chen, and Dan Pei. 2019 · 2019
Earlier work this paper cites.
PyTorch Mobile
Meta. 2019 · 2019
Earlier work this paper cites.
Megatron-LM
NVIDIA. 2019 · 2019
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Vancouver, BC, Canada, 8024–8035
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Earlier work this paper cites.
Low-Memory Neural Network Training: A Technical Report
Nimit Sharad Sohoni, Christopher Richard Aberger, Megan Leszczynski, Jian Zhang, and Christopher Ré. 2019 · 2019
Earlier work this paper cites.
MnasNet: Platform-Aware Neural Architecture Search for Mobile. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE, Long Beach, CA, USA, 2820–2828
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le. 2019 · 2019
Earlier work this paper cites.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Long Beach, CA, USA, 6105–6114
Mingxing Tan and Quoc V. Le. 2019 · 2019
Earlier work this paper cites.
HAQ: Hardware-Aware Automated Quantization With Mixed Precision. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE, Long Beach, CA, USA, 8612–8620
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han. 2019 · 2019
Earlier work this paper cites.
Personalized Dialogue Generation with Diversified Traits
Yinhe Zheng, Guanyi Chen, Minlie Huang, Song Liu, and Xuan Zhu. 2019 · 2019
Earlier work this paper cites.
TinyTL: Reduce Memory, Not Parameters for Efficient On-Device Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Virtual, 13 pages
Han Cai, Chuang Gan, Ligeng Zhu, and Song Han. 2020 · 2020
Earlier work this paper cites.
Server-driven video streaming for deep learning inference. In Proceedings of Annual Conference of ACM’s Special Interest Group on Data Communication (SIGCOMM) . ACM, Virtual, 557–570
Kuntai Du, Ahsan Pervaiz, Xin Yuan, Aakanksha Chowdhery, Qizheng Zhang, Henry Hoffmann, and Junchen Jiang. 2020 · 2020
Earlier work this paper cites.
Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning Approach. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Virtual, 12 pages
Alireza Fallah, Aryan Mokhtari, and Asuman E. Ozdaglar. 2020 · 2020
Cited alongside, same era.
EdgeRec: Recommender System on Edge in Mobile Taobao. In Proceedings of ACM International Conference on Information and Knowledge Management (CIKM) . ACM, Virtual, 2477–2484
Yu Gong, Ziwen Jiang, Yufei Feng, Binbin Hu, Kaiqi Zhao, Qingwen Liu, and Wenwu Ou. 2020 · 2020
Cited alongside, same era.
Retrieval Augmented Language Model Pre-Training. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Virtual, 3929–3938
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Cited alongside, same era.
Federated Visual Classification with Real-World Data Distribution. In Proceedings of European Conference on Computer Vision (ECCV) . Springer, Glasgow, UK, 76–92
Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. 2020 · 2020
Cited alongside, same era.
llama.cpp
Georgi Gerganov. 2023 · 2023
Later among the works it cites.
Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., New Orleans, LA, USA, 13 pages
Yuchao Gu, Xintao Wang, Jay Zhangjie Wu, Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao, Shuning Chang, Weijia Wu, Yixiao Ge, Ying Shan, and Mike Zheng Shou. 2023 · 2023
Later among the works it cites.
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. In Findings of the Association for Computational Linguistics (ACL) . Association for Computational Linguistics, Toronto, Canada, 8003–8017
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023 · 2023
Later among the works it cites.
Rethinking Federated Learning with Domain Shift: A Prototype View. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Vancouver, BC, Canada, 16312–16322
Wenke Huang, Mang Ye, Zekun Shi, He Li, and Bo Du. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SCAFFOLD: Stochastic Controlled Averaging for Federated Learning. In Proceedings of International Conference on Machine Learning (ICML) , Vol. 119. PMLR, Virtual, 5132–5143
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh. 2020 · 2020
Cited alongside, same era.
Generalization through Memorization: Nearest Neighbor Language Models. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Addis Ababa, Ethiopia, 13 pages
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Cited alongside, same era.
A Unified Theory of Decentralized SGD with Changing Topology and Local Updates. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Virtual, 5381–5393
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U. Stich. 2020 · 2020
Cited alongside, same era.
SPINN: synergistic progressive inference of neural networks over device and cloud. In Proceedings of Annual International Conference on Mobile Computing and Networking (MobiCom) . ACM, London, United Kingdom, 37:1–37:15
Stefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis, and Nicholas D. Lane. 2020 · 2020
Cited alongside, same era.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Virtual, 9459–9474
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Cited alongside, same era.
Ensemble Distillation for Robust Model Fusion in Federated Learning. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., virtual, 13 pages
Tao Lin, Lingjing Kong, Sebastian U. Stich, and Martin Jaggi. 2020 · 2020
Cited alongside, same era.
Three Approaches for Personalization with Applications to Federated Learning
Yishay Mansour, Mehryar Mohri, Jae Ro, and Ananda Theertha Suresh. 2020 · 2020
Cited alongside, same era.
DeepSpeed
Microsoft. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Efficient Edge Inference by Selective Query. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Kigali, Rwanda, 25 pages
Anil Kag, Igor Fedorov, Aditya Gangrade, Paul N. Whatmough, and Venkatesh Saligrama. 2023 · 2023
Later among the works it cites.
Fast Inference from Transformers via Speculative Decoding. In International Conference on Machine Learning (ICML) (Proceedings of Machine Learning Research, Vol. 202) . Proceedings of Machine Learning Research (PMLR), Honolulu, Hawaii, USA, 19274–19286
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Later among the works it cites.
DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model Generalization. In Proceedings of ACM Web Conference (WWW) . ACM, Austin, TX, USA, 3077–3085
Zheqi Lv, Wenqiao Zhang, Shengyu Zhang, Kun Kuang, Feng Wang, Yongwei Wang, Zhengyu Chen, Tao Shen, Hongxia Yang, Beng Chin Ooi, and Fei Wu. 2023 · 2023
Later among the works it cites.
LLM-Pruner: On the Structural Pruning of Large Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., New Orleans, LA, USA, 19 pages
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023 · 2023
Later among the works it cites.
ExecuTorch
Meta. 2023 · 2023
Later among the works it cites.
Stanford Alpaca: An Instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
Tabi: An Efficient Multi-Level Inference System for Large Language Models. In Proceedings of European Conference on Computer Systems (EuroSys) . ACM, Rome, Italy, 233–248
Yiding Wang, Kai Chen, Haisheng Tan, and Kun Guo. 2023 · 2023
Later among the works it cites.
Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation. In Findings of the Association for Computational Linguistics (EMNLP) . Association for Computational Linguistics, Singapore, 3909–3925
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui. 2023 · 2023
Later among the works it cites.
Offsite-Tuning: Transfer Learning without Full Model
Guangxuan Xiao, Ji Lin, and Song Han. 2023a · 2023
Later among the works it cites.
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–18
Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024 · 2024
Later among the works it cites.
AutoMix: Automatically Mixing Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Vancouver, BC, Canada, 35 pages
Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Pei Zhou, Aditya Gupta, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang, Shyam Upadhyay, Manaal Faruqui, and Mausam. 2024 · 2024
Later among the works it cites.
Apple Intelligence Foundation Language Models
Apple. 2024b · 2024
Later among the works it cites.
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. In Proceedings of International Conference on Machine Learning (ICML) . OpenReview.net, Vienna, Austria, 42 pages
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu. 2024 · 2024
Later among the works it cites.
Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation Models. In Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational Linguistics, Miami, FL, USA, 12903–12913
Yae Jee Cho, Luyang Liu, Zheng Xu, Aldi Fahrezi, and Gauri Joshi. 2024 · 2024
Later among the works it cites.
AutoXPCR: Automated Multi-Objective Model Selection for Time Series Forecasting. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . ACM, Barcelona, Spain, 806–815
Raphael Fischer and Amal Saadallah. 2024 · 2024
Later among the works it cites.
Delta: A Cloud-assisted Data Enrichment Framework for On-Device Continual Learning. In Proceedings of Annual International Conference on Mobile Computing and Networking (MobiCom) . ACM, Washington D.C., DC, USA, 1408–1423
Chen Gong, Zhenzhe Zheng, Fan Wu, Xiaofeng Jia, and Guihai Chen. 2024 · 2024
Later among the works it cites.
MediaPipe LLM Inference API
Google. 2024 · 2024
Later among the works it cites.
MiniLLM: Knowledge Distillation of Large Language Models. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–24
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024 · 2024
Later among the works it cites.
The Surge of On-Device Large Models: Trends, Impacts, and Recommendations
Tencent Research Institute. 2024 · 2024
Later among the works it cites.
Adversarial Moment-Matching Distillation of Large Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Vancouver, BC, Canada, 1–33
Chen Jia. 2024 · 2024
Later among the works it cites.
DistiLLM: Towards Streamlined Distillation for Large Language Models. In Proceedings of International Conference on Machine Learning (ICML) . OpenReview.net, Vienna, Austria, 1–24
Jongwoo Ko, Sungnyun Kim, Tianyi Chen, and Se-Young Yun. 2024 · 2024
Later among the works it cites.
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration. In Proceedings of the Seventh Annual Conference on Machine Learning and Systems (MLSys) . mlsys.org, Santa Clara, CA, USA, 15 pages
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024 · 2024
Later among the works it cites.
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
Yuejiang Liu and Alexandre Alahi. 2024 · 2024
Later among the works it cites.
Llama Team, AI @ Meta. 2024 · 2024
Later among the works it cites.
torchchat
Meta. 2024 · 2024
Later among the works it cites.
An Emulator for Fine-tuning Large Language Models using Small Language Models. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 16 pages
Eric Mitchell, Rafael Rafailov, Archit Sharma, Chelsea Finn, and Christopher D. Manning. 2024 · 2024
Later among the works it cites.
Orca-Math: Unlocking the potential of SLMs in Grade School Math
Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah. 2024 · 2024
Later among the works it cites.
Orthogonal Adaptation for Modular Customization of Diffusion Models. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Seattle, WA, USA, 7964–7973
Ryan Po, Guandao Yang, Kfir Aberman, and Gordon Wetzstein. 2024 · 2024
Later among the works it cites.
ChatDev: Communicative Agents for Software Development. In Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL) . Association for Computational Linguistics, Bangkok, Thailand, 15174–15186
Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2024 · 2024
Later among the works it cites.
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
Ankit Singh Rawat, Veeranjaneyulu Sadhanala, Afshin Rostamizadeh, Ayan Chakrabarti, Wittawat Jitkrittum, Vladimir Feinberg, Seungyeon Kim, Hrayr Harutyunyan, Nikunj Saunshi, Zachary Nado, Rakesh Shivanna, Sashank J. Reddi, Aditya Krishna Menon, Rohan Anil, and Sanjiv Kumar. 2024 · 2024
Later among the works it cites.
Your Student is Better than Expected: Adaptive Teacher-Student Collaboration for Text-Conditional Diffusion Models. In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Seattle, WA, USA, 9275–9285
Nikita Starodubcev, Dmitry Baranchuk, Artem Fedorov, and Artem Babenko. 2024 · 2024
Later among the works it cites.
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning. In Proceedings of The Twelfth International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 25 pages
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. 2024 · 2024
Later among the works it cites.
Edge-Cloud Routing for Text-to-Image Model with Token-Level Multi-Metric Prediction
Zewei Xin, Qinya Li, Chaoyue Niu, and Fan Wu. 2024 · 2024
Later among the works it cites.
WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–22
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. 2024 · 2024
Later among the works it cites.
Federated Optimization Under Intermittent Client Availability
Yikai Yan, Chaoyue Niu, Yucheng Ding, Zhenzhe Zheng, Shaojie Tang, Qinya Li, Fan Wu, Chengfei Lyu, Yanghe Feng, and Guihai Chen. 2024 · 2024
Later among the works it cites.
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 1–22
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 2024 · 2024
Later among the works it cites.
Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient Reasoning. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Vienna, Austria, 38 pages
Murong Yue, Jie Zhao, Min Zhang, Liang Du, and Ziyu Yao. 2024 · 2024
Later among the works it cites.
Kaiyan Zhang, Jianyu Wang, Ning Ding, Biqing Qi, Ermo Hua, Xingtai Lv, and Bowen Zhou. 2024a · 2024
Later among the works it cites.
Revisiting Knowledge Distillation for Autoregressive Language Models. In Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL) . Association for Computational Linguistics, Bangkok, Thailand, 10900–10913
Qihuang Zhong, Liang Ding, Li Shen, Juhua Liu, Bo Du, and Dacheng Tao. 2024 · 2024
Later among the works it cites.
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., Vancouver, BC, Canada, 4819–4851
Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, and Yu Qiao. 2024 · 2024
Later among the works it cites.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI. 2025 · 2025
Closest in time.
Personalized Language Model Learning on Text Data Without User Identifiers. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . ACM, Toronto, ON, Canada, 12 pages
Yucheng Ding, Yangwenjian Tan, Xiangyu Liu, Chaoyue Niu, Fandong Meng, Jie Zhou, Ning Liu, Fan Wu, and Guihai Chen. 2025 · 2025
Closest in time.
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Bowen Jin, Hansi Zeng, Zhenrui Yue, Dong Wang, Hamed Zamani, and Jiawei Han. 2025 · 2025
Closest in time.
Collaboration of Large Language Models and Small Recommendation Models for Device-Cloud Recommendation. In Proceedings of ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . ACM, Toronto, ON, Canada, 12 pages
Zheqi Lv, Tianyu Zhan, Wenjie Wang, Xinyu Lin, Shengyu Zhang, Wenqiao Zhang, Jiwei Li, Kun Kuang, and Fei Wu. 2025 · 2025
Closest in time.
RouteLLM: Learning to Route LLMs from Preference Data. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Singapore, 16 pages
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. 2025 · 2025
Closest in time.
EmbedLLM: Learning Compact Representations of Large Language Models. In Proceedings of International Conference on Learning Representations (ICLR) . OpenReview.net, Singapore, 14 pages
Richard Zhuang, Tianhao Wu, Zhaojin Wen, Andrew Li, Jiantao Jiao, and Kannan Ramchandran. 2025 · 2025
Closest in time.
Exploiting Shared Representations for Personalized Federated Learning. In Proceedings of International Conference on Machine Learning (ICML) . PMLR, Virtual, 2089–2099
Liam Collins, Hamed Hassani, Aryan Mokhtari, and Sanjay Shakkottai. 2021 · 2099
Closest in time.