Fetching the paper…
Reading the bibliography…
As large language models (LLMs) evolve, deploying them solely in the cloud or compressing them for edge devices has become inadequate due to concerns about latency, privacy, cost, and personalization.
Optimized cost per click in taobao display advertising
Han Zhu, Junqi Jin, Chang Tan, et al · 2017
Earlier work this paper cites.
Auto-tuning neural network quantization framework for collaborative inference between the cloud and edge
Guangli Li, Lei Liu, Xueying Wang, et al · 2018
Earlier work this paper cites.
General data protection regulation
Protection Regulation · 2018
Earlier work this paper cites.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, et al · 2018
Earlier work this paper cites.
Leaf: A benchmark for federated settings
Sebastian Caldas, Sai Meher Karthik Duddu, Peter Wu, et al · 2019
Earlier work this paper cites.
Personalized dialogue generation with diversified traits
Yinhe Zheng, Guanyi Chen, Minlie Huang, et al · 2019
Earlier work this paper cites.
Towards federated learning at scale: System design
Kallista A. Bonawitz, Hubert Eichner, Wolfgang Grieskamp, et al · 2019
Earlier work this paper cites.
Edge cloud offloading algorithms: Issues, methods, and perspectives
Jianyu Wang, Jianli Pan, Flavio Esposito, et al · 2019
Earlier work this paper cites.
Federated visual classification with real-world data distribution
Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown · 2020
Earlier work this paper cites.
Tinybert: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, et al · 2020
Earlier work this paper cites.
Fedml: A research library and benchmark for federated machine learning
Chaoyang He, Songze Li, Jinhyun So, et al · 2020
Earlier work this paper cites.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, et al · 2020
Earlier work this paper cites.
Cloud-edge collaborative intelligent inference based on distributed neural networks in power distribution networks
Hao Luo, Hui Tian, Peng Zhang, et al · 2021
Earlier work this paper cites.
Parameter-efficient domain knowledge integration from multiple sources for biomedical pre-trained language models
Qiuhao Lu, Dejing Dou, and Thien Huu Nguyen · 2021
Earlier work this paper cites.
Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices
Yang Hu, Connor Imes, Xuanang Zhao, et al · 2022
Earlier work this paper cites.
Fedscale: Benchmarking model and system performance of federated learning at scale
Fan Lai, Yinwei Dai, Sanjay Sri Vallabh Singapuram, et al · 2022
Earlier work this paper cites.
FedNLP: Benchmarking federated learning methods for natural language processing tasks
Bill Yuchen Lin, Chaoyang He, Zihang Ze, et al · 2022
Earlier work this paper cites.
pfl-bench: A comprehensive benchmark for personalized federated learning
Daoyuan Chen, Dawei Gao, Weirui Kuang, et al · 2022
Earlier work this paper cites.
On-device model fine-tuning with label correction in recommender systems
Yucheng Ding, Chaoyue Niu, Fan Wu, et al · 2022
Earlier work this paper cites.
Federatedscope-gnn: Towards a unified, comprehensive and efficient package for federated graph learning
Zhen Wang, Weirui Kuang, Yuexiang Xie, et al · 2022
Earlier work this paper cites.
Flower: A friendly federated learning research framework
Daniel J. Beutel, Taner Topal, Akhil Mathur, et al · 2022
Earlier work this paper cites.
GLM: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, et al · 2022
Earlier work this paper cites.
Galactica: A large language model for science
Ross Taylor, Marcin Kardas, Guillem Cucurull, et al · 2022
Earlier work this paper cites.
Top ten technology trends of damo academy, 2022
Alibaba Damo Academy · 2022
Earlier work this paper cites.
Intelligent request strategy design in recommender system
Xufeng Qian, Yue Xu, Fuyu Lv, et al · 2022
Earlier work this paper cites.
Walle: An end-to-end,general-purpose, and large-scale production system for device-cloud collaborative machine learning
Chengfei Lv, Chaoyue Niu, Renjie Gu, et al · 2022
Earlier work this paper cites.
Online security-aware and reliability-guaranteed ai service chains provisioning in edge intelligence cloud
Yu Qiu, Junbin Liang, Victor CM Leung, et al · 2023
Earlier work this paper cites.
Privatelora for efficient privacy preserving llm
Yiming Wang, Yu Lin, Xiaodong Zeng, et al · 2023
Earlier work this paper cites.
Hybrid retrieval-augmented generation for real-time composition assistance
Xuchao Zhang, Menglin Xia, Camille Couturier, et al · 2023
Earlier work this paper cites.
Dc-ccl: Device-cloud collaborative controlled learning for large vision models
Yucheng Ding, Chaoyue Niu, Fan Wu, et al · 2023
Earlier work this paper cites.
Large language models (llms) inference offloading and resource allocation in cloud-edge networks: An active inference approach
Jingcheng Fang, Ying He, F Richard Yu, et al · 2023
Earlier work this paper cites.
Llmcad: Fast and scalable on-device large language model inference
Daliang Xu, Wangsong Yin, Xin Jin, et al · 2023
Earlier work this paper cites.
Speculative decoding with big little decoder
Sehoon Kim, Karttikeya Mangalam, Suhong Moon, et al · 2023
Earlier work this paper cites.
Edgemoe: Fast on-device inference of moe-based large language models
Rongjie Yi, Liwei Guo, Shiyun Wei, et al · 2023
Earlier work this paper cites.
Small models are valuable plug-ins for large language models
Canwen Xu, Yichong Xu, Shuohang Wang, et al · 2023
Earlier work this paper cites.
Mutual enhancement of large and small language models with cross-silo knowledge transfer
Yongheng Deng, Ziqing Qiao, Ju Ren, et al · 2023
Earlier work this paper cites.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, et al · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Earlier work this paper cites.
Duet: A tuning-free device-cloud collaborative parameters generation framework for efficient device model generalization
Zheqi Lv, Wenqiao Zhang, Shengyu Zhang, et al · 2023
Earlier work this paper cites.
Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation
Heming Xia, Tao Ge, Peiyi Wang, et al · 2023
Earlier work this paper cites.
Cloud-edge collaborative inference with network pruning
Mingran Li, Xuejun Zhang, Jiasheng Guo, et al · 2023
Earlier work this paper cites.
High-efficiency device-cloud collaborative transformer model
Penghao Jiang, Ke Xin, Chunxi Li, et al · 2023
Earlier work this paper cites.
Context-aware layer scheduling for seamless neural network inference in cloud-edge systems
Matthias Stammler, Vladimir Sidorenko, Fabian Kreß, et al · 2023
Earlier work this paper cites.
Accelerating llm inference with staged speculative decoding
Benjamin Spector and Chris Re · 2023
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, et al · 2023
Earlier work this paper cites.
Livechat: A large-scale personalized dialogue dataset automatically constructed from live streaming
Jingsheng Gao, Yixin Lian, Ziyi Zhou, et al · 2023
Earlier work this paper cites.
Fedmultimodal: A benchmark for multimodal federated learning
Tiantian Feng, Digbalay Bose, Tuo Zhang, et al · 2023
Earlier work this paper cites.
Tabi: An efficient multi-level inference system for large language models
Yiding Wang, Kai Chen, Haisheng Tan, et al · 2023
Earlier work this paper cites.
Self-knowledge guided retrieval augmentation for large language models
Yile Wang, Peng Li, Maosong Sun, et al · 2023
Earlier work this paper cites.
Toolformer: language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, et al · 2023
Earlier work this paper cites.
Knowledge card: Filling llms’ knowledge gaps with plug-in specialized language models
Shangbin Feng, Weijia Shi, Yuyang Bai, et al · 2023
Earlier work this paper cites.
Device-unimodal cloud-multimodal collaboration for livestreaming content understanding
Yufei Zhu, Chaoyue Niu, Yikai Yan, et al · 2023
Earlier work this paper cites.
Pushing large language models to the 6g edge: Vision, challenges, and opportunities, 2023
Zheng Lin, Guanqiao Qu, Qiyuan Chen, Xianhao Chen, Zhe Chen, and Kaibin Huang · 2023
Earlier work this paper cites.
A bargaining game for personalized, energy efficient split learning over wireless networks
Minsu Kim, Alexander DeRieux, and Walid Saad · 2023
Earlier work this paper cites.
Cloud-device collaborative adaptation to continual changing environments in the real-world
Yulu Gan, Mingjie Pan, Rongyu Zhang, et al · 2023
Earlier work this paper cites.
Efficient edge inference by selective query
Anil Kag, Igor Fedorov, Aditya Gangrade, et al · 2023
Earlier work this paper cites.
Spectr: fast speculative decoding via optimal transport
Ziteng Sun, Ananda Theertha Suresh, Jae Hun Ro, et al · 2023
Earlier work this paper cites.
Language models meet world models: Embodied experiences enhance language models
Jiannan Xiang, Tianhua Tao, Yi Gu, et al · 2023
Earlier work this paper cites.
Minillm: Large language models on consumer gpus
Volodymyr Kuleshov · 2023
Earlier work this paper cites.
Eclm: Efficient edge-cloud collaborative learning with continuous environment adaptation
Yan Zhuang, Zhenzhe Zheng, Yunfeng Shao, et al · 2023
Earlier work this paper cites.
Fast and robust early-exiting framework for autoregressive language models with synchronized parallel decoding
Sangmin Bae, Jongwoo Ko, Hwanjun Song, et al · 2023
Earlier work this paper cites.
Finding the SWEET spot: Analysis and improvement of adaptive inference in low resource settings
Daniel Rotem, Michael Hassid, Jonathan Mamou, et al · 2023
Earlier work this paper cites.
Edge-cloud polarization and collaboration: A comprehensive survey for ai
Jiangchao Yao, Shengyu Zhang, Yang Yao, et al · 2023
Earlier work this paper cites.
Bounding and filling: A fast and flexible framework for image captioning
Zheng Ma, Changxin Wang, Bo Huang, et al · 2023
Earlier work this paper cites.
UPOA: A user preference based latency and energy aware intelligent offloading approach for cloud-edge systems
Jingling Yuan, Yao Xiang, Yuhui Deng, et al · 2023
Earlier work this paper cites.
Llm-based edge intelligence: A comprehensive survey on architectures, applications, security and trustworthiness
Othmane Friha, Mohamed Amine Ferrag, Burak Kantarci, et al · 2024
Earlier work this paper cites.
Cambricon-llm: A chiplet-based hybrid architecture for on-device inference of 70b LLM
Zhongkai Yu, Shengwen Liang, Tianyun Ma, et al · 2024
Earlier work this paper cites.
Hybrid sd: Edge-cloud collaborative inference for stable diffusion models
Chenqian Yan, Songwei Liu, Hongjian Liu, et al · 2024
Earlier work this paper cites.
Crayon: Customized on-device llm via instant adapter blending and edge-server hybrid inference
Jihwan Bang, Juntae Lee, Kyuhong Shim, et al · 2024
Earlier work this paper cites.
Specexec: Massively parallel speculative decoding for interactive llm inference on consumer devices
Ruslan Svirschevski, Avner May, Zhuoming Chen, et al · 2024
Earlier work this paper cites.
Ce-collm: Efficient and adaptive large language models through cloud-edge collaboration
Hongpeng Jin and Yanzhao Wu · 2024
Earlier work this paper cites.
Hybrid slm and llm for edge-cloud collaborative inference
Zixu Hao, Huiqiang Jiang, Shiqi Jiang, et al · 2024
Earlier work this paper cites.
Efficient deployment of large language model across cloud-device systems
Fan Yang, Zehao Wang, Haoyu Zhang, et al · 2024
Earlier work this paper cites.
Yao Yao, Zuchao Li, and Hai Zhao · 2024
Earlier work this paper cites.
Large language models empowered autonomous edge ai for connected intelligence
Yifei Shen, Jiawei Shao, Xinjie Zhang, et al · 2024
Earlier work this paper cites.
Octo-planner: On-device language model for planner-action agents
Wei Chen, Zhiyuan Li, Zhen Guo, et al · 2024
Earlier work this paper cites.
A survey on model compression for large language models
Xunyu Zhu, Jian Li, Yong Liu, et al · 2024
Earlier work this paper cites.
Unleashing the power of edge-cloud generative ai in mobile networks: A survey of aigc services
Minrui Xu, Hongyang Du, Dusit Niyato, et al · 2024
Earlier work this paper cites.
Crayon: Customized on-device LLM via instant adapter blending and edge-server hybrid inference
Jihwan Bang, Juntae Lee, Kyuhong Shim, et al · 2024
Earlier work this paper cites.
Cogenesis: A framework collaborating large and small language models for secure context-aware instruction following
Kaiyan Zhang, Jianyu Wang, Remo Hua, et al · 2024
Earlier work this paper cites.
Kaiyan Zhang, Jianyu Wang, Ning Ding, et al · 2024
Earlier work this paper cites.
Cloud-edge collaborative large model services: Challenges and solutions
Yanghe Pan, Zhou Su, Yuntao Wang, et al · 2024
Earlier work this paper cites.
Kangaroo: Lossless self-speculative decoding for accelerating llms via double early exiting
Fangcheng Liu, Yehui Tang, Zhenhua Liu, et al · 2024
Earlier work this paper cites.
Specexec: Massively parallel speculative decoding for interactive llm inference on consumer devices
Ruslan Svirschevski, Avner May, Zhuoming Chen, et al · 2024
Earlier work this paper cites.
Enhancing on-device llm inference with historical cloud-based llm interactions
Yucheng Ding, Chaoyue Niu, Fan Wu, et al · 2024
Earlier work this paper cites.
Velo: A vector database-assisted cloud-edge collaborative llm qos optimization framework
Zhi Yao, Zhiqing Tang, Jiong Lou, et al · 2024
Earlier work this paper cites.
Cloud-device collaborative learning for multimodal large language models
Guanqun Wang, Jiaming Liu, Chenxuan Li, et al · 2024
Earlier work this paper cites.
Resource allocation for stable llm training in mobile edge computing
Chang Liu and Jun Zhao · 2024
Earlier work this paper cites.
Large language models (llms) inference offloading and resource allocation in cloud-edge computing: An active inference approach
Ying He, Jingcheng Fang, F. Richard Yu, et al · 2024
Earlier work this paper cites.
Hybridrag: Integrating knowledge graphs and vector retrieval augmented generation for efficient information extraction
Bhaskarjit Sarmah, Dhagash Mehta, Benika Hall, et al · 2024
Earlier work this paper cites.
Edge-cloud collaborative motion planning for autonomous driving with large language models
Jiao Chen, Suyan Dai, Fangfang Chen, et al · 2024
Earlier work this paper cites.
Grey-box prompt optimization and fine-tuning for cloud-edge LLM agents, 2024
Ya Liu, Kai Yang, Yu Zhu, et al · 2024
Earlier work this paper cites.
Slim: Speculative decoding with hypothesis reduction
Chi-Heng Lin, Shikhar Tuli, James Seale Smith, et al · 2024
Earlier work this paper cites.
Diffusion-based cloud-edge-device collaborative learning for next poi recommendations
Jing Long, Guanhua Ye, Tong Chen, et al · 2024
Earlier work this paper cites.
An edge-cloud collaboration framework for generative AI service provision with synergetic big cloud model and small edge models
Yuqing Tian, Zhaoyang Zhang, Yuzhi Yang, et al · 2024
Earlier work this paper cites.
Enhanced hybrid inference techniques for scalable on-device llm personalization and cloud integration
Teresa Peng, Liam Liu, Maya Gupta, et al · 2024
Earlier work this paper cites.
Edge-llm: A collaborative framework for large language model serving in edge computing
Fenglong Cai, Dong Yuan, Zhe Yang, et al · 2024
Earlier work this paper cites.
Diet: Customized slimming for incompatible networks in sequential recommendation
Kairui Fu, Shengyu Zhang, Zheqi Lv, et al · 2024
Earlier work this paper cites.
Modelgpt: Unleashing llm’s capabilities for tailored model generation
Zihao Tang, Zheqi Lv, Shengyu Zhang, et al · 2024
Cited alongside, same era.
Intelligent model update strategy for sequential recommendation
Zheqi Lv, Wenqiao Zhang, Zhengyu Chen, et al · 2024
Cited alongside, same era.
Aug-kd: Anchor-based mixup generation for out-of-domain knowledge distillation
Zihao Tang, Zheqi Lv, Shengyu Zhang, et al · 2024
Cited alongside, same era.
Llmco4mr: Llms-aided neural combinatorial optimization for ancient manuscript restoration from fragments with case studies on dunhuang
Yuqing Zhang, Hangqi Li, Shengyu Zhang, et al · 2024
Cited alongside, same era.
Mpod123: One image to 3d content generation using mask-enhanced progressive outline-to-detail optimization
Jimin Xu, Tianbao Wang, Tao Jin, et al · 2024
Cited alongside, same era.
Fuzzy speculative decoding for a tunable accuracy-runtime tradeoff
Maximilian Holsman, Yukun Huang, and Bhuwan Dhingra · 2025
Closest in time.
Duodecoding: Hardware-aware heterogeneous speculative decoding with dynamic multi-sequence drafting
Kai Lv, Honglin Guo, Qipeng Guo, and Xipeng Qiu · 2025
Closest in time.
Longspec: Long-context speculative decoding with efficient drafting and verification
Penghui Yang, Cunxiao Du, Fengzhuo Zhang, et al · 2025
Closest in time.
Collaboration of large language models and small recommendation models for device-cloud recommendation
Zheqi Lv, Tianyu Zhan, Wenjie Wang, et al · 2025
Closest in time.
Backpropagation-free multi-modal on-device model adaptation via cloud-device collaboration
Wei Ji, Li Li, Zheqi Lv, et al · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tao Fan, Yan Kang, Guoqiang Ma, et al · 2024
Cited alongside, same era.
Minillm: Knowledge distillation of large language models
Yuxian Gu, Li Dong, Furu Wei, et al · 2024
Cited alongside, same era.
Improving in-context learning with small language model ensembles
M. Mehdi Mojarradi, Lingyi Yang, Robert McCraith, et al · 2024
Cited alongside, same era.
Kaiyan Zhang, Jianyu Wang, Ning Ding, et al · 2024
Cited alongside, same era.
Modular pluralism: Pluralistic alignment via multi-llm collaboration
Shangbin Feng, Taylor Sorensen, Yuhan Liu, et al · 2024
Cited alongside, same era.
Automix: Automatically mixing language models
Pranjal Aggarwal, Aman Madaan, Ankit Anand, et al · 2024
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, et al · 2024
Cited alongside, same era.
Closest in time.
Reward-guided speculative decoding for efficient llm reasoning
Baohao Liao, Yuhui Xu, Hanze Dong, et al · 2025
Closest in time.
Adaserve: Slo-customized llm serving with fine-grained speculative decoding
Zikun Li, Zhuofu Chen, Remi Delacourt, et al · 2025
Closest in time.
Guiding reasoning in small language models with llm assistance
Yujin Kim, Euiin Yi, Minu Kim, et al · 2025
Closest in time.
Disco: Device-server collaborative llm-based text streaming services
Ting Sun, Penghan Wang, and Fan Lai · 2025
Closest in time.
Pice: A semantic-driven progressive inference system for llm serving in cloud-edge networks
Huiyou Zhan, Xuan Zhang, Haisheng Tan, et al · 2025
Closest in time.
Federated transfer learning for on-device llms efficient fine tuning optimization
Chuantao Li, Bruce Gu, Zhigang Zhao, et al · 2025
Closest in time.
Division-of-thoughts: Harnessing hybrid language model synergy for efficient on-device agents
Chenyang Shao, Xinyuan Hu, Yutang Lin, et al · 2025
Closest in time.
Adaptlink: A heterogeneity-aware adaptive framework for distributed mllm inference
Xinyi Hu, Zihan Chen, Kun Guo, et al · 2025
Closest in time.
A cloud-edge collaborative architecture for multimodal llm-based advanced driver assistance systems in iot networks
Yaqi Hu, Dongdong Ye, Jiawen Kang, et al · 2025
Closest in time.
Opt-tree: Speculative decoding with adaptive draft tree structure
Jikai Wang, Yi Su, Juntao Li, et al · 2025
Closest in time.
Swift: On-the-fly self-speculative decoding for llm inference acceleration
Heming Xia, Yongqi Li, Jun Zhang, et al · 2025
Closest in time.
Mergenet: Knowledge migration across heterogeneous models, tasks, and modalities
Kunxi Li, Tianyu Zhan, Kairui Fu, et al · 2025
Closest in time.
Optimize incompatible parameters through compatibility-aware knowledge integration
Zheqi Lv, Keming Ye, Zishu Wei, et al · 2025
Closest in time.
Forward once for all: Structural parameterized adaptation for efficient cloud-coordinated on-device recommendation
Kairui Fu, Zheqi Lv, Shengyu Zhang, et al · 2025
Closest in time.
Edge vs cloud: How do we balance cost, latency, and quality for large language models over 5g networks?
Minsu Kim, Pinyarash Pinyoanuntapong, Bong-Ho Kim, et al · 2025
Closest in time.
Edge-cloud collaborative computing on distributed intelligence and model optimization: A survey
Jing Liu, Yao Du, Kun Yang, et al · 2025
Closest in time.
Fedcfa: Alleviating simpson’s paradox in model aggregation with counterfactual federated learning
Zhonghua Jiang, Jimin Xu, Shengyu Zhang, et al · 2025
Closest in time.
Llm-empowered embodied agent for memory-augmented task planning in household robotics
Marc Glocker, Peter Hönig, Matthias Hirschmanner, et al · 2025
Closest in time.
Ran Xu, Wenqi Shi, Yuchen Zhuang, et al · 2025
Closest in time.
Knowledge-decoupled synergetic learning: An MLLM based collaborative approach to few-shot multimodal dialogue intention recognition
Bin Chen, Yu Zhang, Hongfei Ye, et al · 2025
Closest in time.
Yu-Neng Chuang, Leisheng Yu, Guanchu Wang, et al · 2025
Closest in time.
Efficient multitask learning in small language models through upside-down reinforcement learning
Yu-Chen Lin, Sanat Sharma, Hari Manikandan, et al · 2025
Closest in time.
Towards harnessing the collaborative power of large and small models for domain tasks
Yang Liu, Bingjie Yan, Tianyuan Zou, et al · 2025
Closest in time.
Slmrec: Distilling large language models into small for sequential recommendation
Wujiang Xu, Qitian Wu, Zujie Liang, et al · 2025
Closest in time.
Routellm: Learning to route llms from preference data
Isaac Ong, Amjad Almahairi, Vincent Wu, et al · 2025
Closest in time.
Llm for mobile: An initial roadmap
Daihang Chen, Yonghui Liu, Mingyi Zhou, et al · 2025
Closest in time.
Edgeshard: Efficient llm inference via collaborative edge computing
Mingjin Zhang, Xiaoming Shen, Jiannong Cao, et al · 2025
Closest in time.
Beyond the cloud: Edge inference for generative large language models in wireless networks
Xinyuan Zhang, Jiangtian Nie, Yudong Huang, et al · 2025
Closest in time.
A cloud-edge collaborative architecture for multimodal llm-based advanced driver assistance systems in iot networks
Yaqi Hu, Dongdong Ye, Jiawen Kang, et al · 2025
Closest in time.
Efficientllm: Scalable pruning-aware pretraining for architecture-agnostic edge language models
Xingrun Xing, Zheng Liu, Shitao Xiao, et al · 2025
Closest in time.
Federated reinforcement learning-empowered task offloading for large models in vehicular edge computing
Huaming Wu, Anqi Gu, and Yonghui Liang · 2025
Closest in time.
Traversal verification for speculative tree decoding
Yepeng Weng, Qiao Hu, Xujie Chen, et al · 2025
Closest in time.
Pipespec: Breaking stage dependencies in hierarchical llm decoding
Bradley McDanel, Sai Qian Zhang, Yunhai Hu, et al · 2025
Closest in time.
Hamburger: Accelerating llm inference via token smashing
Jingyu Liu and Ce Zhang · 2025
Closest in time.
Minions: Cost-efficient collaboration between on-device and cloud language models
Avanika Narayan, Dan Biderman, Sabri Eyuboglu, et al · 2025
Closest in time.
Gemini: A family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, et al · 2025
Closest in time.
A survey on collaborative mechanisms between large and small language models
Yi Chen, JiaHao Zhao, and HaoHao Han · 2025
Closest in time.
Chaoyue Niu, Yucheng Ding, Junhui Lu, et al · 2025
Closest in time.
Transitioning from mlops to llmops: Navigating the unique challenges of large language models
Saurabh Pahune and Zahid Akhtar · 2025
Closest in time.
MixLLM: Dynamic routing in mixed large language models
Xinyuan Wang, Yanchi Liu, Wei Cheng, et al · 2025
Closest in time.
Blade: Enhancing black-box large language models with small domain-specific models
Haitao Li, Qingyao Ai, Jia Chen, et al · 2025
Closest in time.
Citer: Collaborative inference for efficient large language model decoding with token-level routing
Wenhao Zheng, Yixiao Chen, Weitong Zhang, et al · 2025
Closest in time.
Jiaxing Li, Chi Xu, Lianchen Jia, et al · 2025
Closest in time.
Empirical guidelines for deploying llms onto resource-constrained edge devices
Ruiyang Qin, Dancheng Liu, Chenhui Xu, et al · 2025
Closest in time.
Accelerating inference of retrieval-augmented generation via sparse context selection
Yun Zhu, Jia-Chen Gu, Caitlin Sikora, et al · 2025
Closest in time.
Robust implementation of retrieval-augmented generation on edge-based computing-in-memory architectures
Ruiyang Qin, Zheyu Yan, Dewen Zeng, et al · 2025
Closest in time.
Pearl: Parallel speculative decoding with adaptive draft length
Tianyu Liu, Yun Li, Qitan Lv, et al · 2025
Closest in time.
Facil: Flexible DRAM address mapping for soc-pim cooperative on-device LLM inference
Seong Hoon Seo, Junghoon Kim, Donghyun Lee, et al · 2025
Closest in time.
Lincoln: Real-time 50~100b LLM inference on consumer devices with lpddr-interfaced, compute-enabled flash memory
Weiyi Sun, Mingyu Gao, Zhaoshi Li, et al · 2025
Closest in time.
Falcon: Faster and parallel inference of large language models through enhanced semi-autoregressive drafting and custom-designed decoding tree
Xiangxiang Gao, Weisheng Xie, Yiwei Xiang, et al · 2025
Closest in time.
Hierarchical speculative decoding with dynamic window
Shensian Syu and Hung-yi Lee · 2025
Closest in time.
Weak-to-strong search: align large language models via searching over small language models
Zhanhui Zhou, Zhixuan Liu, Jie Liu, et al · 2025
Closest in time.
Flash: Latent-aware semi-autoregressive speculative decoding for multimodal tasks
Zihua Wang, Ruibo Li, Haozhe Du, et al · 2025
Closest in time.
Accelerating greedy coordinate gradient and general prompt optimization via probe sampling
Yiran Zhao, Wenyue Zheng, Tianle Cai, et al · 2025
Closest in time.
Seed: Accelerating reasoning tree construction via scheduled speculative decoding
Zhenglin Wang, Jialong Wu, Yilong Lai, et al · 2025
Closest in time.
Speculative rag: Enhancing retrieval augmented generation through drafting
Zilong Wang, Zifeng Wang, Long Le, et al · 2025
Closest in time.
Speculative knowledge distillation: Bridging the teacher-student gap through interleaved sampling
Wenda Xu, Rujun Han, Zifeng Wang, et al · 2025
Closest in time.
Accelerated diffusion models via speculative sampling
Valentin De Bortoli, Alexandre Galashov, Arthur Gretton, et al · 2025
Closest in time.
Banditspec: Adaptive speculative decoding via bandit algorithms
Cunxiao Du Yunlong Hou, Fengzhuo Zhang et al · 2025
Closest in time.
Fast large language model collaborative decoding via speculation
Jiale Fu, Yuchu Jiang, Junkai Chen, et al · 2025
Closest in time.
Diffusion models are secretly exchangeable: Parallelizing ddpms via autospeculation
Hengyuan Hu, Aniket Das, Dorsa Sadigh, et al · 2025
Closest in time.
Swift: On-the-fly self-speculative decoding for llm inference acceleration
Heming Xia, Yongqi Li, Jun Zhang, et al · 2025
Closest in time.
Efficient inference for large language model-based generative recommendation
Xinyu Lin, Chaoqun Yang, Wenjie Wang, et al · 2025
Closest in time.
An efficient private gpt never autoregressively decodes
Zhengyi Li, Yue Guan, Kang Yang, et al · 2025
Closest in time.
Speculate, then collaborate: Fusing knowledge of language models during decoding
Ziyao Wang, Muneeza Azmat, Ang Li, et al · 2025
Closest in time.
Accelerating inference of retrieval-augmented generation via sparse context selection
Yun Zhu, Jia-Chen Gu, Caitlin Sikora, et al · 2025
Closest in time.
Qwen, An Yang, Baosong Yang, et al · 2025
Closest in time.
Deploying foundation model powered agent services: A survey
Wenchao Xu, Jinyu Chen, Peirong Zheng, et al · 2025
Closest in time.
Early-exit deep neural network - A comprehensive survey
Haseena Rahmath P, Vishal Srivastava, Kuldeep Chaurasia, et al · 2025
Closest in time.
Ecoagent: An efficient edge-cloud collaborative multi-agent framework for mobile automation
Biao Yi, Xavier Hu, Yurun Chen, et al · 2025
Closest in time.
Large-scale multi-agent learning-based cloud–edge collaborative distributed pv data compression and information aggregation for multimodal network in power systems
Junhao Feng, Boyang Huang, Xiaodong Zhou, et al · 2025
Closest in time.
Luoxi models
LuoXi Team · 2025
Closest in time.
Infiguiagent: A multimodal generalist gui agent with native reasoning and reflection
Yuhang Liu, Pengxiang Li, Zishu Wei, et al · 2025
Closest in time.
Hybrid edge-ai framework for intelligent mobile applications: Leveraging large language models for on-device contextual assistance and code-aware automation
Liao Hu · 2025
Closest in time.
Llm bandit: Cost-efficient llm generation via preference-conditioned dynamic routing
Yang Li · 2025
Closest in time.
The moe-empowered edge llms deployment: Architecture, challenges, and opportunities
Ning Li, Song Guo, Tuo Zhang, et al · 2025
Closest in time.
Slimrag: Retrieval without graphs via entity-aware context selection
Jiale Zhang, Jiaxiang Chen, Zhucong Li, et al · 2025
Closest in time.
Arag: Agentic retrieval augmented generation for personalized recommendation
Reza Yousefi Maragheh, Pratheek Vadla, Priyank Gupta, et al · 2025
Closest in time.
Cross-attention speculative decoding
Wei Zhong, Manasa Bharadwaj, Yixiao Wang, et al · 2025
Closest in time.
Graft: Integrating the domain knowledge via efficient parameter synergy for mllms
Yang Dai, Jianxiang An, Tianwei Lin, et al · 2025
Closest in time.
Huaiying Luo and Cheng Ji · 2025
Closest in time.
Chitranshu Harbola and Anupam Purwar · 2025
Closest in time.
A comprehensive survey in llm(-agent) full stack safety: Data, training and deployment
Kun Wang, Guibin Zhang, Zhenhong Zhou, et al · 2025
Closest in time.
Hawkeye:efficient reasoning with model collaboration
Jianshu She, Zhuohao Li, Zhemin Huang, et al · 2025
Closest in time.
Vita-1.5: Towards gpt-4o level real-time vision and speech interaction
Chaoyou Fu, Haojia Lin, Xiong Wang, et al · 2025
Closest in time.
Harnessing multiple large language models: A survey on llm ensemble
Zhijun Chen, Jingzheng Li, Pengpeng Chen, et al · 2025
Closest in time.
DeepSeek-AI, Aixin Liu, Bei Feng, et al · 2025
Closest in time.
Spin: Accelerating large language model inference with heterogeneous speculative models
Fahao Chen, Peng Li, Tom H. Luan, et al · 2025
Closest in time.
Fast and cost-effective speculative edge-cloud decoding with early exits, 2025
Yeshwanth Venkatesha, Souvik Kundu, and Priyadarshini Panda · 2025
Closest in time.
Differentially private low-rank adaptation of large language model using federated learning
Xiao-Yang Liu, Rongyi Zhu, Daochen Zha, et al · 2025
Closest in time.
QLLMS: quantization-adaptive LLM scheduling for partially informed edge serving systems
Miao Hu, Qi He, and Di Wu · 2025
Closest in time.
Enhancing stability and resource efficiency in llm training for edge-assisted mobile systems
Chang Liu and Jun Zhao · 2025
Closest in time.