Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved remarkable progress across domains and applications but face challenges such as high fine-tuning costs, inference latency, limited edge deployability, and reliability concerns.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
TinyBERT: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Earlier work this paper cites.
De-co: A two-step spelling correction model for combating adversarial typos
Zhengxiao Liu, Fali Wang, Zheng Lin, Lei Wang, and Zhiyi Yin. 2020 · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020 · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, and 1 others. 2021 · 2021
Earlier work this paper cites.
{DEBERTA}: {DECODING}-{enhanced} {bert} {with} {disentangled} {attention}
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Macrobert: Maximizing certified region of bert to adversarial word substitutions
Fali Wang, Zheng Lin, Zhengxiao Liu, Mingyu Zheng, Lei Wang, and Daren Zha. 2021 · 2021
Earlier work this paper cites.
Biogpt: generative pre-trained transformer for biomedical text generation and mining
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022 · 2022
Earlier work this paper cites.
Neural Text Sanitization with Explicit Measures of Privacy Risk
Anthi Papadopoulou, Yunhao Yu, Pierre Lison, and Lilja Øvrelid. 2022 · 2022
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023 · 2023
Earlier work this paper cites.
On the privacy risk of in-context learning
Haonan Duan, Adam Dziedzic, Mohammad Yaghini, Nicolas Papernot, and Franziska Boenisch. 2023 · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and 1 others. 2023 · 2023
Earlier work this paper cites.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Dongfu Jiang, Xiang Ren, and Bill Yuchen Lin. 2023 · 2023
Earlier work this paper cites.
Speculative decoding with big little decoder
Sehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik, Michael W Mahoney, Amir Gholami, and Kurt Keutzer. 2023 · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Earlier work this paper cites.
Maximum entropy loss, the silver bullet targeting backdoor attacks in pre-trained language models
Zhengxiao Liu, Bowen Shen, Zheng Lin, Fali Wang, and Weiping Wang. 2023 · 2023
Earlier work this paper cites.
Duet: A tuning-free device-cloud collaborative parameters generation framework for efficient device model generalization
Zheqi Lv, Wenqiao Zhang, Shengyu Zhang, Kun Kuang, Feng Wang, Yongwei Wang, Zhengyu Chen, Tao Shen, Hongxia Yang, Beng Chin Ooi, and 1 others. 2023 · 2023
Earlier work this paper cites.
CombLM: Adapting Black-Box Language Models through Small Fine-Tuned Models
Aitor Ormazabal, Mikel Artetxe, and Eneko Agirre. 2023 · 2023
Earlier work this paper cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023 · 2023
Earlier work this paper cites.
Accelerating llm inference with staged speculative decoding
Benjamin Spector and Chris Re. 2023 · 2023
Earlier work this paper cites.
Spectr: Fast speculative decoding via optimal transport
Ziteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami, Himanshu Jain, and Felix Yu. 2023 · 2023
Earlier work this paper cites.
On provable copyright protection for generative models
Nikhil Vyas, Sham M. Kakade, and Boaz Barak. 2023 · 2023
Earlier work this paper cites.
Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui. 2023 · 2023
Earlier work this paper cites.
Offsite-tuning: Transfer learning without full model
Guangxuan Xiao, Ji Lin, and Song Han. 2023 · 2023
Earlier work this paper cites.
Llmcad: Fast and scalable on-device large language model inference
Daliang Xu, Wangsong Yin, Xin Jin, Ying Zhang, Shiyun Wei, Mengwei Xu, and Xuanzhe Liu. 2023 · 2023
Earlier work this paper cites.
Perplexed by perplexity: Perplexity-based data pruning with small reference models
Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L Leavitt, and Mansheej Paul. 2024 · 2024
Earlier work this paper cites.
General collaborative framework between large language model and experts for universal information extraction
K Bao and Ning Wang. 2024 · 2024
Earlier work this paper cites.
Think big, generate quick: Llm-to-slm for fast autoregressive decoding
Benjamin Bergner, Andrii Skliar, Amelie Royer, Tijmen Blankevoort, Yuki Asano, and Babak Ehteshami Bejnordi. 2024 · 2024
Earlier work this paper cites.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D Lee, Deming Chen, and Tri Dao. 2024 · 2024
Earlier work this paper cites.
What is the role of small models in the llm era: A survey
Lihu Chen and Gaël Varoquaux. 2024 · 2024
Earlier work this paper cites.
Llama guard 3 vision: Safeguarding human-ai image understanding conversations
Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan, Eric Smith, Javier Rando, Yiming Zhang, Kate Plawiak, Zacharie Delpierre Coudert, Kartikeya Upasani, and Mahesh Pasupuleti. 2024 · 2024
Earlier work this paper cites.
Casper: Prompt Sanitization for Protecting User Privacy in Web-Based Large Language Models
Chun Jie Chong, Chenxi Hou, Zhihao Yao, and Seyed Mohammadjavad Seyed Talebi. 2024 · 2024
Earlier work this paper cites.
Yuwei Du, Jie Feng, Jie Zhao, Jian Yuan, and Yong Li. 2024 · 2024
Earlier work this paper cites.
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, and 1 others. 2024 · 2024
Earlier work this paper cites.
Combining small language models and large language models for zero-shot nl2sql
Ju Fan, Zihui Gu, Songyue Zhang, Yuxin Zhang, Zui Chen, Lei Cao, Guoliang Li, Samuel Madden, Xiaoyong Du, and Nan Tang. 2024 · 2024
Earlier work this paper cites.
Detecting hallucinations in large language models using semantic entropy
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024 · 2024
Earlier work this paper cites.
Llama guard 3-1b-int4: Compact and efficient safeguard for human-ai conversations
Igor Fedorov, Kate Plawiak, Lemeng Wu, Tarek Elgamal, Naveen Suda, Eric Smith, Hongyuan Zhan, Jianfeng Chi, Yuriy Hulovatyy, Kimish Patel, and 1 others. 2024 · 2024
Earlier work this paper cites.
Enhancing guardrails for safe and secure healthcare ai
Ananya Gangavarapu. 2024 · 2024
Earlier work this paper cites.
FedPT: Federated Proxy-Tuning of Large Language Models on Resource-Constrained Edge Devices
Zhidong Gao, Yu Zhang, Zhenxiao Zhang, Yanmin Gong, and Yuanxiong Guo. 2024 · 2024
Earlier work this paper cites.
MiniLLM: Knowledge distillation of large language models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024 · 2024
Earlier work this paper cites.
Apple intelligence foundation language models
Tom Gunter, Zirui Wang, Chong Wang, Ruoming Pang, Andy Narayanan, Aonan Zhang, Bowen Zhang, Chen Chen, Chung-Cheng Chiu, David Qiu, and 1 others. 2024 · 2024
Earlier work this paper cites.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. 2024 · 2024
Earlier work this paper cites.
Hybrid slm and llm for edge-cloud collaborative inference
Zixu Hao, Huiqiang Jiang, Shiqi Jiang, Ju Ren, and Ting Cao. 2024 · 2024
Earlier work this paper cites.
Can LLMs get help from other LLMs without revealing private information?
Florian Hartmann, Duc-Hieu Tran, Peter Kairouz, Victor Cărbune, and Blaise Aguera Y Arcas. 2024 · 2024
Earlier work this paper cites.
CPT: Consistent Proxy Tuning for Black-box Optimization
Yuanyang He, Zitong Huang, Xinxing Xu, Rick Siow Mong Goh, Salman Khan, Wangmeng Zuo, Yong Liu, and Chun-Mei Feng. 2024 · 2024
Earlier work this paper cites.
Calm: Contrasting large and small language models to verify grounded generation
I-Hung Hsu, Zifeng Wang, Long Le, Lesly Miculicich Werlen, Nanyun Peng, Chen-Yu Lee, and Tomas Pfister. 2024 · 2024
Cited alongside, same era.
On efficient distillation from llms to slms
Metod Jazbec, Menglin Xia, Ankur Mallick, Daniel Madrigal, Dongge Han, Samuel Kessler, and Victor Rühle. 2024 · 2024
Cited alongside, same era.
Synergistic weak-strong collaboration by aligning preferences
Yizhu Jiao, Xuchao Zhang, Zhaoyang Wang, Yubo Ma, Zhun Deng, Rujia Wang, Chetan Bansal, Saravan Rajmohan, Jiawei Han, and Huaxiu Yao. 2024 · 2024
Cited alongside, same era.
BioMistral: A collection of open-source pretrained large language models for medical domains
Yanis Labrak, Adrien Bazoge, Emmanuel Morin, Pierre-Antoine Gourraud, Mickael Rouvier, and Richard Dufour. 2024 · 2024
Cited alongside, same era.
Can small language models help large language models reason better?: Lm-guided chain-of-thought
Jooyoung Lee, Fan Yang, Thanh Tran, Qian Hu, Emre Barut, and Kai-Wei Chang. 2024 · 2024
Cited alongside, same era.
RemoteRAG: A privacy-preserving LLM cloud RAG service
Yihang Cheng, Lan Zhang, Junyang Wang, Mu Yuan, and Yunhao Yao. 2025 · 2025
Closest in time.
Llamafirewall: An open source guardrail system for building secure ai agents
Sahana Chennabasappa, Cyrus Nikolaidis, Daniel Song, David Molnar, Stephanie Ding, Shengye Wan, Spencer Whitman, Lauren Deason, Nicholas Doucette, Abraham Montilla, and 1 others. 2025 · 2025
Closest in time.
Crosslm: A data-free collaborative fine-tuning framework for large and small language models
Yongheng Deng, Ziqing Qiao, Ye Zhang, Zhenya Ma, Yang Liu, and Ju Ren. 2025 · 2025
Closest in time.
BEST-route: Adaptive LLM routing with test-time optimal compute
Dujian Ding, Ankur Mallick, Shaokun Zhang, Chi Wang, Daniel Madrigal, Mirian Del Carmen Hipolito Garcia, Menglin Xia, Laks V. S. Lakshmanan, Qingyun Wu, and Victor Rühle. 2025 · 2025
Closest in time.
LoRA-x: Bridging foundation models with training-free cross-model adaptation
Farzad Farhadzadeh, Debasmit Das, Shubhankar Borse, and Fatih Porikli. 2025 · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Synergizing large language models and pre-trained smaller models for conversational intent discovery
Jinggui Liang, Lizi Liao, Hao Fei, and Jing Jiang. 2024 · 2024
Cited alongside, same era.
Large model strategic thinking, small model efficiency: Transferring theory of mind in large language models
Nunzio Lorè, Alireza Sepehr Ilami, and Babak Heydari. 2024 · 2024
Cited alongside, same era.
Small language models: Survey, measurements, and insights
Zhenyan Lu, Xiang Li, Dongqi Cai, Rongjie Yi, Fangming Liu, Xiwen Zhang, Nicholas D Lane, and Mengwei Xu. 2024 · 2024
Cited alongside, same era.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, and 1 others. 2024 · 2024
Cited alongside, same era.
An emulator for fine-tuning large language models using small language models
Eric Mitchell, Rafael Rafailov, Archit Sharma, Chelsea Finn, and Christopher D Manning. 2024 · 2024
Cited alongside, same era.
Improving in-context learning with small language model ensembles
M. Mehdi Mojarradi, Lingyi Yang, Robert McCraith, and Adam Mahdi. 2024 · 2024
Cited alongside, same era.
When in doubt, cascade: Towards building efficient and capable guardrails
Manish Nagireddy, Inkit Padhi, Soumya Ghosh, and Prasanna Sattigeri. 2024 · 2024
Cited alongside, same era.
Closest in time.
Elm: Ensemble of language models for predicting tumor group from pathology reports
Lovedeep Gondara, Jonathan Simkin, Shebnum Devji, Gregory Arbour, and Raymond Ng. 2025 · 2025
Closest in time.
Distilled pretraining: A modern lens of data, in-context learning and test-time scaling
Sachin Goyal, David Lopez-Paz, and Kartik Ahuja. 2025 · 2025
Closest in time.
MiniPLM: Knowledge distillation for pre-training language models
Yuxian Gu, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2025 · 2025
Closest in time.
Accelerating diffusion llms via adaptive parallel decoding
Daniel Israel, Guy Van den Broeck, and Aditya Grover. 2025 · 2025
Closest in time.
Moe 2 : Optimizing collaborative inference for edge large language models
Lyudong Jin, Yanning Zhang, Yanhan Li, Shurong Wang, Howard H Yang, Jian Wu, and Meng Zhang. 2025 · 2025
Closest in time.
Distilling llm agent into small models with retrieval and code tools
Minki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho, and Sung Ju Hwang. 2025 · 2025
Closest in time.
Plug-in and fine-tuning: Bridging the gap between small language models and large language models
Kyeonghyun Kim, Jinhee Jang, Juhwan Choi, Yoonji Lee, Kyohoon Jin, and YoungBin Kim. 2025 · 2025
Closest in time.
Hispec: Hierarchical speculative decoding for llms
Avinash Kumar, Sujay Sanghavi, and Poulami Das. 2025 · 2025
Closest in time.
Learning with less: Knowledge distillation from large language models via unlabeled data
Juanhui Li, Sreyashi Nag, Hui Liu, Xianfeng Tang, Sheikh Muhammad Sarwar, Limeng Cui, Hansu Gu, Suhang Wang, Qi He, and Jiliang Tang. 2025a · 2025
Closest in time.
Synergistic augmentation: Enhancing cross-domain zero-shot slot filling with small model-assisted large language models
Weizhen Li, Junbao Huang, Peijie Huang, Yuhong Xu, and Jiekun Fan. 2025d · 2025
Closest in time.
Small models struggle to learn from strong reasoners
Yuetai Li, Xiang Yue, Zhangchen Xu, Fengqing Jiang, Luyao Niu, Bill Yuchen Lin, Bhaskar Ramasubramanian, and Radha Poovendran. 2025g · 2025
Closest in time.
Collaboration of large language models and small recommendation models for device-cloud recommendation
Zheqi Lv, Tianyu Zhan, Wenjie Wang, Xinyu Lin, Shengyu Zhang, Wenqiao Zhang, Jiwei Li, Kun Kuang, and Fei Wu. 2025 · 2025
Closest in time.
Causal distillation: Transferring structured explanations from large to compact language models
Aggrey Muhebwa and Khalid K Osman. 2025 · 2025
Closest in time.
Chaoyue Niu, Yucheng Ding, Junhui Lu, Zhengxiang Huang, Hang Zeng, Yutong Dai, Xuezhen Tu, Chengfei Lv, Fan Wu, and Guihai Chen. 2025 · 2025
Closest in time.
RouteLLM: Learning to route LLMs from preference data
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. 2025 · 2025
Closest in time.
LlamaDuo: LLMOps pipeline for seamless migration from service LLMs to small-scale local LLMs
Chansung Park, Juyong Jiang, Fan Wang, Sayak Paul, and Jing Tang. 2025 · 2025
Closest in time.
Pre-training distillation for large language models: A design space exploration
Hao Peng, Xin Lv, Yushi Bai, Zijun Yao, Jiajie Zhang, Lei Hou, and Juanzi Li. 2025 · 2025
Closest in time.
Quang PM Pham, Khoi TN Nguyen, Nhi H Doan, Cuong A Pham, Kentaro Inui, and Dezhen Song. 2025 · 2025
Closest in time.
Mixture of small and large models for Chinese spelling check
Ziheng Qiao, Houquan Zhou, and Zhenghua Li. 2025 · 2025
Closest in time.
Gatekeeper: Improving model cascades through confidence tuning
Stephan Rabanser, Nathalie Rauschmayr, Achin Kulshrestha, Petra Poklukar, Wittawat Jitkrittum, Sean Augenstein, Congchao Wang, and Federico Tombari. 2025 · 2025
Closest in time.
Division-of-thoughts: Harnessing hybrid language model synergy for efficient on-device agents
Chenyang Shao, Xinyuan Hu, Yutang Lin, and Fengli Xu. 2025 · 2025
Closest in time.
Hawkeye: Model collaboration for efficient reasoning
Jianshu She, Zhuohao Li, Zhemin Huang, Qi Li, Peiran Xu, Haonan Li, and Qirong Ho. 2025 · 2025
Closest in time.
Mobillama: Towards accurate & lightweight fully transparent GPT
Omkar Chakradhar Thawakar, Ashmal Vayani, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Michael Felsberg, Timothy Baldwin, Eric P. Xing, and Fahad Shahbaz Khan. 2025 · 2025
Closest in time.
Beyond answers: Transferring reasoning capabilities to smaller llms using multi-teacher knowledge distillation
Yijun Tian, Yikun Han, Xiusi Chen, Wei Wang, and Nitesh V Chawla. 2025 · 2025
Closest in time.
Guarding your conversations: Privacy gatekeepers for secure interactions with cloud-based ai models
GodsGift Uzor, Hasan Al-Qudah, Ynes Ineza, and Abdul Serwadda. 2025 · 2025
Closest in time.
Phlora: data-free post-hoc low-rank adapter extraction from full-rank checkpoint
Bhoomit Vasani, Jack FitzGerald, Anjie Fang, and Sushmit Vaish. 2025 · 2025
Closest in time.
Fast and cost-effective speculative edge-cloud decoding with early exits
Yeshwanth Venkatesha, Souvik Kundu, and Priyadarshini Panda. 2025 · 2025
Closest in time.
Mixllm: Dynamic routing in mixed large language models
Xinyuan Wang, Yanchi Liu, Wei Cheng, Xujiang Zhao, Zhengzhang Chen, Wenchao Yu, Yanjie Fu, and Haifeng Chen. 2025c · 2025
Closest in time.
ThinkGuard: Deliberative slow thinking leads to cautious guardrails
Xiaofei Wen, Wenxuan Zhou, Wenjie Jacky Mo, and Muhao Chen. 2025 · 2025
Closest in time.
Image corruption-inspired membership inference attacks against large vision-language models
Zongyu Wu, Minhua Lin, Zhiwei Zhang, Fali Wang, Xianren Zhang, Xiang Zhang, and Suhang Wang. 2025 · 2025
Closest in time.
Cross-LoRA: A Data-Free LoRA Transfer Framework across Heterogeneous LLMs
Feifan Xia, Mingyang Liao, Yuyang Fang, Defang Li, Yantong Xie, Weikang Li, Yang Li, Deguo Xia, and Jizhou Huang. 2025 · 2025
Closest in time.
Ran Xu, Wenqi Shi, Yuchen Zhuang, Yue Yu, Joyce C Ho, Haoyu Wang, and Carl Yang. 2025 · 2025
Closest in time.
Hades: Hardware accelerated decoding for efficient speculation in large language models
Ze Yang, Yihong Jin, and Xinhe Xu. 2025 · 2025
Closest in time.
Efficient construction of model family through progressive training using model expansion
Kazuki Yano, Sho Takase, Sosuke Kobayashi, Shun Kiyono, and Jun Suzuki. 2025 · 2025
Closest in time.
GradOT: Training-free gradient-preserving offsite-tuning for large language models
Kai Yao, Zhaorui Tan, Penglei Gao, Lichun Li, Kaixin Wu, Yinggui Wang, Yuan Zhao, Yixin Ji, Jianke Zhu, and Wei Wang. 2025 · 2025
Closest in time.
Shieldgemma 2: Robust and tractable image content moderation
Wenjun Zeng, Dana Kurniawan, Ryan Mullins, Yuchi Liu, Tamoghna Saha, Dirichi Ike-Njoku, Jindong Gu, Yiwen Song, Cai Xu, Jingjing Zhou, and 1 others. 2025 · 2025
Closest in time.
Pice: A semantic-driven progressive inference system for llm serving in cloud-edge networks
Huiyou Zhan, Xuan Zhang, Haisheng Tan, Han Tian, Dongping Yong, Junyang Zhang, and Xiang-Yang Li. 2025 · 2025
Closest in time.
Weak-to-strong jailbreaking on large language models
Xuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du, Lei Li, Yu-Xiang Wang, and William Yang Wang. 2025 · 2025
Closest in time.
Cross-attention speculative decoding
Wei Zhong, Manasa Bharadwaj, Yixiao Wang, Nikhil Verma, Yipeng Ji, and Chul Lee. 2025 · 2025
Closest in time.
Efficient llm inference over heterogeneous edge networks with speculative decoding
Bingjie Zhu, Zhixiong Chen, Liqiang Zhao, Hyundong Shin, and Arumugam Nallanathan. 2025 · 2025
Closest in time.