Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) deliver powerful AI capabilities but face deployment challenges due to high resource costs and latency, whereas Small Language Models (SLMs) offer efficiency and deployability at the cost of reduced performance.
“Albert: A lite bert for self-supervised learning of language representations”
Zhenzhong Lan et al · 2019
Earlier work this paper cites.
“Language models are few-shot learners”
Tom Brown, Benjamin Mann and Nick Ryder · 2020
Earlier work this paper cites.
“Speculative Decoding for Fast and Accurate Language Generation”
Jiannan Chen, Xia Liu and Di He · 2023
Earlier work this paper cites.
“Mutual enhancement of large and small language models with cross-silo knowledge transfer”
Yongheng Deng et al · 2023
Earlier work this paper cites.
“Design, Implementation, and Practical Evaluation of a Voice Recognition Based IoT Home Automation System for Low-Resource Languages and Resource-Constrained Edge IoT Devices: A System for Galician and Mobile Opportunistic Scenarios”
Iván Froiz-Miguez, Paula Fraga-Lamas and Tiago Fernández-CaraméS · 2023
Earlier work this paper cites.
“MiniLLM: Knowledge distillation of large language models”
Yuxian Gu, Li Dong, Furu Wei and Minlie Huang · 2023
Earlier work this paper cites.
“Adaptive Cascading for Efficient Large Language Model Inference”
Siddharth Gupta, Samyam Rajbhandari and Di Zhao · 2023
Earlier work this paper cites.
“Tiny machine learning: Progress and futures [feature]”
Ji Lin et al · 2023
Earlier work this paper cites.
“A comprehensive overview of large language models”
Humza Naveed et al · 2023
Earlier work this paper cites.
“GPT-4 Technical Report” https://openai.com/research/gpt-4 , 2023
OpenAI · 2023
Earlier work this paper cites.
“Aizip Works with SoftBank Corp. to Launch Customized Small Language Model Solutions for Privacy-Critical Enterprise Applications”, 2024
Aizip · 2024
Earlier work this paper cites.
“Llm for mobile: An initial roadmap”
Daihang Chen et al · 2024
Earlier work this paper cites.
“Hybrid llm: Cost-efficient and quality-aware query routing”
Dujian Ding et al · 2024
Earlier work this paper cites.
“Hymba: A hybrid-head architecture for small language models”
Xin Dong et al · 2024
Earlier work this paper cites.
“FedCoLLM: A Parameter-Efficient Federated Co-tuning Framework for Large and Small Language Models”
Tao Fan et al · 2024
Earlier work this paper cites.
“Modular pluralism: Pluralistic alignment via multi-llm collaboration”
Shangbin Feng et al · 2024
Earlier work this paper cites.
“Llm-based edge intelligence: A comprehensive survey on architectures, applications, security and trustworthiness”
Othmane Friha et al · 2024
Earlier work this paper cites.
“MiniLLM: Compressing LLMs with Targeted Distillation”
Jiawei Gu, Yanan Ren and Yang Lin · 2024
Earlier work this paper cites.
“Smoothie: Label free language model routing”
Neel Guha et al · 2024
Earlier work this paper cites.
“Hybrid slm and llm for edge-cloud collaborative inference”
Zixu Hao et al · 2024
Earlier work this paper cites.
“Auxiliary task demands mask the capabilities of smaller language models”
Jennifer Hu and Michael Frank · 2024
Earlier work this paper cites.
“Introducing Apple Intelligence” https://www.apple.com/newsroom/ , 2024
Apple Inc · 2024
Earlier work this paper cites.
“The State of Edge AI”, https://peri-labs.github.io/docs/assets/files/The_State_of_Edge_AI.pdf , 2024
A. Jayant, M. Sheldon, S. Kim and S. Shrivastava · 2024
Earlier work this paper cites.
“CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration”
Hongpeng Jin and Yanzhao Wu · 2024
Earlier work this paper cites.
“Large language models (LLMs) for semantic communication in edge-based IoT networks”
Alakesh Kalita · 2024
Earlier work this paper cites.
“Thoughtful things: Building human-centric smart devices with small language models”
Evan King et al · 2024
Earlier work this paper cites.
“A Contemporary Overview: Trends and Applications of Large Language Models on Mobile Devices”
Lianjun Liu, Hongli An, Pengxuan Chen and Longxiang Ye · 2024
Earlier work this paper cites.
“A multimodal generative AI copilot for human pathology”
Ming Lu et al · 2024
Earlier work this paper cites.
“Pack of llms: Model fusion at test-time via perplexity optimization”
Costas Mavromatis, Petros Karypis and George Karypis · 2024
Earlier work this paper cites.
“Routoo: Learning to Route to Large Language Models Effectively”
Alireza Mohammadshahi, Arshad Shaikh and Majid Yazdani · 2024
Earlier work this paper cites.
“Improving In-Context Learning with Small Language Model Ensembles”
M Mojarradi, Lingyi Yang, Robert McCraith and Adam Mahdi · 2024
Earlier work this paper cites.
“RouteLLM: Learning to Route LLMs with Preference Data. arXiv. org”, 2024
I Ong et al · 2024
Earlier work this paper cites.
Tianyu Peng and Jiajun Zhang · 2024
Cited alongside, same era.
“Large language models meet user interfaces: The case of provisioning feedback”
Stanislav Pozdniakov et al · 2024
Cited alongside, same era.
“Vilbias: A framework for bias detection using linguistic and visual cues”
Shaina Raza et al · 2024
Cited alongside, same era.
“Bias Amplification in Language Model Evolution: An Iterated Learning Perspective”
Yi Ren et al · 2024
Cited alongside, same era.
“ProFuser: Progressive Fusion of Large Language Models”
Tianyuan Shi et al · 2024
“LLM-Empowered Embodied Agent for Memory-Augmented Task Planning in Household Robotics”
M. Glocker, P. Hönig and M. Hirschmanner · 2025
Closest in time.
Chitranshu Harbola and Anupam Purwar · 2025
Closest in time.
Daniel Hendriks, Philipp Spitzer, Niklas Kühl and Gerhard Satzger · 2025
Closest in time.
“Dynamic Low-Rank Sparse Adaptation for Large Language Models”
Weizhong Huang et al · 2025
Closest in time.
“From Large to Small: The Rise of Small Language Models (SLMs) in Text Analytics”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Llumnix: Dynamic scheduling for large language model serving”
Biao Sun et al · 2024
Cited alongside, same era.
“Small language model is a good guide for large language model in Chinese entity relation extraction”
Xuemei Tang, Jun Wang and Qi Su · 2024
Cited alongside, same era.
“HarmonyOS and Pangu Model Integration” Huawei Developer Conference, 2024
Huawei Technologies · 2024
Cited alongside, same era.
“Robots saving lives: A literature review about search and rescue (sar) in harsh environments”
Kailin Tong et al · 2024
Cited alongside, same era.
“Knowledge fusion of large language models”
Fanqi Wan et al · 2024
Cited alongside, same era.
“LLM-SLM Collaboration for Efficient NLP Systems: Pipelines and Triggers”
Fan Wang, Li Zhang and Jian Hu · 2024
Cited alongside, same era.
“Personalized large language models”
Stanisław Woźniak et al · 2024
Cited alongside, same era.
Akshi Kumar · 2025
Closest in time.
“Multi-Agent Geospatial Copilots for Remote Sensing Workflows”
Chaehong Lee et al · 2025
Closest in time.
“Life-Cycle Routing Vulnerabilities of LLM Router”
Qiqi Lin et al · 2025
Closest in time.
“Efficient Multitask Learning in Small Language Models Through Upside-Down Reinforcement Learning”
Y Lin, S Sharma and H Manikandan · 2025
Closest in time.
“Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks”
Yang Liu et al · 2025
Closest in time.
Zheqi Lv et al · 2025
Closest in time.
“Minions: Cost-efficient Collaboration Between On-device and Cloud Language Models”
Avanika Narayan et al · 2025
Closest in time.
Chaoyue Niu et al · 2025
Closest in time.
“Transitioning from MLOps to LLMOps: Navigating the Unique Challenges of Large Language Models”
Saurabh Pahune and Zahid Akhtar · 2025
Closest in time.
“MathFusion: Enhancing Mathematic Problem-solving of LLM through Instruction Fusion”
Qizhi Pei et al · 2025
Closest in time.
“Mobile edge intelligence for large language models: A contemporary survey”
Guanqiao Qu et al · 2025
Closest in time.
“Improving consistency in large language models through chain of guidance”
Harsh Raj, Vipul Gupta, Domenic Rosati and Subhabrata Majumdar · 2025
Closest in time.
“Division-of-thoughts: Harnessing hybrid language model synergy for efficient on-device agents”
Chenyang Shao, Xinyuan Hu, Yutang Lin and Fengli Xu · 2025
Closest in time.
“Hawkeye: Efficient Reasoning with Model Collaboration”
J She, Z Li and Z Huang · 2025
Closest in time.
Y Shen, C Fu and S Dong · 2025
Closest in time.
“Small Language Models (SLMs) Can Still Pack a Punch: A survey”
Shreyas Subramanian, Vikram Elango and Mecit Gungor · 2025
Closest in time.
Clovis Varangot-Reille et al · 2025
Closest in time.
“Rema: Learning to meta-think for LLMs with multi-agent reinforcement learning”
Zhen Wan, Yuyu Li and Yachao Song · 2025
Closest in time.
“Mixllm: Dynamic routing in mixed large language models”
Xinyuan Wang et al · 2025
Closest in time.
“Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding”
Ziyao Wang et al · 2025
Closest in time.
Ran Xu et al · 2025
Closest in time.
“Collaborative Stance Detection via Small-Large Language Model Consistency Verification”
Yu Yan et al · 2025
Closest in time.
“Toward Super Agent System with Hybrid AI Routers”
Yuhang Yao et al · 2025
Closest in time.
“LLMOps in Production: 457 Case Studies of What Actually Works”, 2025
ZenML Blog · 2025
Closest in time.
“On Efficient Deployment of LLMs in Edge Devices: Trends and Challenges”
Qi Zhang, Yuan Sun and Tong Liu · 2025
Closest in time.
Y Zhang, S Qiao and J Zhang · 2025
Closest in time.
Wenhao Zheng et al · 2025
Closest in time.
“A review on edge large language models: Design, execution, and applications”
Yue Zheng et al · 2025
Closest in time.