Fetching the paper…
Reading the bibliography…
Large language model (LLM) inference at the network edge is a promising serving paradigm that leverages distributed edge resources to run inference near users and enhance privacy.
S. M. Johnson, “Optimal two-and three-stage production schedules with setup times included,” Naval research logistics quarterly , vol. 1, no. 1, pp. 61–68, 1954
1954
Earlier work this paper cites.
J. Du, L. Zhao, J. Feng, and X. Chu, “Computation offloading and resource allocation in mixed fog/cloud computing systems with min-max fairness guarantee,” IEEE Transactions on Communications , vol. 66, no. 4, pp. 1594–1608, 2017
2017
Earlier work this paper cites.
X. Hu, L. Wang, K.-K. Wong, M. Tao, Y. Zhang, and Z. Zheng, “Edge and central cloud computing: A perfect pairing for high energy efficiency and low-latency,” IEEE Transactions on Wireless Communications , vol. 19, no. 2, pp. 1070–1083, 2020
2020
Earlier work this paper cites.
M. Yao, L. Chen, J. Zhang, J. Huang, and J. Wu, “Loading cost-aware model caching and request routing for cooperative edge inference,” in ICC 2022-IEEE International Conference on Communications . IEEE, 2022, pp. 2327–2332
2022
Earlier work this paper cites.
W. Shi, S. Zhou, Z. Niu, M. Jiang, and L. Geng, “Multiuser co-inference with batch processing capable edge server,” IEEE Transactions on Wireless Communications , vol. 22, no. 1, pp. 286–300, 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Liu, Q. Lan, and K. Huang, “Resource allocation for multiuser edge inference with batching and early exiting,” IEEE Journal on Selected Areas in Communications , vol. 41, no. 4, pp. 1186–1200, 2023
2023
Earlier work this paper cites.
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 274–19 286
2023
Earlier work this paper cites.
S. Kim, K. Mangalam, S. Moon, J. Malik, M. W. Mahoney, A. Gholami, and K. Keutzer, “Speculative decoding with big little decoder,” Advances in Neural Information Processing Systems , vol. 36, pp. 39 236–39 256, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang, “Flexgen: High-throughput generative inference of large language models with a single gpu,” in International Conference on Machine Learning . PMLR, 2023, pp. 31 094–31 116
2023
Earlier work this paper cites.
Z. Chen, W. Yi, Y. Liu, and A. Nallanathan, “Knowledge-aided federated learning for energy-limited wireless networks,” IEEE Transactions on Communications , vol. 71, no. 6, pp. 3368–3386, 2023
2023
Earlier work this paper cites.
Y. Xu, J. Sun, S. Zhou, and Z. Niu, “Smdp-based dynamic batching for efficient inference on gpu-based platforms,” in ICC 2023-IEEE International Conference on Communications . IEEE, 2023, pp. 5483–5489
2023
Cited alongside, same era.
Z. Chen, W. Yi, and A. Nallanathan, “Exploring representativity in device scheduling for wireless federated learning,” IEEE Transactions on Wireless Communications , vol. 23, no. 1, pp. 720–735, 2023
2023
Cited alongside, same era.
H. Zhou, C. Hu, Y. Yuan, Y. Cui, Y. Jin, C. Chen, H. Wu, D. Yuan, L. Jiang, D. Wu et al. , “Large language model (llm) for telecommunications: A comprehensive survey on principles, key techniques, and opportunities,” IEEE Communications Surveys & Tutorials , vol. 27, no. 3, pp. 1955–2005, 2024
2024
Cited alongside, same era.
Y. Cang, M. Chen, and K. Huang, “Joint batching and scheduling for high-throughput multiuser edge ai with asynchronous task arrivals,” IEEE Transactions on Wireless Communications , vol. 23, no. 10, pp. 13 782–13 795, 2024
H. Li, Y. Liu, Y. Cheng, S. Ray, K. Du, and J. Jiang, “Eloquent: A more robust transmission scheme for llm token streaming,” in Proceedings of the 2024 SIGCOMM Workshop on Networks for AI Computing , 2024, pp. 34–40
2024
Later among the works it cites.
H. Oh, K. Kim, J. Kim, S. Kim, J. Lee, D.-s. Chang, and J. Seo, “Exegpt: Constraint-aware resource scheduling for llm inference,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2024, pp. 369–384
2024
Later among the works it cites.
W. Zhao, W. Jing, Z. Lu, and X. Wen, “Edge and terminal cooperation enabled llm deployment optimization in wireless network,” in 2024 IEEE/CIC International Conference on Communications in China (ICCC Workshops) . IEEE, 2024, pp. 220–225
2024
Later among the works it cites.
G. Qu, Q. Chen, W. Wei, Z. Lin, X. Chen, and K. Huang, “Mobile edge intelligence for large language models: A contemporary survey,” IEEE Communications Surveys & Tutorials , 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
X. Li and S. Bi, “Optimal ai model splitting and resource allocation for device-edge co-inference in multi-user wireless sensing systems,” IEEE Transactions on Wireless Communications , vol. 23, no. 9, pp. 11 094–11 108, 2024
2024
Cited alongside, same era.
X. Zhang, J. Nie, Y. Huang, G. Xie, Z. Xiong, J. Liu, D. Niyato, and X. S. Shen, “Beyond the cloud: Edge inference for generative large language models in wireless networks,” IEEE Transactions on Wireless Communications , 2024
2024
Cited alongside, same era.
M. Zhang, X. Shen, J. Cao, Z. Cui, and S. Jiang, “Edgeshard: Efficient llm inference via collaborative edge computing,” IEEE Internet of Things Journal , 2024
2024
Cited alongside, same era.
Y. He, J. Fang, F. R. Yu, and V. C. Leung, “Large language models (llms) inference offloading and resource allocation in cloud-edge computing: An active inference approach,” IEEE Transactions on Mobile Computing , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, Z. Wang, Z. Zhang, R. Y. Y. Wong, A. Zhu, L. Yang, X. Shi et al. , “Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , 2024, pp. 932–949
2024
Cited alongside, same era.
NVIDIA Corporation, “NVIDIA Jetson TX2 Developer Kit,” https://developer.nvidia.com/embedded/jetson-tx2, 2017, accessed: 2024-12-19
2024
Cited alongside, same era.
B. Butler, S. Yu, A. Mazaheri, and A. Jannesari, “Pipeinfer: Accelerating llm inference using asynchronous pipelined speculation,” in SC24: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2024, pp. 1–19
2024
Cited alongside, same era.
2025
Closest in time.
Z. Lin, G. Qu, Q. Chen, X. Chen, Z. Chen, and K. Huang, “Pushing large language models to the 6g edge: Vision, challenges, and opportunities,” IEEE Communications Magazine , vol. 63, no. 9, pp. 52–59, 2025
2025
Closest in time.
M. Hu, Q. He, and D. Wu, “Qllms: Quantization-adaptive llm scheduling for partially informed edge serving systems,” in IEEE INFOCOM 2025 - IEEE Conference on Computer Communications , 2025, pp. 1–10
2025
Closest in time.
K. Zhang, H. He, S. Song, J. Zhang, and K. B. Letaief, “Communication-efficient distributed on-device llm inference over wireless networks,” IEEE Journal of Selected Topics in Signal Processing , pp. 1–16, 2025
2025
Closest in time.
G. Xie, Z. Xiong, R. Xie, X. Deng, S. Guo, M. Guizani, and Z. Han, “Mixture of experts-enabled parallel scheduling and processing for vehicular generative ai services,” IEEE Transactions on Cognitive Communications and Networking , 2025
2025
Closest in time.
B. Zhu, Z. Chen, L. Zhao, H. Shin, and A. Nallanathan, “Joint caching and inference for large language models in wireless networks,” in ICC 2025-IEEE International Conference on Communications . IEEE, 2025, pp. 6285–6290
2025
Closest in time.
J. A. Lab, “Deploying small language models on jetson: End-to-end tutorial,” https://www.jetson-ai-lab.com/tutorial_slm.html, 2024, accessed: 2025-06-17
2025
Closest in time.
F. Chen, P. Li, T. H. Luan, Z. Su, and J. Deng, “Spin: Accelerating large language model inference with heterogeneous speculative models,” in IEEE INFOCOM 2025-IEEE Conference on Computer Communications . IEEE, 2025, pp. 1–10
2025
Closest in time.