Fetching the paper…
Reading the bibliography…
With the rapid growth in the number of large language model (LLM) users, it is difficult for bandwidth-constrained cloud servers to simultaneously process massive LLM services in real-time.
Cloud computing: state-of-the-art and research challenges
Q. Zhang, L. Cheng, and R. Boutaba · 2010
Earlier work this paper cites.
Upper-confidence-bound algorithms for active learning in multi-armed bandits
A. Carpentier, A. Lazaric, M. Ghavamzadeh, R. Munos, and P. Auer · 2011
Earlier work this paper cites.
Combinatorial multi-armed bandit with general reward functions
W. Chen, W. Hu, F. Li, J. Li, Y. Liu, and P. Lu · 2016
Earlier work this paper cites.
Edge computing: Vision and challenges
W. Shi, J. Cao, Q. Zhang, Y. Li, and L. Xu · 2016
Earlier work this paper cites.
Osmotic computing: A new paradigm for edge/cloud integration
M. Villari, M. Fazio, S. Dustdar, O. Rana, and R. Ranjan · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
{ \{ MArk
C. Zhang, M. Yu, W. Wang, and F. Yan · 2019
Earlier work this paper cites.
Large-scale many-objective deployment optimization of edge servers
B. Cao, S. Fan, J. Zhao, S. Tian, Z. Zheng, Y. Yan, and P. Yang · 2021
Earlier work this paper cites.
Transformer in transformer
K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang · 2021
Earlier work this paper cites.
Combinatorial multi-armed bandits for resource allocation
J. Zuo and C. Joe-Wong · 2021
Earlier work this paper cites.
Sdps and robust satisfiability of promise csp
J. Brakensiek, V. Guruswami, and S. Sandeep · 2023
Cited alongside, same era.
Netgpt: A native-ai network architecture beyond provisioning personalized generative services
Y. Chen, R. Li, Z. Zhao, C. Peng, J. Wu, E. Hossain, and H. Zhang · 2023
Cited alongside, same era.
Large language models (llms) inference offloading and resource allocation in cloud-edge networks: An active inference approach
J. Fang, Y. He, F. R. Yu, J. Li, and V. C. Leung · 2023
Cited alongside, same era.
Model-as-a-service (maas): A survey
W. Gan, S. Wan, and S. Y. Philip · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica · 2023
Cited alongside, same era.
An empirical analysis and resource footprint study of deploying large language models on edge devices
N. Dhar, B. Deng, D. Lo, X. Wu, L. Zhao, and K. Suo · 2024
Closest in time.
Creating edge ai from cloud-based llms
Q. Dong, X. Chen, and M. Satyanarayanan · 2024
Closest in time.
Diffusion-based reinforcement learning for edge-enabled ai-generated content services
H. Du, Z. Li, D. Niyato, J. Kang, Z. Xiong, H. Huang, and S. Mao · 2024
Closest in time.
Deferred continuous batching in resource-efficient large language model serving
Y. He, Y. Lu, and G. Alonso · 2024
Closest in time.
S. Hu, R. Deng, X. Du, Z. Lu, Q. Duan, Y. He, S.-C. Huang, and J. Wu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Lin, G. Qu, Q. Chen, X. Chen, Z. Chen, and K. Huang · 2023
Cited alongside, same era.
Llm-pruner: On the structural pruning of large language models
X. Ma, G. Fang, and X. Wang · 2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang · 2023
Cited alongside, same era.
A survey on model compression for large language models
X. Zhu, J. Li, Y. Liu, C. Ma, and W. Wang · 2023
Cited alongside, same era.
Distributed inference and fine-tuning of large language models over the internet
A. Borzunov, M. Ryabinin, A. Chumachenko, D. Baranchuk, T. Dettmers, Y. Belkada, P. Samygin, and C. A. Raffel · 2024
Cited alongside, same era.
A survey on effective invocation methods of massive llm services
C. Wang, B. Zhang, D. Sui, Z. Tum, X. Liu, and J. Kang
Cited in the paper.
Toward scalable generative ai via mixture of experts in mobile edge networks
J. Wang, H. Du, D. Niyato, J. Kang, Z. Xiong, D. I. Kim, and K. B. Letaief
Cited in the paper.
S. Li, X. Lin, H. Xu, K. Hua, X. Jin, G. Li, and J. Li · 2024
Closest in time.
Y. Liu, X. Peng, J. Cao, L. Dai, X. Liu, W. Liu, and M. Wang · 2024
Closest in time.
Exegpt: Constraint-aware resource scheduling for llm inference
H. Oh, K. Kim, J. Kim, S. Kim, J. Lee, D.-s. Chang, and J. Seo · 2024
Closest in time.
Characterizing power management opportunities for llms in the cloud
P. Patel, E. Choukse, C. Zhang, Í. Goiri, B. Warrier, N. Mahalingam, and R. Bianchini · 2024
Closest in time.
Unleashing the power of edge-cloud generative ai in mobile networks: A survey of aigc services
M. Xu, H. Du, D. Niyato, J. Kang, Z. Xiong, S. Mao, Z. Han, A. Jamalipour, D. I. Kim, X. Shen, et al · 2024
Closest in time.