Fetching the paper…
Reading the bibliography…
With the advancement of serverless computing, running machine learning (ML) inference services over a serverless platform has been advocated, given its labor-free scalability and cost effectiveness.
K. Kotani, T. Yoshimi, and H. Isahara, “A machine learning approach to measurement of text readability for efl learners using various linguistic features.” Online Submission , 2011
2011
Earlier work this paper cites.
J. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl, “Algorithms for hyper-parameter optimization,” Advances in neural information processing systems , vol. 24, 2011
2011
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Elloumi and A. Lambert, “Global solution of non-convex quadratically constrained quadratic programs,” Optimization methods and software , vol. 34, no. 1, pp. 98–114, 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Ali, R. Pinciroli, F. Yan, and E. Smirni, “Batch: machine learning inference serving on serverless platforms with adaptive batching,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–15
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
H. Zhang, Y. Tang, A. Khandelwal, J. Chen, and I. Stoica, “Caerus: { \{ NIMBLE } \} task scheduling for serverless analytics,” in 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21) , 2021, pp. 653–669
2021
Earlier work this paper cites.
M. Yu, Z. Jiang, H. C. Ng, W. Wang, R. Chen, and B. Li, “Gillis: Serving large neural networks in serverless functions with automatic model partitioning,” in 2021 IEEE 41st International Conference on Distributed Computing Systems (ICDCS) . IEEE, 2021, pp. 138–148
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J. Jarachanthan, L. Chen, F. Xu, and B. Li, “Amps-inf: Automatic model partitioning for serverless inference with cost efficiency,” in Proceedings of the 50th International Conference on Parallel Processing , 2021, pp. 1–12
2021
Earlier work this paper cites.
F. Romero, Q. Li, N. J. Yadwadkar, and C. Kozyrakis, “ { \{ INFaaS } \} : Automated model-less inference serving,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21) , 2021, pp. 397–411
2021
Cited alongside, same era.
2021
Cited alongside, same era.
J. Li, L. Zhao, Y. Yang, K. Zhan, and K. Li, “Tetris: Memory-efficient serverless inference through tensor sharing,” in 2022 USENIX Annual Technical Conference (USENIX ATC 22) , 2022
2022
Cited alongside, same era.
A. Mampage, S. Karunasekera, and R. Buyya, “A holistic view on resource management in serverless computing environments: Taxonomy and future directions,” ACM Computing Surveys (CSUR) , vol. 54, no. 11s, pp. 1–36, 2022
2022
Cited alongside, same era.
S. Kounev, N. Herbst, C. L. Abad, A. Iosup, I. Foster, P. Shenoy, O. Rana, and A. A. Chien, “Serverless computing: What it is, and what it is not?” Communications of the ACM , vol. 66, no. 9, pp. 80–92, 2023
2023
Later among the works it cites.
J. Li, Y. Jiang, Y. Zhu, C. Wang, and H. Xu, “Accelerating distributed { \{ MoE } \} training and inference with lina,” in 2023 USENIX Annual Technical Conference (USENIX ATC 23) , 2023, pp. 945–959
2023
Later among the works it cites.
M. Zhai, J. He, Z. Ma, Z. Zong, R. Zhang, and J. Zhai, “ { \{ SmartMoE } \} : Efficiently training { \{ Sparsely-Activated } \} models through combining offline and online parallelization,” in 2023 USENIX Annual Technical Conference (USENIX ATC 23) , 2023, pp. 961–975
2023
Later among the works it cites.
S. Shi, X. Pan, X. Chu, and B. Li, “Pipemoe: Accelerating mixture-of-experts through adaptive pipelining,” in IEEE INFOCOM 2023-IEEE Conference on Computer Communications . IEEE, 2023, pp. 1–10
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research , vol. 23, no. 120, pp. 1–39, 2022
2022
Cited alongside, same era.
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat et al. , “Glam: Efficient scaling of language models with mixture-of-experts,” in International Conference on Machine Learning . PMLR, 2022, pp. 5547–5569
2022
Cited alongside, same era.
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y. Aminabadi, A. A. Awan, J. Rasley, and Y. He, “Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale,” in International Conference on Machine Learning . PMLR, 2022, pp. 18 332–18 346
2022
Cited alongside, same era.
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. M. Dai, Q. V. Le, J. Laudon et al. , “Mixture-of-experts with expert choice routing,” Advances in Neural Information Processing Systems , vol. 35, pp. 7103–7114, 2022
2022
Cited alongside, same era.
J. He, J. Zhai, T. Antunes, H. Wang, F. Luo, S. Shi, and Q. Li, “Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models,” in Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2022, pp. 120–134
2022
Cited alongside, same era.
Y. Yang, L. Zhao, Y. Li, H. Zhang, J. Li, M. Zhao, X. Chen, and K. Li, “Infless: a native serverless system for low-latency, high-throughput inference,” in Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2022, pp. 768–781
2022
Cited alongside, same era.
C. Jin, Z. Zhang, X. Xiang, S. Zou, G. Huang, X. Liu, and X. Jin, “Ditto: Efficient serverless analytics with elastic parallelism,” in Proceedings of the ACM SIGCOMM 2023 Conference , 2023, pp. 406–419
2023
Cited alongside, same era.
Aws lambda. [Online]. Available: https://aws.amazon.com/lambda/
Cited in the paper.
Later among the works it cites.
Z. Zhang, D. Yang, Y. Xia, L. Ding, D. Tao, X. Zhou, and D. Cheng, “Mpipemoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism,” in 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2023, pp. 167–177
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Nie, X. Miao, Z. Wang, Z. Yang, J. Xue, L. Ma, G. Cao, and B. Cui, “Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,” Proceedings of the ACM on Management of Data , vol. 1, no. 1, pp. 1–19, 2023
2023
Later among the works it cites.
W. Wang, Z. Lai, S. Li, W. Liu, K. Ge, Y. Liu, A. Shen, and D. Li, “Prophet: Fine-grained load balancing for parallel training of large-scale moe models,” in 2023 IEEE International Conference on Cluster Computing (CLUSTER) . IEEE, 2023, pp. 82–94
2023
Later among the works it cites.
2024
Later among the works it cites.
Y. Fu, L. Xue, Y. Huang, A.-O. Brabete, D. Ustiugov, Y. Patel, and L. Mai, “ { \{ ServerlessLLM } \} : { \{ Low-Latency } \} serverless inference for large language models,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) , 2024, pp. 135–153
2024
Later among the works it cites.