Fetching the paper…
Reading the bibliography…
Recent innovations in generative large language models (LLMs) have made their applications and use-cases ubiquitous.
W. Zhu, “Analysis of JSQ Policy on Soft Real-time Scheduling in Cluster,” in HPCAsia , 2000
2000
Earlier work this paper cites.
R. Kumar, K. I. Farkas, N. P. Jouppi, P. Ranganathan, and D. M. Tullsen, “Single-ISA Heterogeneous Multi-core Architectures: The Potential for Processor Power Reduction,” in MICRO , 2003
2003
Earlier work this paper cites.
V. Gupta, M. Harchol Balter, K. Sigman, and W. Whitt, “Analysis of Join-the-Shortest-Queue Routing for Web Server Farms,” Performance Evaluation , 2007
2007
Earlier work this paper cites.
M. Isard, M. Budiu, Y. Yu, A. Birrell, and D. Fetterly, “Dryad: Distributed Data-Parallel Programs from Sequential Building Blocks,” in EuroSys , 2007
2007
Earlier work this paper cites.
J. Dean and S. Ghemawat, “MapReduce: Simplified Data Processing on Large Clusters,” Communications of the ACM , 2008
2008
Earlier work this paper cites.
J. Chen and L. K. John, “Efficient Program Scheduling for Heterogeneous Multi-core Processors,” in DAC , 2009
2009
Earlier work this paper cites.
K. Shvachko, H. Kuang, S. Radia, and R. Chansler, “The Hadoop Distributed File System,” in MSST , 2010
2010
Earlier work this paper cites.
K. Van Craeynest, A. Jaleel, L. Eeckhout, P. Narvaez, and J. Emer, “Scheduling Heterogeneous Multi-cores Through Performance Impact Estimation (PIE),” ACM SIGARCH Computer Architecture News , 2012
2012
Earlier work this paper cites.
M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauly, M. J. Franklin, S. Shenker, and I. Stoica, “Resilient Distributed Datasets: A Fault-Tolerant Abstraction for In-Memory Cluster Computing,” in NSDI , 2012
2012
Earlier work this paper cites.
K. Ousterhout, P. Wendell, M. Zaharia, and I. Stoica, “Sparrow: Distributed, Low Latency Scheduling,” in SOSP , 2013
2013
Earlier work this paper cites.
M. Schwarzkopf, A. Konwinski, M. Abd-El-Malek, and J. Wilkes, “Omega: Flexible, Scalable Schedulers for Large Compute Clusters,” in EuroSys , 2013
2013
Earlier work this paper cites.
E. Boutin, J. Ekanayake, W. Lin, B. Shi, J. Zhou, Z. Qian, M. Wu, and L. Zhou, “Apollo: Scalable and Coordinated Scheduling for Cloud-scale Computing,” in OSDI , 2014
2014
Earlier work this paper cites.
M. E. Haque, Y. H. Eom, Y. He, S. Elnikety, R. Bianchini, and K. S. McKinley, “Few-to-Many: Incremental Parallelism for Reducing Tail Latency in Interactive Services,” ACM SIGPLAN Notices , 2015
2015
Earlier work this paper cites.
M. E. Haque, Y. He, S. Elnikety, T. D. Nguyen, R. Bianchini, and K. S. McKinley, “Exploiting Heterogeneity for Tail Latency and Energy Efficiency,” in MICRO , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is All You Need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
H. Yang, Q. Chen, M. Riaz, Z. Luan, L. Tang, and J. Mars, “PowerChief: Intelligent Power Allocation for Multi-stage Applications to Improve Responsiveness on Power Constrained CMP,” in ISCA , 2017
2017
Earlier work this paper cites.
L. A. Barroso, U. Hölzle, and P. Ranganathan, “The Datacenter as a Computer: Designing Warehouse-Scale Machines,” Synthesis Lectures on Computer Architecture , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan et al. , “Ray: A Distributed Framework for Emerging AI Applications,” in OSDI , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in NAACL , 2019
2019
Earlier work this paper cites.
R. S. Kannan, L. Subramanian, A. Raju, J. Ahn, J. Mars, and L. Tang, “GrandSLAm: Guaranteeing SLAs for Jobs in Microservices Execution Frameworks,” in EuroSys , 2019
2019
Cited alongside, same era.
Y. Kwon, Y. Lee, and M. Rhu, “TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning,” in MICRO , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” OpenAI blog , 2019
2019
Cited alongside, same era.
C. Zhang, M. Yu, W. Wang, and F. Yan, “MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving,” in USENIX ATC , 2019
2022
Later among the works it cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A Distributed Serving System for Transformer-Based Generative Models,” in OSDI , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Azure Public Dataset: Azure LLM Inference Trace 2023. GitHub. [Online]. Available: https://github.com/Azure/AzurePublicDataset/blob/master/AzureLLMInferenceDataset2023.md
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2020
Cited alongside, same era.
D. Crankshaw, G.-E. Sela, X. Mo, C. Zumar, I. Stoica, J. Gonzalez, and A. Tumanov, “InferLine: Latency-aware Provisioning and Scaling for Prediction Serving Pipelines,” in SoCC , 2020
2020
Cited alongside, same era.
A. Gujarati, R. Karimi, S. Alzayat, W. Hao, A. Kaufmann, Y. Vigfusson, and J. Mace, “Serving DNNs like Clockwork: Performance Predictability from the Bottom Up,” in OSDI , 2020
2020
Cited alongside, same era.
U. Gupta, S. Hsia, V. Saraph, X. Wang, B. Reagen, G.-Y. Wei, H.-H. S. Lee, D. Brooks, and C.-J. Wu, “DeepRecSys: A System for Optimizing End-to-end At-scale Neural Recommendation Inference,” in ISCA , 2020
2020
Cited alongside, same era.
R. Hwang, T. Kim, Y. Kwon, and M. Rhu, “Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized Recommendations,” in ISCA , 2020
2020
Cited alongside, same era.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. v. Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, “Transformers: State-of-the-art Natural Language Processing,” in EMNLP , 2020
2020
Cited alongside, same era.
Y. Hu, R. Ghosh, and R. Govindan, “Scrooge: A Cost-effective Deep Learning Inference System,” in SoCC , 2021
2021
Cited alongside, same era.
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
M. Javaheripi and S. Bubeck, “Phi-2: The Surprising Power of Small Language Models,” Microsoft Research Blog , 2023
2023
Closest in time.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient Memory Management for Large Language Model Serving with PagedAttention,” in SOSP , 2023
2023
Closest in time.
Z. Li, L. Zheng, Y. Zhong, V. Liu, Y. Sheng, X. Jin, Y. Huang, Z. Chen, H. Zhang, J. E. Gonzalez, and I. Stoica, “AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving,” in OSDI , 2023
2023
Closest in time.
Z. Liu, J. Wang, T. Dao, T. Zhou, B. Yuan, Z. Song, A. Shrivastava, C. Zhang, Y. Tian, C. Re, and B. Chen, “Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time,” in ICML , 2023
2023
Closest in time.
2023
Closest in time.
P. Patel, Z. Gong, S. Rizvi, E. Choukse, P. Misra, T. Anderson, and A. Sriraman, “Towards Improved Power Management in Cloud GPUs,” in IEEE CAL , 2023
2023
Closest in time.
2023
Closest in time.
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently Scaling Transformer Inference,” in MLSys , 2023
2023
Closest in time.
2023
Closest in time.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang, “FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU,” in ICML , 2023
2023
Closest in time.
2023
Closest in time.
P. Patel, E. Choukse, C. Zhang, Í. Goiri, B. Warrier, N. Mahalingam, and R. Bianchini, “Characterizing Power Management Opportunities for LLMs in the Cloud,” in ASPLOS , 2024
2024
Closest in time.