Fetching the paper…
Reading the bibliography…
The growing demand for Large Language Models (LLMs) in applications such as content generation, intelligent chatbots, and sentiment analysis poses considerable challenges for LLM service providers.
The penn treebank: Annotating predicate argument structure
Marcus, M., Kim, G., Marcinkiewicz, M. A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B · 1994
Earlier work this paper cites.
Roofline: an insightful visual performance model for multicore architectures
Williams, S., Waterman, A., and Patterson, D · 2009
Earlier work this paper cites.
Standardized Assessment of Reading Performance: The New International Reading Speed Texts IReST
Trauzettel-Klosinski, S., Dietz, K., and the IReST Study Group · 2012
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding, 2016
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference, 2017
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions, 2019
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale, 2019
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?, 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Earlier work this paper cites.
A white paper on neural network quantization, 2021
Nagel, M., Fournarakis, M., Amjad, R. A., Bondarenko, Y., van Baalen, M., and Blankevoort, T · 2021
Earlier work this paper cites.
Demystifying the nvidia ampere architecture through microbenchmarking and instruction-level analysis, 2022
Abdelkhalik, H., Arafa, Y., Santhi, N., and Badawy, A.-H · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Llm.int8(): 8-bit matrix multiplication for transformers at scale, 2022
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces, 2022
Gu, A., Goel, K., and Ré, C · 2022
Earlier work this paper cites.
Fp8 formats for deep learning, 2022
Micikevicius, P., Stosic, D., Burgess, N., Cornea, M., Dubey, P., Grisenthwaite, R., Ha, S., Heinecke, A., Judd, P., Kamalu, J., Mellempudi, N., Oberman, S., Shoeybi, M., Siu, M., and Wu, H · 2022
Cited alongside, same era.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2022
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers, 2022
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Cited alongside, same era.
Orca: A distributed serving system for Transformer-Based generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Cited alongside, same era.
Splitwise: Efficient generative llm inference using phase splitting, 2023
Patel, P., Choukse, E., Zhang, C., Íñigo Goiri, Shah, A., Maleki, S., and Bianchini, R · 2023
Closest in time.
Microscaling data formats for deep learning, 2023
Rouhani, B. D., Zhao, R., More, A., Hall, M., Khodamoradi, A., Deng, S., Choudhary, D., Cornea, M., Dellinger, E., Denolf, K., Dusan, S., Elango, V., Golub, M., Heinecke, A., James-Roxby, P., Jani, D., Kolhe, G., Langhammer, M., Li, A., Melnick, L., Mesmakhosroshahi, M., Rodriguez, A., Schulte, M., Shafipour, R., Shao, L., Siu, M., Dubey, P., Micikevicius, P., Naumov, M., Verrilli, C., Wittig, R., Burger, D., and Chung, E · 2023
Closest in time.
Omniquant: Omnidirectionally calibrated quantization for large language models, 2023
Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., and Luo, P · 2023
Closest in time.
High-throughput generative inference of large language models with a single gpu
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Fu, D. Y., Xie, Z., Chen, B., Barrett, C. W., Gonzalez, J., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Closest in time.
CUTLASS, January 2023
Thakkar, V., Ramani, P., Cecka, C., Shivam, A., Lu, H., Yan, E., Kosaian, J., Hoemmen, M., Wu, H., Kerr, A., Nicely, M., Merrill, D., Blasig, D., Qiao, F., Majcher, P., Springer, P., Hohnerbach, M., Wang, J., and Gupta, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Quip: 2-bit quantization of large language models with guarantees, 2023
Chee, J., Cai, Y., Kuleshov, V., and Sa, C. D · 2023
Cited alongside, same era.
Dissecting batching effects in gpt inference, May 2023
Chen, L · 2023
Cited alongside, same era.
Punica: Multi-tenant lora serving, 2023
Chen, L., Ye, Z., Wu, Y., Zhuo, D., Ceze, L., and Krishnamurthy, A · 2023
Cited alongside, same era.
Number of chatgpt users, Jul 2023
Duarte, F · 2023
Cited alongside, same era.
Chatgpt costs 700,000 to run daily, openai may go bankrupt in 2024, Aug 2023
Elimian, G · 2023
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers, 2023
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
Gu, A. and Dao, T · 2023
Cited alongside, same era.
Closest in time.
Attention is all you need, 2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2023
Closest in time.
Outlier suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling, 2023
Wei, X., Zhang, Y., Li, Y., Zhang, X., Gong, R., Guo, J., and Liu, X · 2023
Closest in time.
List of qualcomm snapdragon systems on chips — Wikipedia, the free encyclopedia, 2023
Wikipedia contributors · 2023
Closest in time.
Zeroquant-fp: A leap forward in llms post-training w4a8 quantization using floating-point formats, 2023
Wu, X., Yao, Z., and He, Y · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models, 2023
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Closest in time.
Rptq: Reorder-based post-training quantization for large language models, 2023
Yuan, Z., Niu, L., Liu, J., Liu, W., Wang, X., Shang, Y., Sun, G., Wu, Q., Wu, J., and Wu, B · 2023
Closest in time.
H 2 o: Heavy-hitter oracle for efficient generative inference of large language models, 2023
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., Wang, Z., and Chen, B · 2023
Closest in time.
Sysmol: A hardware-software co-design framework for ultra-low and fine-grained mixed-precision neural networks, 2023
Zhou, C., Richard, V., Savarese, P., Hassman, Z., Maire, M., DiBrino, M., and Li, Y · 2023
Closest in time.
Taming throughput-latency tradeoff in llm inference with sarathi-serve, 2024
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., Tumanov, A., and Ramjee, R · 2024
Closest in time.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models, 2024
Dai, D., Deng, C., Zhao, C., Xu, R. X., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., Xie, Z., Li, Y. K., Huang, P., Luo, F., Ruan, C., Sui, Z., and Liang, W · 2024
Closest in time.
Mixtral of experts, 2024
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.
Accelerating self-attentions for llm serving with flashinfer, February 2024
Ye, Z., Chen, L., Lai, R., Zhao, Y., Zheng, S., Shao, J., Hou, B., Jin, H., Zuo, Y., Yin, L., Chen, T., and Ceze, L · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving, 2024
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Closest in time.