Fetching the paper…
Reading the bibliography…
As inference on Large Language Models (LLMs) emerges as an important workload in machine learning applications, weight quantization has become a standard technique for efficient GPU deployment.
Optimizing parallel reduction in cuda
Harris, M. et al · 2007
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
NVIDIA A100 tensor core GPU architecture
NVIDIA · 2020
Earlier work this paper cites.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2022
Earlier work this paper cites.
LLM.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Who says elephants can’t run: Bringing large scale moe models into cloud scale production
Kim, Y. J., Henry, R., Fahim, R., and Awadalla, H. H · 2022
Earlier work this paper cites.
Dissecting tensor cores via microbenchmarks: Latency, throughput and numeric behaviors
Sun, W., Li, A., Geng, T., Stuijk, S., and Corporaal, H · 2022
Earlier work this paper cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Cited alongside, same era.
OPT: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Quip: 2-bit quantization of large language models with guarantees, 2023
Chee, J., Cai, Y., Kuleshov, V., and Sa, C. D · 2023
Cited alongside, same era.
Spqr: A sparse-quantized representation for near-lossless llm weight compression
Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D · 2023
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Flexgen: High-throughput generative inference of large language models with a single gpu
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Later among the works it cites.
The Falcon family of large language models
TII UAE · 2023
Later among the works it cites.
Quarot: Outlier-free 4-bit inference in rotated llms
Ashkboos, S., Mohtashami, A., Croci, M. L., Li, B., Jaggi, M., Alistarh, D., Hoefler, T., and Hensman, J · 2024
Closest in time.
Faster and lighter llms: A survey on current challenges and way forward
Chavan, A., Magazine, R., Kushwaha, S., Debbah, M., and Gupta, D · 2024
Closest in time.
Extreme compression of large language models via additive quantization
Egiazarian, V., Panferov, A., Kuznedelev, D., Frantar, E., Babenko, A., and Alistarh, D · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Cited alongside, same era.
Stream-k: Work-centric parallel decomposition for dense matrix-matrix multiplication on the GPU
Osama, M., Merrill, D., Cecka, C., Garland, M., and Owens, J. D · 2023
Cited alongside, same era.
Omniquant: Omnidirectionally calibrated quantization for large language models, 2023
Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., and Luo, P · 2023
Cited alongside, same era.
Nvidia a10 datasheet
NVIDIA
Cited in the paper.
Nvidia instruction set
NVIDIA
Cited in the paper.
Efficient GEMM in CUDA
NVIDIA
Cited in the paper.
Closest in time.
Exllamav2: A memory efficient fork of hf transformers optimized for llama models
ExLlamaV2 · 2024
Closest in time.
Chatgpt creators openai are generating 100 billion words per day, ceo says, 2024
Griffin, A · 2024
Closest in time.
vllm project pull request #6612: Add support for awq marlin
vLLM Project Contributors · 2024
Closest in time.
Qqq: Quality quattuor-bit quantization for large language models, 2024
Zhang, Y., Zhang, P., Huang, M., Xiang, J., Wang, Y., Wang, C., Zhang, Y., Yu, L., Liu, C., and Lin, W · 2024
Closest in time.