Fetching the paper…
Reading the bibliography…
Quantization can accelerate large language model (LLM) inference.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 1905
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library, 2019
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Earlier work this paper cites.
Aˆ 3: Accelerating attention mechanisms in neural networks with approximation
Ham, T. J., Jung, S. J., Kim, S., Oh, Y. H., Park, Y., Song, Y., Park, J.-H., Lee, S., Park, K., Lee, J. W., et al · 2020
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing, 2020
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Earlier work this paper cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Zadeh, A. H., Edo, I., Awad, O. M., and Moshovos, A · 2020
Earlier work this paper cites.
Elsa: Hardware-software co-design for efficient, lightweight self-attention mechanism in neural networks
Ham, T. J., Lee, Y., Seo, S. H., Kim, S., Choi, H., Jung, S. J., and Lee, J. W · 2021
Earlier work this paper cites.
Edgebert: Sentence-level energy optimizations for latency-aware multi-task nlp inference
Tambe, T., Hooper, C., Pentecost, L., Jia, T., Yang, E.-Y., Donato, M., Sanh, V., Whatmough, P., Rush, A. M., Brooks, D., et al · 2021
Earlier work this paper cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Wang, H., Zhang, Z., and Han, S · 2021
Earlier work this paper cites.
GPT3.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
An algorithm–hardware co-optimized framework for accelerating n: M sparse transformers
Fang, C., Zhou, A., and Wang, Z · 2022
Cited alongside, same era.
GPTQ: Accurate post-training compression for generative pretrained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation
Hong, S., Moon, S., Kim, J., Lee, S., Kim, M., Lee, D., and Kim, J.-Y · 2022
Cited alongside, same era.
Who says elephants can’t run: Bringing large scale moe models into cloud scale production
Kim, Y. J., Henry, R., Fahim, R., and Awadalla, H. H · 2022
Cited alongside, same era.
Dota: detect and omit weak attentions for scalable transformer acceleration
Qu, Z., Liu, L., Tu, F., Chen, Z., Ding, Y., and Xie, Y · 2022
Cited alongside, same era.
Omniquant: Omnidirectionally calibrated quantization for large language models
Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K. Z., Gao, P., Qiao, Y., and Luo, P · 2023
Later among the works it cites.
MLC-LLM, 2023
team, M · 2023
Later among the works it cites.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Later among the works it cites.
Dual grained quantization: Efficient fine-grained quantization for llm
Zhang, L., Fei, W., Wu, W., He, Y., Lou, Z., and Zhou, H · 2023
Later among the works it cites.
Atom: Low-bit quantization for efficient and accurate llm serving
Zhao, Y., Lin, C.-Y., Zhu, K., Ye, Z., Chen, L., Zheng, S., Ceze, L., Krishnamurthy, A., Chen, T., and Kasikci, B · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Outlier suppression: Pushing the limit of low-bit transformer language models
Wei, X., Zhang, Y., Zhang, X., Gong, R., Zhang, S., Zhang, Q., Yu, F., and Liu, X · 2022
Cited alongside, same era.
Orca: A distributed serving system for Transformer-Based generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
TensorRT-LLM: A TensorRT Toolbox for Optimized Large Language Model Inference, 2023
NVIDIA · 2023
Cited alongside, same era.
Microscaling data formats for deep learning
Rouhani, B. D., Zhao, R., More, A., Hall, M., Khodamoradi, A., Deng, S., Choudhary, D., Cornea, M., Dellinger, E., Denolf, K., et al · 2023
Cited alongside, same era.
Efficiently programming large language models using sglang, 2023
Zheng, L., Yin, L., Xie, Z., Huang, J., Sun, C., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., Barrett, C., and Sheng, Y · 2023
Later among the works it cites.
Quarot: Outlier-free 4-bit inference in rotated llms
Ashkboos, S., Mohtashami, A., Croci, M. L., Li, B., Jaggi, M., Alistarh, D., Hoefler, T., and Hensman, J · 2024
Closest in time.
Quip: 2-bit quantization of large language models with guarantees, 2024
Chee, J., Cai, Y., Kuleshov, V., and Sa, C. D · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Squeezellm: Dense-and-sparse quantization, 2024
Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., and Keutzer, K · 2024
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S · 2024
Closest in time.
Yi: Open foundation models by 01.ai, 2024
Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, S., Yang, S., Yu, T., Xie, W., Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu, Z., and Dai, Z · 2024
Closest in time.