Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) from the GPT family have become extremely popular, leading to a race towards reducing their inference costs to allow for efficient local computation.
PiQA: An algebra for querying protein data sets
Tata, S. and Patel, J. M · 2003
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
A systematic classification of knowledge, reasoning, and context within the ARC dataset
Boratko, M., Padigela, H., Mikkilineni, D., Yuvraj, P., Das, R., McCallum, A., Chang, M., Fokoue-Nkoutche, A., Kapanipathi, P., Mattei, N., et al · 2018
Earlier work this paper cites.
Pact: Parameterized clipping activation for quantized neural networks
Choi, J., Wang, Z., Venkataramani, S., Chuang, P. I.-J., Srinivasan, V., and Gopalakrishnan, K · 2018
Earlier work this paper cites.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P · 2020
Earlier work this paper cites.
A framework for few-shot language model evaluation
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., et al · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
FlashAttention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Cited alongside, same era.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2022
Cited alongside, same era.
LLM.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Cited alongside, same era.
Squeezellm: Dense-and-sparse quantization
Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., and Keutzer, K · 2023
Closest in time.
Owq: Lessons learned from activation outliers for weight quantization in large language models
Lee, C., Jin, J., Kim, T., Kim, H., and Park, E · 2023
Closest in time.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Closest in time.
Fptq: Fine-grained post-training quantization for large language models
Li, Q., Zhang, Y., Li, L., Yao, P., Zhang, B., Chu, X., Sun, Y., Du, L., and Xie, Y · 2023
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Cited alongside, same era.
OPT: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Spqr: A sparse-quantized representation for near-lossless llm weight compression
Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D · 2023
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
Nvidia nsight compute
NVIDIA
Cited in the paper.
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Closest in time.
Nvidia cutlass library, 2023
NVIDIA · 2023
Closest in time.
Omniquant: Omnidirectionally calibrated quantization for large language models, 2023
Shao, W., Chen, M., Zhang, Z., Xu, P., Zhao, L., Li, Z., Zhang, K., Gao, P., Qiao, Y., and Luo, P · 2023
Closest in time.
The Falcon family of large language models
TII UAE · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Rptq: Reorder-based post-training quantization for large language models
Yuan, Z., Niu, L., Liu, J., Liu, W., Wang, X., Shang, Y., Sun, G., Wu, Q., Wu, J., and Wu, B · 2023
Closest in time.