Fetching the paper…
Reading the bibliography…
Large language models have high compute, latency, and memory requirements.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Earlier work this paper cites.
Sparsednn: Fast sparse deep learning inference on cpus
Wang, Z · 2021
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebron, F., and Sanghai, S · 2023
Earlier work this paper cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
LLM-FP4: 4-bit floating-point quantized transformers
Liu, S.-y., Liu, Z., Huang, X., Dong, P., and Cheng, K.-T · 2023
Cited alongside, same era.
Llm-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Cited alongside, same era.
Efficient llm inference on cpus
Shen, H., Chang, H., Dong, B., Luo, Y., and Meng, H · 2023
Cited alongside, same era.
Xia, H., Zheng, Z., Li, Y., Zhuang, D., Zhou, Z., Qiu, X., Li, Y., Lin, W., and Song, S. L · 2023
Cited alongside, same era.
llama.cpp, 2024
Gerganov, G · 2024
Later among the works it cites.
Intel oneapi deep neural network library, 2024
Intel · 2024
Later among the works it cites.
Exploiting intel advanced matrix extensions (amx) for large language model inference
Kim, H., Ye, G., Wang, N., Yazdanbakhsh, A., and Kim, N. S · 2024
Later among the works it cites.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., and Han, S · 2024
Later among the works it cites.
Shears: Unstructured sparsity with neural low-rank adapter search
Muñoz, J. P., Yuan, J., and Jain, N · 2024
Later among the works it cites.
SQFT: Low-cost model adaptation in low-precision sparse foundation models
Munoz, J. P., Yuan, J., and Jain, N · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H2o: heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Cited alongside, same era.
A survey on model compression for large language models
Zhu, X., Li, J., Liu, Y., Ma, C., and Wang, W · 2023
Cited alongside, same era.
Shadowllm: Predictor-based contextual sparsity for large language models
Akhauri, Y., AbouElhamayed, A. F., Dotzel, J., Zhang, Z., Rush, A. M., Huda, S., and Abdelfattah, M. S · 2024
Cited alongside, same era.
Everybody prune now: Structured pruning of llms with only forward passes
Dery, L., Kolawole, S., Kagey, J.-F., Smith, V., Neubig, G., and Talwalkar, A · 2024
Cited alongside, same era.
Learning from students: Applying t-distributions to explore accurate and efficient formats for llms
Dotzel, J., Chen, Y., Kotb, B., Prasad, S., Wu, G., Li, S., Abdelfattah, M. S., and Zhang, Z · 2024
Cited alongside, same era.
Marlin: Mixed-precision auto-regressive parallel inference on large language models
Frantar, E., Castro, R. L., Chen, J., Hoefler, T., and Alistarh, D · 2024
Cited alongside, same era.
Quick: Quantization-aware interleaving and conflict-free kernel for efficient llm inference, 2024b
Kim, T., Lee, J., Ahn, D., Kim, S., Choi, J., Kim, M., and Kim, H
Cited in the paper.
Neuralmagic/deepsparse: Sparsity-aware deep learning inference runtime for cpus, 2024
Neuralmagic · 2024
Later among the works it cites.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2024
Later among the works it cites.
Think: Thinner key cache by query-driven pruning
Xu, Y., Jie, Z., Dong, H., Wang, L., Lu, X., Zhou, A., Saha, A., Xiong, C., and Sahoo, D · 2024
Later among the works it cites.
Subgen: Token generation in sublinear time and memory
Zandieh, A., Han, I., Mirrokni, V., and Karbasi, A · 2024
Later among the works it cites.