Fetching the paper…
Reading the bibliography…
Recently, considerable efforts have been directed towards compressing Large Language Models (LLMs), which showcase groundbreaking capabilities across diverse applications but entail significant deployment costs due to their large sizes.
The Penn Treebank: annotating predicate argument structure
Marcus, M., Kim, G., Marcinkiewicz, M. A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B · 1994
Earlier work this paper cites.
PiQA: An algebra for querying protein data sets
Tata, S. and Patel, J. M · 2003
Earlier work this paper cites.
Hacker’s Delight
Warren, H. S · 2012
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Mcdnn: An approximation-based execution framework for deep stream processing under resource constraints
Han, S., Shen, H., Philipose, M., Agarwal, S., Wolman, A., and Krishnamurthy, A · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Work-in-progress: towards efficient quantized neural network inference on mobile devices
Umuroglu, Y. and Jahre, M · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the AI2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Automatic generation of high-performance quantized machine learning kernels
Cowan, M., Moreau, T., Chen, T., Bornholt, J., and Ceze, L · 2020
Earlier work this paper cites.
ALERT: Accurate learning for energy and timeliness
Wan, C., Santriaji, M., Rogers, E., Hoffmann, H., Maire, M., and Lu, S · 2020
Earlier work this paper cites.
Model-Switching: Dealing with fluctuating workloads in Machine-Learning-as-a-Service systems
Zhang, J., Elnikety, S., Zarar, S., Gupta, A., and Garg, S · 2020
Earlier work this paper cites.
A survey of quantization methods for efficient neural network inference, 2021
Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Earlier work this paper cites.
WinoGrande: an adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Any-precision deep neural networks
Yu, H., Li, H., Shi, H., Huang, T. S., and Hua, G · 2021
Cited alongside, same era.
OPT: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Cited alongside, same era.
Decoupled knowledge distillation
Zhao, B., Cui, Q., Song, R., Qiu, Y., and Liang, J · 2022
Cited alongside, same era.
QuIP: 2-bit quantization of large language models with guarantees
Chee, J., Cai, Y., Kuleshov, V., and Sa, C. D · 2023
Cited alongside, same era.
Accelerating large language model decoding with speculative sampling, 2023
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
AWQ: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Later among the works it cites.
LLM-QAT: Data-free quantization aware training for large language models, 2023
Liu, Z., Oguz, B., Zhao, C., Chang, E., Stock, P., Mehdad, Y., Shi, Y., Krishnamoorthi, R., and Chandra, V · 2023
Later among the works it cites.
LLM-Pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Later among the works it cites.
SpecInfer: Accelerating generative large language model serving with speculative inference and token tree verification, 2023
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., Shi, C., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2023
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
QLoRA: Efficient finetuning of quantized LLMs
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Cited alongside, same era.
SparseGPT: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Cited alongside, same era.
OPTQ: Accurate quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Cited alongside, same era.
Distilling Step-by-Step! outperforming larger language models with less training data and smaller model sizes
Hsieh, C., Li, C., Yeh, C., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C., and Pfister, T · 2023
Cited alongside, same era.
Mistral 7B, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Cited alongside, same era.
Later among the works it cites.
What matters in the structured pruning of generative language models?, 2023
Santacroce, M., Wen, Z., Shen, Y., and Li, Y · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
LoRAPrune: Pruning meets low-rank parameter-efficient fine-tuning, 2023
Zhang, M., Chen, H., Shen, C., Yang, Z., Ou, L., Yu, X., and Zhuang, B · 2023
Later among the works it cites.
Generalized knowledge distillation for auto-regressive language models
Agarwal, R., Vieillard, N., Zhou, Y., Stanczyk, P., Ramos, S., Geist, M., and Bachem, O · 2024
Closest in time.
SpQR: A sparse-quantized representation for near-lossless llm weight compression
Dettmers, T., Svirschevski, R., Egiazarian, V., Kuznedelev, D., Frantar, E., Ashkboos, S., Borzunov, A., Hoefler, T., and Alistarh, D · 2024
Closest in time.
MiniLLM: Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2024
Closest in time.
Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models
Lee, C., Jin, J., Kim, T., Kim, H., and Park, E · 2024
Closest in time.
LUT-GEMM: Quantized matrix multiplication based on LUTs for efficient inference in large-scale generative language models
Park, G., Park, B., Kim, M., Lee, S., Kim, J., Kwon, B., Kwon, S. J., Kim, B., Lee, Y., and Lee, D · 2024
Closest in time.