Fetching the paper…
Reading the bibliography…
Compressing high-capability Large Language Models (LLMs) has emerged as a favored strategy for resource-efficient inferences.
Piqa: An algebra for querying protein data sets
Tata, S. and Patel, J. M · 2003
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
A systematic classification of knowledge, reasoning, and context within the arc dataset
Boratko, M., Padigela, H., Mikkilineni, D., Yuvraj, P., Das, R., McCallum, A., Chang, M., Fokoue-Nkoutche, A., Kapanipathi, P., Mattei, N., et al · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Dettmers, T. and Zettlemoyer, L · 2019
Earlier work this paper cites.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Earlier work this paper cites.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Mostafa, H. and Wang, X · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Multi-prize lottery ticket hypothesis: Finding accurate binary neural networks by pruning a randomly weighted network
Diffenderfer, J. and Kailkhura, B · 2020
Earlier work this paper cites.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Dynamic model pruning with feedback
Lin, T., Stich, S. U., Barba, L., Dmitriev, D., and Jaggi, M · 2020
Earlier work this paper cites.
Nvidia a100 tensor core gpu architecture
Nvidia · 2020
Cited alongside, same era.
Woodfisher: Efficient second-order approximation for neural network compression
Singh, S. P. and Alistarh, D · 2020
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Learning n: m fine-grained structured sparse neural networks from scratch
Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H · 2021
Cited alongside, same era.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Later among the works it cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Li, J., Cheng, X., Zhao, W. X., Nie, J., and rong Wen, J · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Later among the works it cites.
Llm-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Later among the works it cites.
Mo, L., Wang, B., Chen, M., and Sun, H · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal brain compression: A framework for accurate post-training quantization and pruning
Frantar, E. and Alistarh, D · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
Training your sparse neural network better with any mask
Jaiswal, A. K., Ma, H., Chen, T., Ding, Y., and Wang, Z · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Cited alongside, same era.
Intriguing properties of quantization at scale
Ahmadian, A., Dash, S., Chen, H., Venkitesh, B., Gou, S., Blunsom, P., Üstün, A., and Hooker, S · 2023
Cited alongside, same era.
Later among the works it cites.
The cost of compression: Investigating the impact of compression on parametric knowledge in language models
Namburi, S. S. S., Sreedhar, M., Srinivasan, S., and Sala, F · 2023
Later among the works it cites.
Latent jailbreak: A benchmark for evaluating text safety and output robustness of large language models
Qiu, H., Zhang, S., Li, A., He, H., and Lan, Z · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2023
Later among the works it cites.
Baby llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty
Timiryasov, I. and Tastet, J · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Later among the works it cites.
Tensorgpt: Efficient compression of the embedding layer in llms based on the tensor-train decomposition
Xu, M., Xu, Y. L., and Mandic, D. P · 2023
Later among the works it cites.
Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity
Yin, L., Wu, Y., Zhang, Z., Hsieh, C.-Y., Wang, Y., Jia, Y., Pechenizkiy, M., Liang, Y., Wang, Z., and Liu, S · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Gong, N. Z., Zhang, Y., et al · 2023
Later among the works it cites.
Pruning for protection: Increasing jailbreak resistance in aligned llms without fine-tuning
Hasan, A., Rugina, I., and Wang, A · 2024
Closest in time.
Trustllm: Trustworthiness in large language models
Sun, L., Huang, Y., Wang, H., Wu, S., Zhang, Q., Gao, C., Huang, Y., Lyu, W., Zhang, Y., Li, X., et al · 2024
Closest in time.
Quip#: with lattice codebooks, 2023
Tseng, A., Chee, J., Sun, Q., Kuleshov, V., and De Sa, C · 2024
Closest in time.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Wei, B., Huang, K., Huang, Y., Xie, T., Qi, X., Xia, M., Mittal, P., Wang, M., and Henderson, P · 2024
Closest in time.