Fetching the paper…
Reading the bibliography…
Large language models (LLMs) show excellent performance in difficult tasks, but they often require massive memories and computational resources.
Building a large annotated corpus of English: the penn treebank
Marcus, M · 1993
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Le bras, R., Gao, J., and Choi, Y · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Compressing pre-trained language models by matrix decomposition
Noach, M. and Goldberg, Y · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Le Bras, R., Bhagavatula, C., and Choi, Y · 2020
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A · 2020
Earlier work this paper cites.
Drone: Data-aware low-rank compression for large nlp models
Chen, P. H., Yu, H.-F., Dhillon, I. S., and Hsieh, C.-J · 2021
Earlier work this paper cites.
Gpt3.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Language model compression with weighted low-rank factorization
Hsu, Y.-C., Hua, T., Chang, S., Lou, Q., Shen, Y., and Jin, H · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W · 2022
Cited alongside, same era.
Numerical optimizations for weighted low-rank estimation on language model
Hua, T., Hsu, Y.-C., Wang, F., Lou, Q., Shen, Y., and Jin, H · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Cited alongside, same era.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Later among the works it cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Later among the works it cites.
Lord: Low rank decomposition of monolingual code llms for one-shot compression
Kaushal, A., Vaidhya, T., and Rish, I · 2023
Later among the works it cites.
Llm-pruner: On the structural pruning of large language models
Ma, X., Fang, G., and Wang, X · 2023
Later among the works it cites.
Low-rank prune-and-factorize for language model compression
Ren, S. and Zhu, K. Q · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Cited alongside, same era.
Structured pruning learns compact and accurate models
Xia, M., Zhong, Z., and Chen, D · 2022
Cited alongside, same era.
Gradient-based intra-attention pruning on pre-trained language models
Yang, Z., Cui, Y., Yao, X., and Wang, S · 2022
Cited alongside, same era.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Yazdani Aminabadi, R., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Cited alongside, same era.
Glm-130b: An open bilingual pre-trained model
Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., et al · 2022
Cited alongside, same era.
Lorashear: Efficient large language model structured pruning and knowledge recovery
Chen, T., Ding, T., Yadav, B., Zharkov, I., and Liang, L · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Can language models teach? teacher explanations improve student performance via personalization
Saha, S., Hase, P., and Bansal, M · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
Sun, M., Liu, Z., Bair, A., and Kolter, J. Z · 2023
Later among the works it cites.
A framework for few-shot language model evaluation
Sutawika, L., Gao, L., Schoelkopf, H., Biderman, S., Tow, J., Abbasi, B., ben fattori, Lovering, C., farzanehnakhaee70, Phang, J., Thite, A., Fazz, Aflah, Muennighoff, N., Wang, T., sdtblck, nopperl, gakada, tttyuntian, researcher2, Chris, Etxaniz, J., Kasner, Z., Khalid, Hsu, J., AndyZwei, Ammanamanchi, P. S., Groeneveld, D., Smith, E., and Tang, E · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
SCOTT: Self-consistent chain-of-thought distillation
Wang, P., Wang, Z., Li, Z., Gao, Y., Yin, B., and Ren, X · 2023
Later among the works it cites.
Junk dna hypothesis: A task-centric angle of llm pre-trained weights through sparsity
Yin, L., Liu, S., Jaiswal, A., Kundu, S., and Wang, Z · 2023
Later among the works it cites.
Compressing transformers: Features are low-rank, but weights are not!
Yu, H. and Wu, J · 2023
Later among the works it cites.
Pruning meets low-rank parameter-efficient fine-tuning
Zhang, M., Shen, C., Yang, Z., Ou, L., Yu, X., Zhuang, B., et al · 2023
Later among the works it cites.