Fetching the paper…
Reading the bibliography…
Low Rank Decomposition of matrix - splitting a large matrix into a product of two smaller matrix offers a means for compression that reduces the parameters of a model without sparsification, and hence delivering more speedup on modern hardware.
An updated set of basic linear algebra subprograms (blas)
L Susan Blackford, Antoine Petitet, Roldan Pozo, Karin Remington, R Clint Whaley, James Demmel, Jack Dongarra, Iain Duff, Sven Hammarling, Greg Henry, et al · 2002
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library , chapter ., pp.
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2019
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need, 2019
Noam Shazeer · 2019
Earlier work this paper cites.
Compressing pre-trained language models by matrix decomposition
Matan Ben Noach and Yoav Goldberg · 2020
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Earlier work this paper cites.
Train large, then compress: Rethinking model size for efficient training and inference of transformers
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E. Gonzalez · 2020
Earlier work this paper cites.
Glu variants improve transformer, 2020
Noam Shazeer · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Earlier work this paper cites.
Drone: Data-aware low-rank compression for large nlp models
Patrick Chen, Hsiang-Fu Yu, Inderjit Dhillon, and Cho-Jui Hsieh · 2021
Earlier work this paper cites.
Pruning and quantization for deep neural network acceleration: A survey
Tailin Liang, John Glossner, Lei Wang, Shaobo Shi, and Xiaotong Zhang · 2021
Earlier work this paper cites.
Self-attention does not need o ( n 2 ) o(n^{2}) memory, 2021
Markus N. Rabe and Charles Staats · 2021
Earlier work this paper cites.
The stack smol, 2022
Project Bigcode · 2022
Earlier work this paper cites.
Creating sparse gpt-3 models with iterative pruning, 11 2022
Team Cerebras · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with IO-awareness
Tri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher Re · 2022
Earlier work this paper cites.
The case for 4-bit precision: k-bit inference scaling laws, 2022
Tim Dettmers and Luke Zettlemoyer · 2022
Cited alongside, same era.
GPT3.int8(): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Kronecker decomposition for GPT compression
Ali Edalati, Marzieh Tahaei, Ahmad Rashid, Vahid Nia, James Clark, and Mehdi Rezagholizadeh · 2022
Cited alongside, same era.
Rank diminishing in deep neural networks
Ruili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao, Michael Jordan, and Zheng-Jun Zha · 2022
Cited alongside, same era.
Language model compression with weighted low-rank factorization
Yen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou, Yilin Shen, and Hongxia Jin · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Impossible distillation: from low-quality model to high-quality dataset & model for summarization and paraphrasing, 2023
Jaehun Jung, Peter West, Liwei Jiang, Faeze Brahman, Ximing Lu, Jillian Fisher, Taylor Sorensen, and Yejin Choi · 2023
Closest in time.
Owq: Lessons learned from activation outliers for weight quantization in large language models, 2023
Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, and Eunhyeok Park · 2023
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration, 2023
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han · 2023
Closest in time.
Llm-pruner: On the structural pruning of large language models, 2023
Xinyin Ma, Gongfan Fang, and Xinchao Wang · 2023
Closest in time.
Compute unified device architecture (cuda)
Corporation NVIDIA · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Numerical optimizations for weighted low-rank estimation on language models
Ting Hua, Yen-Chang Hsu, Felicity Wang, Qian Lou, Yilin Shen, and Hongxia Jin · 2022
Cited alongside, same era.
The stack: 3 tb of permissively licensed source code, 2022
Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries · 2022
Cited alongside, same era.
Lut-gemm: Quantized matrix multiplication based on luts for efficient inference in large-scale generative language models, 2022
Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, and Dongsoo Lee · 2022
Cited alongside, same era.
KroneckerBERT: Significant compression of pre-trained language models through kronecker decomposition and knowledge distillation
Marzieh Tahaei, Ella Charlaix, Vahid Nia, Ali Ghodsi, and Mehdi Rezagholizadeh · 2022
Cited alongside, same era.
Outlier suppression: Pushing the limit of low-bit transformer language models
Xiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong, Shanghang Zhang, Qi Zhang, Fengwei Yu, and Xianglong Liu · 2022
Cited alongside, same era.
Gkd: Generalized knowledge distillation for auto-regressive sequence models, 2023
Rishabh Agarwal, Nino Vieillard, Piotr Stanczyk, Sabela Ramos, Matthieu Geist, and Olivier Bachem · 2023
Cited alongside, same era.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Closest in time.
The impact of ai on developer productivity: Evidence from github copilot, 2023
Sida Peng, Eirini Kalliamvakou, Peter Cihon, and Mert Demirer · 2023
Closest in time.
Low-rank prune-and-factorize for language model compression, 2023
Siyu Ren and Kenny Q. Zhu · 2023
Closest in time.
Code llama: Open foundation models for code, 2023
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve · 2023
Closest in time.
Omniquant: Omnidirectionally calibrated quantization for large language models, 2023
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo · 2023
Closest in time.
Pangu-coder2: Boosting large language models for code with ranking feedback, 2023
Bo Shen, Jiaxin Zhang, Taihong Chen, Daoguang Zan, Bing Geng, An Fu, Muhan Zeng, Ailun Yu, Jichuan Ji, Jingyang Zhao, Yuenan Guo, and Qianxiang Wang · 2023
Closest in time.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Zeroquant-fp: A leap forward in llms post-training w4a8 quantization using floating-point formats, 2023
Xiaoxia Wu, Zhewei Yao, and Yuxiong He · 2023
Closest in time.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Closest in time.
Zeroquant-v2: Exploring post-training quantization in llms from comprehensive study to low rank compensation, 2023
Zhewei Yao, Xiaoxia Wu, Cheng Li, Stephen Youn, and Yuxiong He · 2023
Closest in time.
Compressing transformers: Features are low-rank, but weights are not!
Hao Yu and Jianxin Wu · 2023
Closest in time.
Rptq: Reorder-based post-training quantization for large language models, 2023
Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu · 2023
Closest in time.
Lora-fa: Memory-efficient low-rank adaptation for large language models fine-tuning, 2023
Longteng Zhang, Lin Zhang, Shaohuai Shi, Xiaowen Chu, and Bo Li · 2023
Closest in time.
A survey on model compression for large language models, 2023
Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang · 2023
Closest in time.