Fetching the paper…
Reading the bibliography…
How can we compress language models without sacrificing accuracy? The number of compression algorithms for language models is rapidly growing to benefit from remarkable advances of recent language models without side effects due to the gigantic size of language models, such as increased carbon emissions and expensive maintenance fees.
NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture Search. In KDD 2021 , Feida Zhu, Beng Chin Ooi, and Chunyan Miao (Eds.). ACM, 1933–1943
Jin Xu, Xu Tan, Renqian Luo, Kaitao Song, Jian Li, Tao Qin, and Tie-Yan Liu. 2021 · 1943
Earlier work this paper cites.
Optimal Brain Damage. In NeurIPS , David S. Touretzky (Ed.). Morgan Kaufmann
Yann LeCun, John S. Denker, and Sara A. Solla. 1989 · 1989
Earlier work this paper cites.
Second Order Derivatives for Network Pruning: Optimal Brain Surgeon. In NeurIPS 1992] , Stephen Jose Hanson, Jack D. Cowan, and C. Lee Giles (Eds.). Morgan Kaufmann, 164–171
Babak Hassibi and David G. Stork. 1992 · 1992
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases. In IWP 2005 . AFNLP
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville. 2013 · 2013
Earlier work this paper cites.
Learning both Weights and Connections for Efficient Neural Networks
Song Han, Jeff Pool, John Tran, and William J. Dally. 2015 · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In ICML , Francis R. Bach and David M. Blei (Eds.)
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Layer Normalization
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Long Short-Term Memory-Networks for Machine Reading. In EMNLP 2016 , Jian Su, Xavier Carreras, and Kevin Duh (Eds.). ACL
Jianpeng Cheng, Li Dong, and Mirella Lapata. 2016 · 2016
Earlier work this paper cites.
Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear Units
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
A Decomposable Attention Model for Natural Language Inference. In EMNLP 2016 , Jian Su, Xavier Carreras, and Kevin Duh (Eds.). ACL
Ankur P. Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text. In EMNLP 2016 , Jian Su, Xavier Carreras, and Kevin Duh (Eds.). ACL
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
Daniel M. Cer, Mona T. Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In ACL 2017 , Regina Barzilay and Min-Yen Kan (Eds.). 1601–1611
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Learning Sparse Neural Networks through L 0 {}_{\mbox{0}} Regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In NeurIPS 2017 , Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.)
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Deep Learning using Rectified Linear Units (ReLU)
Abien Fred Agarap. 2018 · 2018
Earlier work this paper cites.
Know What You Don’t Know: Unanswerable Questions for SQuAD. In ACL 2018 , Iryna Gurevych and Yusuke Miyao (Eds.). 784–789
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In NAACL-HLT 2018 , Marilyn A. Walker, Heng Ji, and Amanda Stent (Eds.). ACL, 1112–1122
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Earlier work this paper cites.
Efficient 8-Bit Quantization of Transformer Neural Machine Language Translation Model
Aishwarya Bhandare, Vamsi Sripathi, Deepthi Karkada, Vivek Menon, Sun Choi, Kushal Datta, and Vikram A. Saletore. 2019 · 2019
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Recurrent Stacking of Layers for Compact Neural Machine Translation Models. In AAAI 2019 . 6292–6299
Raj Dabre and Atsushi Fujita. 2019 · 2019
Earlier work this paper cites.
Universal Transformers. In ICLR 2019
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT 2019 , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). ACL, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks. In ICLR
Jonathan Frankle and Michael Carbin. 2019 · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP. In ICML 2019 . PMLR, 2790–2799
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Natural Questions: a Benchmark for Question Answering Research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Are Sixteen Heads Really Better than One?. In NeurIPS 2019 , Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds.). 14014–14024
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Earlier work this paper cites.
Fast Transformer Decoding: One Write-Head is All You Need
Noam Shazeer. 2019 · 2019
Earlier work this paper cites.
The Evolved Transformer. In ICML 2019 (PMLR, Vol. 97) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). 5877–5886
David R. So, Quoc V. Le, and Chen Liang. 2019 · 2019
Earlier work this paper cites.
Patient Knowledge Distillation for BERT Model Compression. In IJCNLP 2019 , Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). ACL, 4322–4331
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In ICLR 2019
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Tied Transformers: Neural Machine Translation with Shared Encoder and Decoder. In AAAI 2019 . 5466–5473
Yingce Xia, Tianyu He, Xu Tan, Fei Tian, Di He, and Tao Qin. 2019 · 2019
Earlier work this paper cites.
Q8BERT: Quantized 8Bit BERT. In NeurIPS 2019 . 36–39
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 2019
Earlier work this paper cites.
Knowledge Distillation from Internal Representations. In AAAI 2020 . 7350–7357
Gustavo Aguilar, Yuan Ling, Yu Zhang, Benjamin Yao, Xing Fan, and Chenlei Guo. 2020 · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners. In NeurIPS 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search. In IJCAI 2020 , Christian Bessiere (Ed.). 2463–2469
Daoyuan Chen, Yaliang Li, Minghui Qiu, Zhen Wang, Bofang Li, Bolin Ding, Hongbo Deng, Jun Huang, Wei Lin, and Jingren Zhou. 2020b · 2020
Earlier work this paper cites.
A comprehensive survey on model compression and acceleration
Tejalal Choudhary, Vipul Kumar Mishra, Anurag Goswami, and Jagannathan Sarangapani. 2020 · 2020
Earlier work this paper cites.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. In ICLR 2020
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Earlier work this paper cites.
Multi-Head Attention: Collaborate Instead of Concatenate
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2020 · 2020
Earlier work this paper cites.
Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing. In NeurIPS 2020
Zihang Dai, Guokun Lai, Yiming Yang, and Quoc Le. 2020 · 2020
Earlier work this paper cites.
Analyzing Redundancy in Pretrained Transformer Models. In EMNLP , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.)
Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov. 2020 · 2020
Earlier work this paper cites.
Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive Survey
Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie. 2020 · 2020
Earlier work this paper cites.
PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector Elimination. In ICML 2020 (PMLR, Vol. 119) . 3690–3699
Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy, Yogish Sabharwal, and Ashish Verma. 2020 · 2020
Earlier work this paper cites.
DynaBERT: Dynamic BERT with Adaptive Width and Depth. In NeurIPS 2020 , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. 2020 · 2020
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding. In EMNLP 2020 (ACL) , Trevor Cohn, Yulan He, and Yang Liu (Eds.). 4163–4174
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Earlier work this paper cites.
Reformer: The Efficient Transformer. In ICLR 2020
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Earlier work this paper cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In ICLR 2020
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Earlier work this paper cites.
AUBER: Automated BERT Regularization
Hyun Dong Lee, Seongmin Lee, and U Kang. 2020 · 2020
Earlier work this paper cites.
BERT-EMD: Many-to-Many Layer Mapping for BERT Compression with Earth Mover’s Distance. In EMNLP 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). ACL, 3009–3018
Jianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu, Min Yang, and Yaohong Jin. 2020 · 2020
Earlier work this paper cites.
Pruning Redundant Mappings in Transformer Models via Spectral-Normalized Identity Prior. In EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). ACL, 719–730
Zi Lin, Jeremiah Z. Liu, Zi Yang, Nan Hua, and Dan Roth. 2020 · 2020
Earlier work this paper cites.
Compressing Pre-trained Language Models by Matrix Decomposition. In IJCNLP 2020 , Kam-Fai Wong, Kevin Knight, and Hua Wu (Eds.). ACL, 884–889
Matan Ben Noach and Yoav Goldberg. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
Movement Pruning: Adaptive Sparsity by Fine-Tuning. In NeurIPS 2020 , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Victor Sanh, Thomas Wolf, and Alexander M. Rush. 2020 · 2020
Earlier work this paper cites.
Winning the Lottery with Continuous Sparsification. In NeurIPS , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Pedro Savarese, Hugo Silva, and Michael Maire. 2020 · 2020
Cited alongside, same era.
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT. In AAAI 2020 . 8815–8821
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer. 2020 · 2020
Cited alongside, same era.
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices. In ACL 2020 , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault (Eds.). 2158–2170
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Cited alongside, same era.
HAT: Hardware-Aware Transformers for Efficient Natural Language Processing. In ACL 2020 . 7675–7688
Hanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai, Ligeng Zhu, Chuang Gan, and Song Han. 2020d · 2020
Cited alongside, same era.
Linformer: Self-Attention with Linear Complexity
BiBERT: Accurate Fully Binarized BERT. In ICLR 2022
Haotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan, Aishan Liu, Qingqing Dang, Ziwei Liu, and Xianglong Liu. 2022 · 2022
Later among the works it cites.
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, and et al. 2022 · 2022
Later among the works it cites.
MKQ-BERT: Quantized BERT with 4-bits Weights and Activations
Hanlin Tang, Xipeng Zhang, Kai Liu, Jianchen Zhu, and Zhanhui Kang. 2022 · 2022
Later among the works it cites.
Compression of Generative Pre-trained Language Models via Quantization. In ACL 2022 , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). 4821–4836
Chaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, and Ngai Wong. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. 2020a · 2020
Cited alongside, same era.
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. In NeurIPS 2020 , Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.)
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020b · 2020
Cited alongside, same era.
Structured Pruning of Large Language Models. In EMNLP 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). ACL, 6151–6162
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2020c · 2020
Cited alongside, same era.
Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
Hao Wu, Patrick Judd, Xiaojie Zhang, Mikhail Isaev, and Paulius Micikevicius. 2020a · 2020
Cited alongside, same era.
Why Skip If You Can Combine: A Simple Knowledge Distillation Technique for Intermediate Layers. In EMNLP 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). ACL, 1016–1021
Yimeng Wu, Peyman Passban, Mehdi Rezagholizadeh, and Qun Liu. 2020b · 2020
Cited alongside, same era.
GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference. In MICRO 2020 . IEEE, 811–824
Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. 2020 · 2020
Cited alongside, same era.
TernaryBERT: Distillation-aware Ultra-low Bit BERT. In EMNLP 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). ACL, 509–521
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
An Investigation on Different Underlying Quantization Schemes for Pre-trained Language Models. In NLPCC 2020 (LNCS, Vol. 12430) , Xiaodan Zhu, Min Zhang, Yu Hong, and Ruifang He (Eds.). Springer, 359–371
Zihan Zhao, Yuncong Liu, Lu Chen, Qi Liu, Rao Ma, and Kai Yu. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Exploring extreme parameter compression for pre-trained language models. In ICLR 2022
Benyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang, and Qun Liu. 2022 · 2022
Later among the works it cites.
Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models. In NeurIPS
Xiuying Wei, Yunchen Zhang, Xiangguo Zhang, Ruihao Gong, Shanghang Zhang, Qi Zhang, Fengwei Yu, and Xianglong Liu. 2022 · 2022
Later among the works it cites.
Structured Pruning Learns Compact and Accurate Models. In ACL 2022 , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). 1513–1528
Mengzhou Xia, Zexuan Zhong, and Danqi Chen. 2022 · 2022
Later among the works it cites.
From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model Compression. In Thirty-Sixth AAAI Conference on Artificial Intelligence, AAAI 2022, Thirty-Fourth Conference on Innovative Applications of Artificial Intelligence, IAAI 2022, The Twelveth Symposium on Educational Advances in Artificial Intelligence, EAAI 2022 Virtual Event, February 22 - March 1, 2022 . AAAI Press, 11547–11555
Runxin Xu, Fuli Luo, Chengyu Wang, Baobao Chang, Jun Huang, Songfang Huang, and Fei Huang. [n. d.] · 2022
Later among the works it cites.
LEAP: Learnable Pruning for Transformer-based Models
Zhewei Yao, Xiaoxia Wu, Linjian Ma, Sheng Shen, Kurt Keutzer, Michael W. Mahoney, and Yuxiong He. 2022b · 2022
Later among the works it cites.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models. In ACL 2022 . ACL
Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. 2022 · 2022
Later among the works it cites.
OPT: Open Pre-trained Transformer Language Models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023 · 2023
Later among the works it cites.
QuantEase: Optimization-based Quantization for Language Models - An Efficient and Intuitive Algorithm
Kayhan Behdin, Ayan Acharya, Aman Gupta, Sathiya Keerthi Selvaraj, and Rahul Mazumder. 2023 · 2023
Later among the works it cites.
Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing
Yelysei Bondarenko, Markus Nagel, and Tijmen Blankevoort. 2023 · 2023
Later among the works it cites.
INT2.1: Towards Fine-Tunable Quantized Large Language Models with Error Correction through Low-Rank Adaptation
Yuji Chai, John Gkountouras, Glenn G. Ko, David Brooks, and Gu-Yeon Wei. 2023 · 2023
Later among the works it cites.
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher De Sa. 2023 · 2023
Later among the works it cites.
Parameter-Efficient Fine-Tuning Design Spaces. In ICLR
Jiaao Chen, Aston Zhang, Xingjian Shi, Mu Li, Alex Smola, and Diyi Yang. 2023 · 2023
Later among the works it cites.
QLoRA: Efficient Finetuning of Quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023a · 2023
Later among the works it cites.
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. 2023b · 2023
Later among the works it cites.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias Frantar and Dan Alistarh. 2023 · 2023
Later among the works it cites.
OPTQ: Accurate Quantization for Generative Pre-trained Transformers. In ICLR 2023
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2023 · 2023
Later among the works it cites.
PreQuant: A Task-agnostic Quantization Approach for Pre-trained Language Models. In ACL 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). 8065–8079
Zhuocheng Gong, Jiahao Liu, Qifan Wang, Yang Yang, Jingang Wang, Wei Wu, Yunsen Xian, Dongyan Zhao, and Rui Yan. 2023 · 2023
Later among the works it cites.
Knowledge Distillation of Large Language Models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2023 · 2023
Later among the works it cites.
How To Train Your (Compressed) Large Language Model
Ananya Harsh Jha, Dirk Groeneveld, Emma Strubell, and Iz Beltagy. 2023 · 2023
Later among the works it cites.
Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer Quantization
Jeonghoon Kim, Jung Hyun Lee, Sungdong Kim, Joonsuk Park, Kang Min Yoo, Se Jung Kwon, and Dongsoo Lee. 2023d · 2023
Later among the works it cites.
SqueezeLLM: Dense-and-Sparse Quantization
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, and Kurt Keutzer. 2023b · 2023
Later among the works it cites.
Full Stack Optimization of Transformer Inference: a Survey
Sehoon Kim, Coleman Hooper, Thanakul Wattanawong, Minwoo Kang, Ruohan Yan, Hasan Genc, Grace Dinh, Qijing Huang, Kurt Keutzer, Michael W. Mahoney, Yakun Sophia Shao, and Amir Gholami. 2023c · 2023
Later among the works it cites.
FineQuant: Unlocking Efficiency with Fine-Grained Weight-Only Quantization for LLMs
Young Jin Kim, Rawn Henry, Raffy Fahim, and Hany Hassan Awadalla. 2023a · 2023
Later among the works it cites.
ZipLM: Hardware-Aware Structured Pruning of Language Models
Eldar Kurtic, Elias Frantar, and Dan Alistarh. 2023 · 2023
Later among the works it cites.
OWQ: Lessons learned from activation outliers for weight quantization in large language models
Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, and Eunhyeok Park. 2023 · 2023
Later among the works it cites.
Norm Tweaking: High-performance Low-bit Quantization of Large Language Models
Liang Li, Qingyuan Li, Bo Zhang, and Xiangxiang Chu. 2023a · 2023
Later among the works it cites.
LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models
Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, and Tuo Zhao. 2023b · 2023
Later among the works it cites.
LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models
Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, and Tuo Zhao. 2023c · 2023
Later among the works it cites.
Less is More: Task-aware Layer-wise Distillation for Language Model Compression. In ICML 2023 (PMLR, Vol. 202) . 20852–20867
Chen Liang, Simiao Zuo, Qingru Zhang, Pengcheng He, Weizhu Chen, and Tuo Zhao. 2023 · 2023
Later among the works it cites.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023a · 2023
Later among the works it cites.
Understanding Parameter Sharing in Transformers
Ye Lin, Mingxuan Wang, Zhexi Zhang, Xiaohui Wang, Tong Xiao, and Jingbo Zhu. 2023b · 2023
Later among the works it cites.
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. 2023 · 2023
Later among the works it cites.
LightFormer: Light-weight Transformer Using SVD-based Weight Transfer and Parameter Sharing. In ACL 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). 10323–10335
Xiuqing Lv, Peng Zhang, Sunzhu Li, Guobing Gan, and Yueheng Sun. 2023 · 2023
Later among the works it cites.
LLM-Pruner: On the Structural Pruning of Large Language Models
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023 · 2023
Later among the works it cites.
Gradient-Free Structured Pruning with Unlabeled Data. In ICML (PMLR, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). 26326–26341
Azade Nova, Hanjun Dai, and Dale Schuurmans. 2023 · 2023
Later among the works it cites.
On the effect of dropping layers of pre-trained transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2023 · 2023
Later among the works it cites.
A Simple and Effective Pruning Approach for Large Language Models
Mingjie Sun, Zhuang Liu, Anna Bair, and J. Zico Kolter. 2023 · 2023
Later among the works it cites.
Structured Pruning for Efficient Generative Pre-trained Language Models. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023 , Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, 10880–10895
Chaofan Tao, Lu Hou, Haoli Bai, Jiansheng Wei, Xin Jiang, Qun Liu, Ping Luo, and Ngai Wong. 2023 · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a · 2023
Later among the works it cites.
DyLoRA: Parameter-Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank Adaptation. In EACL 2023 , Andreas Vlachos and Isabelle Augenstein (Eds.). ACL, 3266–3279
Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, and Ali Ghodsi. 2023 · 2023
Later among the works it cites.
Understanding INT4 Quantization for Transformer Models: Latency Speedup, Composability, and Failure Cases
Xiaoxia Wu, Cheng Li, Reza Yazdani Aminabadi, Zhewei Yao, and Yuxiong He. 2023a · 2023
Later among the works it cites.
ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats
Xiaoxia Wu, Zhewei Yao, and Yuxiong He. 2023b · 2023
Later among the works it cites.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. In ICML 2023 (PMLR, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). 38087–38099
Guangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu, Julien Demouth, and Song Han. 2023 · 2023
Later among the works it cites.
ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation
Zhewei Yao, Xiaoxia Wu, Cheng Li, Stephen Youn, and Yuxiong He. 2023 · 2023
Later among the works it cites.
Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
Lu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh, Yaqing Wang, Yiling Jia, Mykola Pechenizkiy, Yi Liang, Zhangyang Wang, and Shiwei Liu. 2023 · 2023
Later among the works it cites.
RPTQ: Reorder-based Post-training Quantization for Large Language Models
Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu. 2023 · 2023
Later among the works it cites.
NUPES : Non-Uniform Post-Training Quantization via Power Exponent Search
Edouard Yvinec, Arnaud Dapogny, and Kevin Bailly. 2023a · 2023
Later among the works it cites.
PowerQuant: Automorphism Search for Non-Uniform Quantization. In ICLR 2023
Edouard Yvinec, Arnaud Dapogny, Matthieu Cord, and Kevin Bailly. 2023b · 2023
Later among the works it cites.
LoRAPrune: Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning
Mingyang Zhang, Hao Chen, Chunhua Shen, Zhen Yang, Linlin Ou, Xinyi Yu, and Bohan Zhuang. 2023b · 2023
Later among the works it cites.
Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023a · 2023
Later among the works it cites.
Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models. In ICLR
Seungcheol Park, Hojun Choi, and U Kang. 2024 · 2024
Closest in time.