Fetching the paper…
Reading the bibliography…
Despite achieving state-of-the-art performance on many NLP tasks, the high energy cost and long inference delay prevent Transformer-based pretrained language models (PLMs) from seeing broader adoption including for edge and mobile computing.
Distilling task-specific knowledge from bert into simple neural networks
Tang, R.; Lu, Y.; Liu, L.; et al. 2019 · 1903
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Turc, I.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 1908
Earlier work this paper cites.
On the effectiveness of low-rank matrix factorization for lstm model compression
Winata, G. I.; Madotto, A.; Shin, J.; et al. 2019 · 1908
Earlier work this paper cites.
Quantifying the carbon emissions of machine learning
Lacoste, A.; Luccioni, A.; Schmidt, V.; and Dandres, T. 2019 · 1910
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
Zafrir, O.; Boudoukh, G.; Izsak, P.; and Wasserblat, M. 2019 · 1910
Earlier work this paper cites.
Optimal Brain Damage
LeCun, Y.; Denker, J. S.; and Solla, S. A. 1989 · 1989
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A.; and Teboulle, M. 2003 · 2003
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Yuan, M.; and Lin, Y. 2006 · 2006
Earlier work this paper cites.
Decompositions of a Higher-Order Tensor in Block Terms - Part II: Definitions and Uniqueness
Lathauwer, L. D. 2008 · 2008
Earlier work this paper cites.
Fastformers: Highly efficient transformer models for natural language understanding
Kim, Y. J.; and Awadalla, H. H. 2020 · 2010
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y.; Léonard, N.; and Courville, A. 2013 · 2013
Earlier work this paper cites.
Low-rank matrix factorization for Deep Neural Network training with high-dimensional output targets
Sainath, T. N.; Kingsbury, B.; Sindhwani, V.; et al. 2013 · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; Dean, J.; et al. 2015 · 2015
Earlier work this paper cites.
Big/little deep neural network for ultra low power inference
Park, E.; Kim, D.; Kim, S.; et al. 2015 · 2015
Earlier work this paper cites.
Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding
Han, S.; Mao, H.; and Dally, W. J. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
Compression of Neural Machine Translation Models via Pruning
See, A.; Luong, M.; and Manning, C. D. 2016 · 2016
Earlier work this paper cites.
BranchyNet: Fast inference via early exiting from deep neural networks
Teerapittayanon, S.; McDanel, B.; and Kung, H. T. 2016 · 2016
Earlier work this paper cites.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Finn, C.; Abbeel, P.; and Levine, S. 2017 · 2017
Earlier work this paper cites.
Neural Networks Compression for Language Modeling
Grachev, A. M.; Ignatov, D. I.; and Savchenko, A. V. 2017 · 2017
Earlier work this paper cites.
Exploring Sparsity in Recurrent Neural Networks
Narang, S.; Diamos, G.; Sengupta, S.; and Elsen, E. 2017 · 2017
Earlier work this paper cites.
Block-sparse recurrent neural networks
Narang, S.; Undersander, E.; and Diamos, G. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; et al. 2017 · 2017
Earlier work this paper cites.
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Jacob, B.; Kligys, S.; Chen, B.; et al. 2018 · 2018
Earlier work this paper cites.
Learning Sparse Neural Networks through L_0 Regularization
Louizos, C.; Welling, M.; and Kingma, D. P. 2018 · 2018
Earlier work this paper cites.
Deep Contextualized Word Representations
Peters, M. E.; Neumann, M.; Iyyer, M.; et al. 2018 · 2018
Earlier work this paper cites.
Is Robustness the Cost of Accuracy? - A Comprehensive Study on the Robustness of 18 Deep Image Classification Models
Su, D.; Zhang, H.; Chen, H.; et al. 2018 · 2018
Earlier work this paper cites.
mixup: Beyond Empirical Risk Minimization
Zhang, H.; Cissé, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2018 · 2018
Earlier work this paper cites.
Recurrent Stacking of Layers for Compact Neural Machine Translation Models
Dabre, R.; and Fujita, A. 2019 · 2019
Earlier work this paper cites.
Universal Transformers
Dehghani, M.; Gouws, S.; Vinyals, O.; et al. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision
Dong, Z.; Yao, Z.; Gholami, A.; et al. 2019 · 2019
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Frankle, J.; and Carbin, M. 2019 · 2019
Earlier work this paper cites.
Shallow-Deep Networks: Understanding and Mitigating Network Overthinking
Kaya, Y.; Hong, S.; and Dumitras, T. 2019 · 2019
Earlier work this paper cites.
DARTS: Differentiable Architecture Search
Liu, H.; Simonyan, K.; and Yang, Y. 2019 · 2019
Earlier work this paper cites.
A Tensorized Transformer for Language Modeling
Ma, X.; Zhang, P.; Zhang, S.; et al. 2019 · 2019
Earlier work this paper cites.
Are Sixteen Heads Really Better than One?
Michel, P.; Levy, O.; and Neubig, G. 2019 · 2019
Cited alongside, same era.
Patient Knowledge Distillation for BERT Model Compression
Sun, S.; Cheng, Y.; Gan, Z.; and Liu, J. 2019 · 2019
Cited alongside, same era.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Voita, E.; Talbot, D.; Moiseev, F.; et al. 2019 · 2019
Cited alongside, same era.
Tied Transformers: Neural Machine Translation with Shared Encoder and Decoder
Xia, Y.; He, T.; Tan, X.; et al. 2019 · 2019
Cited alongside, same era.
Knowledge Distillation from Internal Representations
Aguilar, G.; Ling, Y.; Zhang, Y.; et al. 2020 · 2020
Cited alongside, same era.
The Lottery Ticket Hypothesis for Pre-trained BERT Networks
Chen, T.; Frankle, J.; Chang, S.; et al. 2020 · 2020
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Bender, E. M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S. 2021 · 2021
Later among the works it cites.
DRONE: Data-aware Low-rank Compression for Large NLP Models
Chen, P.-H.; Yu, H.-F.; Dhillon, I.; and Hsieh, C.-J. 2021 · 2021
Later among the works it cites.
What do Compressed Large Language Models Forget? Robustness Challenges in Model Compression
Du, M.; Mukherjee, S.; Cheng, Y.; et al. 2021 · 2021
Later among the works it cites.
Romebert: Robust training of multi-exit bert
Geng, S.; Gao, P.; Fu, Z.; and Zhang, Y. 2021 · 2021
Later among the works it cites.
Parameter-Efficient Transfer Learning with Diff Pruning
Guo, D.; Rush, A. M.; and Kim, Y. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reducing Transformer Depth on Demand with Structured Dropout
Fan, A.; Grave, E.; and Joulin, A. 2020 · 2020
Cited alongside, same era.
Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning
Gordon, M. A.; Duh, K.; and Andrews, N. 2020 · 2020
Cited alongside, same era.
PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector Elimination
Goyal, S.; Choudhury, A. R.; Raje, S.; et al. 2020 · 2020
Cited alongside, same era.
Towards the systematic reporting of the energy and carbon footprints of machine learning
Henderson, P.; Hu, J.; Romoff, J.; et al. 2020 · 2020
Cited alongside, same era.
DynaBERT: Dynamic BERT with Adaptive Width and Depth
Hou, L.; Huang, Z.; Shang, L.; et al. 2020 · 2020
Cited alongside, same era.
TinyBERT: Distilling BERT for Natural Language Understanding
Jiao, X.; Yin, Y.; Shang, L.; et al. 2020 · 2020
Cited alongside, same era.
Pre-trained models: Past, present and future
Han, X.; Zhang, Z.; Ding, N.; et al. 2021 · 2021
Later among the works it cites.
Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
Kim, G.; and Cho, K. 2021 · 2021
Later among the works it cites.
I-BERT: Integer-only BERT Quantization
Kim, S.; Gholami, A.; Yao, Z.; et al. 2021 · 2021
Later among the works it cites.
Block Pruning For Faster Transformers
Lagunas, F.; Charlaix, E.; Sanh, V.; and Rush, A. M. 2021 · 2021
Later among the works it cites.
MixKD: Towards Efficient Distillation of Large-scale Language Models
Liang, K. J.; Hao, W.; Shen, D.; et al. 2021 · 2021
Later among the works it cites.
A Global Past-Future Early Exit Method for Accelerating Inference of Pre-trained Language Models
Liao, K.; Zhang, Y.; Ren, X.; et al. 2021 · 2021
Later among the works it cites.
Towards Efficient NLP: A Standard Evaluation and A Strong Baseline
Liu, X.; Sun, T.; He, J.; et al. 2021 · 2021
Later among the works it cites.
Accelerating sparse deep neural networks
Mishra, A.; Latorre, J. A.; Pool, J.; et al. 2021 · 2021
Later among the works it cites.
Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers
Reid, M.; Marrese-Taylor, E.; and Matsuo, Y. 2021 · 2021
Later among the works it cites.
Consistent accelerated inference via confident adaptive transformers
Schuster, T.; Fisch, A.; Jaakkola, T. S.; and Barzilay, R. 2021 · 2021
Later among the works it cites.
Does Knowledge Distillation Really Work?
Stanton, S.; Izmailov, P.; Kirichenko, P.; et al. 2021 · 2021
Later among the works it cites.
Early Exiting with Ensemble Internal Classifiers
Sun, T.; Zhou, Y.; Liu, X.; et al. 2021 · 2021
Later among the works it cites.
Tahaei, M. S.; Charlaix, E.; Nia, V. P.; et al. 2021 · 2021
Later among the works it cites.
Lessons on parameter sharing across layers in transformers
Takase, S.; and Kiyono, S. 2021 · 2021
Later among the works it cites.
MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers
Wang, W.; Bao, H.; Huang, S.; et al. 2021 · 2021
Later among the works it cites.
One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers
Wu, C.; Wu, F.; and Huang, Y. 2021 · 2021
Later among the works it cites.
BERxiT: Early Exiting for BERT with Better Fine-Tuning and Extension to Regression
Xin, J.; Tang, R.; Yu, Y.; and Lin, J. 2021 · 2021
Later among the works it cites.
TR-BERT: Dynamic Token Reduction for Accelerating BERT Inference
Ye, D.; Lin, Y.; Huang, Y.; and Sun, M. 2021 · 2021
Later among the works it cites.
LeeBERT: Learned Early Exit for BERT with cross-level optimization
Zhu, W. 2021 · 2021
Later among the works it cites.
Kronecker Decomposition for GPT Compression
Edalati, A.; Tahaei, M. S.; Rashid, A.; et al. 2022 · 2022
Closest in time.
Transkimmer: Transformer Learns to Layer-wise Skim
Guan, Y.; Li, Z.; Leng, J.; et al. 2022 · 2022
Closest in time.
Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm
Huang, S.; Xu, D.; Yen, I. E.; et al. 2022 · 2022
Closest in time.
Learned Token Pruning for Transformers
Kim, S.; Shen, S.; Thorsley, D.; et al. 2022 · 2022
Closest in time.
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
Le Scao, T.; Fan, A.; Akiki, C.; et al. 2022 · 2022
Closest in time.
Multi-Granularity Structural Knowledge Distillation for Language Model Compression
Liu, C.; Tao, C.; Feng, J.; and Zhao, D. 2022 · 2022
Closest in time.
Knowledge Distillation with Reptile Meta-Learning for Pretrained Language Model Compression
Ma, X.; Wang, J.; Yu, L.; and Zhang, X. 2022 · 2022
Closest in time.
Compression of Generative Pre-trained Language Models via Quantization
Tao, C.; Hou, L.; Zhang, W.; et al. 2022 · 2022
Closest in time.
SkipBERT: Efficient Inference with Shallow Layer Skipping
Wang, J.; Chen, K.; Chen, G.; et al. 2022 · 2022
Closest in time.
Structured Pruning Learns Compact and Accurate Models
Xia, M.; Zhong, Z.; and Chen, D. 2022 · 2022
Closest in time.
PCEE-BERT: Accelerating BERT Inference via Patient and Confident Early Exiting
Zhang, Z.; Zhu, W.; Zhang, J.; et al. 2022 · 2022
Closest in time.
BERT Learns to Teach: Knowledge Distillation with Meta Learning
Zhou, W.; Xu, C.; and McAuley, J. J. 2022 · 2022
Closest in time.