Fetching the paper…
Reading the bibliography…
Pruning is an effective method to reduce the memory footprint and computational cost associated with large natural language processing models.
Linformer: Self-attention with linear complexity
Wang, S., Li, B., Khabsa, M., Fang, H., and Ma, H · 2006
Earlier work this paper cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Wang, H., Zhang, Z., and Han, S · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Dong, X., Chen, S., and Pan, S · 2017
Earlier work this paper cites.
First quora dataset release: Question pairs, 2017
Iyer, S., Dandekar, N., and Csernai, K · 2017
Earlier work this paper cites.
Thinet: A filter level pruning method for deep neural network compression
Luo, J.-H., Wu, J., and Lin, W · 2017
Earlier work this paper cites.
Block-sparse recurrent neural networks
Narang, S., Undersander, E., and Diamos, G · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Williams, A., Nangia, N., and Bowman, S. R · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Amc: Automl for model compression and acceleration on mobile devices
He, Y., Lin, J., Liu, Z., Wang, H., Li, L.-J., and Han, S · 2018
Earlier work this paper cites.
Data-driven sparse structure selection for deep neural networks
Huang, Z. and Wang, N · 2018
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H · 2018
Earlier work this paper cites.
Accelerating convolutional networks via global & dynamic filter pruning
Lin, S., Ji, R., Li, Y., Wu, Y., Huang, F., and Zhang, B · 2018
Earlier work this paper cites.
Nisp: Pruning networks using neuron importance score propagation
Yu, R., Li, A., Chen, C.-F., Lai, J.-H., Morariu, V. I., Han, X., Gao, M., Lin, C.-Y., and Davis, L. S · 2018
Cited alongside, same era.
Efficient 8-bit quantization of transformer neural machine language translation model
Bhandare, A., Sripathi, V., Karkada, D., Menon, V., Choi, S., Datta, K., and Saletore, V · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S · 2019
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
Fan, A., Grave, E., and Joulin, A · 2019
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
The lottery ticket hypothesis for pre-trained BERT networks
Chen, T., Frankle, J., Chang, S., Liu, S., Zhang, Y., Wang, Z., and Carbin, M · 2020
Later among the works it cites.
Training with quantization noise for extreme fixed-point compression
Fan, A., Stock, P., Graham, B., Grave, E., Gribonval, R., Jegou, H., and Joulin, A · 2020
Later among the works it cites.
SqueezeBERT: What can computer vision teach NLP about efficient neural networks?
Iandola, F. N., Shaw, A. E., Krishna, R., and Keutzer, K. W · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
TinyBERT: Distilling BERT for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2019
Cited alongside, same era.
ALBERT: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Cited alongside, same era.
DistilBERT, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Megatron-LM: Training multi-billion parameter language models using gpu model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Cited alongside, same era.
Patient knowledge distillation for bert model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Cited alongside, same era.
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Later among the works it cites.
Lookahead: a far-sighted alternative of magnitude-based pruning
Park, S., Lee, J., Mo, S., and Shin, J · 2020
Later among the works it cites.
When BERT plays the lottery, all tickets are winning
Prasanna, S., Rogers, A., and Rumshisky, A · 2020
Later among the works it cites.
Poor man’s BERT: Smaller and faster transformer models
Sajjad, H., Dalvi, F., Durrani, N., and Nakov, P · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A. M · 2020
Later among the works it cites.
Q-BERT: Hessian based ultra low precision quantization of bert
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Later among the works it cites.
MobileBERT: a compact task-agnostic BERT for resource-limited devices
Sun, Z., Yu, H., Song, X., Liu, R., Yang, Y., and Zhou, D · 2020
Later among the works it cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Zadeh, A. H., Edo, I., Awad, O. M., and Moshovos, A · 2020
Later among the works it cites.
Ternarybert: Distillation-aware ultra-low bit bert
Zhang, W., Hou, L., Yin, Y., Shang, L., Chen, X., Jiang, X., and Liu, Q · 2020
Later among the works it cites.
Masking as an efficient alternative to finetuning for pretrained language models
Zhao, M., Lin, T., Mi, F., Jaggi, M., and Schütze, H · 2020
Later among the works it cites.
Block pruning for faster transformers
Lagunas, F., Charlaix, E., Sanh, V., and Rush, A. M · 2021
Closest in time.
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, the World’s Largest and Most Powerful Generative Language Model
Microsoft and Nvidia · 2021
Closest in time.
What’s hidden in a one-layer randomly weighted transformer?
Shen, S., Yao, Z., Kiela, D., Keutzer, K., and Mahoney, M. W · 2021
Closest in time.
Hessian-aware pruning and optimal neural implant
Yu, S., Yao, Z., Gholami, A., Dong, Z., Mahoney, M. W., and Keutzer, K · 2021
Closest in time.