Fetching the paper…
Reading the bibliography…
In this work, we present a novel approach for simultaneous knowledge transfer and model compression called Weight Squeezing.
Tensorized embedding layers for efficient model compression
Khrulkov, V.; Hrinchuk, O.; Mirvakhabova, L.; and Oseledets, I. 2019 · 1901
Earlier work this paper cites.
Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
Tang, R.; Lu, Y.; Liu, L.; Mou, L.; Vechtomova, O.; and Lin, J. 2019 · 1903
Earlier work this paper cites.
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Turc, I.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 1908
Earlier work this paper cites.
Reweighted Proximal Pruning for Large-Scale Language Representation
Guo, F.-M.; Liu, S.; Mungall, F. S.; Lin, X.; and Wang, Y. 2019 · 1909
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding
Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2019 · 1909
Earlier work this paper cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2019 · 1909
Earlier work this paper cites.
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
Shen, S.; Dong, Z.; Ye, J.; Ma, L.; Yao, Z.; Gholami, A.; Mahoney, M. W.; and Keutzer, K. 2019 · 1909
Earlier work this paper cites.
Extreme Language Model Compression with Optimal Subwords and Shared Projections
Zhao, S.; Gupta, R.; Song, Y.; and Zhou, D. 2019 · 1909
Earlier work this paper cites.
Structured Pruning of a BERT-based Question Answering Model
McCarley, J.; Chakravarti, R.; and Sil, A. 2019 · 1910
Earlier work this paper cites.
Distilling Transformers into Simple Neural Networks with Unlabeled Transfer Data
Mukherjee, S.; and Awadallah, A. H. 2019 · 1910
Earlier work this paper cites.
Zafrir, O.; Boudoukh, G.; Izsak, P.; and Wasserblat, M. 2019 · 1910
Earlier work this paper cites.
AdaBERT: Task-Adaptive BERT Compression with Differentiable Neural Architecture Search
Chen, D.; Li, Y.; Qiu, M.; Wang, Z.; Bofang Li, B. D.; Deng, H.; Huang, J.; Lin, W.; and Zhou, J. 2020 · 2001
Earlier work this paper cites.
Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
Ganesh, P.; Chen, Y.; Lou, X.; Khan, M. A.; Yang, Y.; Chen, D.; Winslett, M.; Sajjad, H.; and Nakov, P. 2020 · 2002
Cited alongside, same era.
Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning
Gordon, M. A.; Duh, K.; and Andrews, N. 2020 · 2002
Cited alongside, same era.
Pre-trained Models for Natural Language Processing: A Survey
Qiu, X.; Sun, T.; Xu, Y.; Shao, Y.; Dai, N.; and Huang, X. 2020 · 2003
Cited alongside, same era.
LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression
Mao, Y.; Wang, Y.; Wu, C.; Zhang, C.; Wang, Y.; Yang, Y.; Zhang, Q.; Tong, Y.; and Bai, J. 2020 · 2004
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D.; and Ba, J. 2015 · 2015
Later among the works it cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Later among the works it cites.
Universal Language Model Fine-tuning for Text Classification
Howard, J.; and Ruder, S. 2018 · 2018
Later among the works it cites.
Online embedding compression for text classification using low rank matrix factorization
Acharya, A.; Goel, R.; Metallinou, A.; and Dhillon, I. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Later among the works it cites.
Reducing Transformer Depth on Demand with Structured Dropout
Fan, A.; Grave, E.; and Joulin, A. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sajjad, H.; Dalvi, F.; Durrani, N.; and Nakov, P. 2020 · 2004
Cited alongside, same era.
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
Sun, Z.; Yu, H.; Song, X.; Liu, R.; Yang, Y.; and Zhou, D. 2020 · 2004
Cited alongside, same era.
Tensor-train decomposition
Oseledets, I. V. 2011 · 2011
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G. S.; and Dean, J. 2013 · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
Ba, J.; and Caruana, R. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Cited alongside, same era.
Fitnets: Hints for thin deep nets
Romero, A.; Ballas, N.; Kahou, S. E.; Chassang, A.; Gatta, C.; and Bengio, Y. 2014 · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; ; and Dean, J. 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 2019
Later among the works it cites.
Compressing Word Embeddings via Deep Compositional Code Learning
Shu, R.; and Nakayama, H. 2019 · 2019
Later among the works it cites.
Patient Knowledge Distillation for BERT Model Compression
Sun, S.; Cheng, Y.; Gan, Z.; and Liu, J. 2019 · 2019
Later among the works it cites.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Voita, E.; Talbot, D.; Moiseev, F.; Sennrich, R.; and Titov, I. 2019 · 2019
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, R. S. 2019 · 2019
Later among the works it cites.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Katharopoulos, A.; Vyas, A.; Pappas, N.; and Fleuret, F. 2020 · 2020
Closest in time.