Fetching the paper…
Reading the bibliography…
A common technique for compressing a neural network is to compute the $k$-rank $\ell_2$ approximation $A_{k,2}$ of the matrix $A\in\mathbb{R}^{n\times d}$ that corresponds to a fully connected layer (or embedding layer).
Distilling task-specific knowledge from bert into simple neural networks
Tang, R.; Lu, Y.; Liu, L.; Mou, L.; Vechtomova, O.; and Lin, J. 2019 · 1903
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019b · 1907
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Sun, S.; Cheng, Y.; Gan, Z.; and Liu, J. 2019 · 1908
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
Fan, A.; Grave, E.; and Joulin, A. 2019 · 1909
Earlier work this paper cites.
Reweighted proximal pruning for large-scale language representation
Guo, F.-M.; Liu, S.; Mungall, F. S.; Lin, X.; and Wang, Y. 2019 · 1909
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2019 · 1909
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2019 · 1909
Earlier work this paper cites.
Extreme language model compression with optimal subwords and shared projections
Zhao, S.; Gupta, R.; Song, Y.; and Zhou, D. 2019 · 1909
Earlier work this paper cites.
Introduction to coresets: Accurate coresets
Jubran, I.; Maalouf, A.; and Feldman, D. 2019 · 1910
Earlier work this paper cites.
Pruning a bert-based question answering model
McCarley, J. S. 2019 · 1910
Earlier work this paper cites.
Distilling transformers into simple neural networks with unlabeled transfer data
Mukherjee, S.; and Awadallah, A. H. 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2019 · 1910
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
Structured pruning of large language models
Wang, Z.; Wohlwend, J.; and Lei, T. 2019 · 1910
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. 2019 · 1910
Earlier work this paper cites.
Zafrir, O.; Boudoukh, G.; Izsak, P.; and Wasserblat, M. 2019 · 1910
Earlier work this paper cites.
Attentive student meets multi-task teacher: Improved knowledge distillation for pretrained models
Liu, L.; Wang, H.; Lin, J.; Socher, R.; and Xiong, C. 2019a · 1911
Earlier work this paper cites.
The approximation of one matrix by another of lower rank
Eckart, C.; and Young, G. 1936 · 1936
Earlier work this paper cites.
The Ellipsoid Method
Grötschel, M.; Lovász, L.; and Schrijver, A. 1993 · 1993
Earlier work this paper cites.
Oriented principal component analysis for large margin classifiers
Bermejo, S.; and Cabestany, J. 2001 · 2001
Earlier work this paper cites.
Compressing BERT: Studying the effects of weight pruning on transfer learning
Gordon, M. A.; Duh, K.; and Andrews, N. 2020 · 2002
Earlier work this paper cites.
Optimally sparse representation in general (nonorthogonal) dictionaries via ℓ 1 \ell_{1} minimization
Donoho, D. L.; and Elad, M. 2003 · 2003
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Sun, Z.; Yu, H.; Song, X.; Liu, R.; Yang, Y.; and Zhou, D. 2020 · 2004
Cited alongside, same era.
R 1-PCA: rotational invariant L 1-norm principal component analysis for robust subspace factorization
Ding, C.; Zhou, D.; He, X.; and Zha, H. 2006 · 2006
Cited alongside, same era.
Compressed sensing
Donoho, D. L. 2006 · 2006
Cited alongside, same era.
Faster PAC Learning and Smaller Coresets via Smoothed Analysis
Maalouf, A.; Jubran, I.; Tukan, M.; and Feldman, D. 2020 · 2006
Cited alongside, same era.
Coresets for Near-Convex Functions
Tukan, M.; Maalouf, A.; and Feldman, D. 2020 · 2006
Cited alongside, same era.
Training deep nets with sublinear memory cost
Chen, T.; Xu, B.; Zhang, C.; and Guestrin, C. 2016 · 2016
Later among the works it cites.
The fast cauchy transform and faster robust linear regression
Clarkson, K. L.; Drineas, P.; Magdon-Ismail, M.; Mahoney, M. W.; Meng, X.; and Woodruff, D. P. 2016 · 2016
Later among the works it cites.
Dimensionality reduction of massive sparse datasets using coresets
Feldman, D.; Volkov, M.; and Rus, D. 2016 · 2016
Later among the works it cites.
Low-rank approximation and regression in input sparsity time
Clarkson, K. L.; and Woodruff, D. P. 2017 · 2017
Later among the works it cites.
The reversible residual network: Backpropagation without storing activations
Gomez, A. N.; Ren, M.; Urtasun, R.; and Grosse, R. B. 2017 · 2017
Later among the works it cites.
Automatic differentiation in PyTorch
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient subspace approximation algorithms
Shyamalkumar, N. D.; and Varadarajan, K. 2007 · 2007
Cited alongside, same era.
Coresets and sketches for high dimensional subspace approximation problems
Feldman, D.; Monemizadeh, M.; Sohler, C.; and Woodruff, D. P. 2010 · 2010
Cited alongside, same era.
A unified framework for approximating and clustering data
Feldman, D.; and Langberg, M. 2011 · 2011
Cited alongside, same era.
Scikit-learn: Machine Learning in Python
Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; Vanderplas, J.; Passos, A.; Cournapeau, D.; Brucher, M.; Perrot, M.; and Duchesnay, E. 2011 · 2011
Cited alongside, same era.
The NumPy array: a structure for efficient numerical computation
Van Der Walt, S.; Colbert, S. C.; and Varoquaux, G. 2011 · 2011
Cited alongside, same era.
Matrix computations , volume 3
Golub, G. H.; and Van Loan, C. F. 2012 · 2012
Cited alongside, same era.
On the Sensitivity of Shape Fitting Problems
Varadarajan, K.; and Xiao, X. 2012 · 2012
Cited alongside, same era.
Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; and Lerer, A. 2017 · 2017
Later among the works it cites.
On compressing deep models by low rank and sparse decomposition
Yu, X.; Liu, T.; Wang, X.; and Tao, D. 2017 · 2017
Later among the works it cites.
Deep contextualized word representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Later among the works it cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. 2018 · 2018
Later among the works it cites.
Online embedding compression for text classification using low rank matrix factorization
Acharya, A.; Goel, R.; Metallinou, A.; and Dhillon, I. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Later among the works it cites.
All The Ways You Can Compress BERT
Gordon, M. A. 2019 · 2019
Later among the works it cites.
Fast and accurate least-mean-squares solvers
Maalouf, A.; Jubran, I.; and Feldman, D. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Michel, P.; Levy, O.; and Neubig, G. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R. R.; and Le, Q. V. 2019 · 2019
Later among the works it cites.
Open source code for all the algorithms presented in this paper
Code. 2020 · 2020
Closest in time.
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
Shen, S.; Dong, Z.; Ye, J.; Ma, L.; Yao, Z.; Gholami, A.; Mahoney, M. W.; and Keutzer, K. 2020 · 2020
Closest in time.
Mean squared error — Wikipedia, The Free Encyclopedia
Wikipedia. 2020 · 2020
Closest in time.
Tight sensitivity bounds for smaller coresets
Maalouf, A.; Statman, A.; and Feldman, D. 2020 · 2061
Closest in time.