Fetching the paper…
Reading the bibliography…
We rigorously evaluate three state-of-the-art techniques for inducing sparsity in deep neural networks on two large-scale learning tasks: Transformer trained on WMT 2014 English-to-German, and ResNet-50 trained on ImageNet.
Bayesian Variable Selection in Linear Regression
Mitchell, T. J. and Beauchamp, J. J · 1988
Earlier work this paper cites.
Optimal Brain Damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D. G · 1992
Earlier work this paper cites.
Sparse Connection and Pruning in Large Dynamic Artificial Neural Networks
Ström, N · 1997
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Memory Bounded Deep Convolutional Networks
Collins, M. D. and Kohli, P · 2014
Earlier work this paper cites.
Stochastic Backpropagation and Approximate Inference in Deep Generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Learning both Weights and Connections for Efficient Neural Network
Han, S., Pool, J., Tran, J., and Dally, W. J · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Earlier work this paper cites.
Dynamic Network Surgery for Efficient DNNs
Guo, Y., Yao, A., and Chen, Y · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Pruning Convolutional Neural Networks for Resource Efficient Transfer Learning
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
Wavenet: A Generative Model for Raw Audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A. W., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Wide Residual Networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Deep Rewiring: Training Very Sparse Deep Networks
Bellec, G., Kappel, D., Maass, W., and Legenstein, R. A · 2017
Cited alongside, same era.
Block-sparse gpu kernels
Gray, S., Radford, A., and Kingma, D. P · 2017
Cited alongside, same era.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G. F., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Cited alongside, same era.
Runtime neural pruning
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Later among the works it cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Later among the works it cites.
Compressing Neural Networks using the Variational Information Bottleneck
Dai, B., Zhu, C., and Wipf, D. P · 2018
Later among the works it cites.
The Lottery Ticket Hypothesis: Training Pruned Neural Networks
Frankle, J. and Carbin, M · 2018
Later among the works it cites.
AMC: automl for model compression and acceleration on mobile devices
He, Y., Lin, J., Liu, Z., Wang, H., Li, L., and Han, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, J., Rao, Y., Lu, J., and Zhou, J · 2017
Cited alongside, same era.
Learning Efficient Convolutional Networks through Network Slimming
Liu, Z., Li, J., Shen, Z., Huang, G., Yan, S., and Zhang, C · 2017
Cited alongside, same era.
Bayesian Compression for Deep Learning
Louizos, C., Ullrich, K., and Welling, M · 2017
Cited alongside, same era.
Thinet: A Filter Level Pruning Method for Deep Neural Network Compression
Luo, J., Wu, J., and Lin, W · 2017
Cited alongside, same era.
Variational Dropout Sparsifies Deep Neural Networks
Molchanov, D., Ashukha, A., and Vetrov, D. P · 2017
Cited alongside, same era.
Exploring Sparsity in Recurrent Neural Networks
Narang, S., Diamos, G. F., Sengupta, S., and Elsen, E · 2017
Cited alongside, same era.
Soft Weight-Sharing for Neural Network Compression
Ullrich, K., Meeds, E., and Welling, M · 2017
Cited alongside, same era.
Efficient Neural Audio Synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., van den Oord, A., Dieleman, S., and Kavukcuoglu, K · 2018
Later among the works it cites.
Rethinking the Value of Network Pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2018
Later among the works it cites.
Scalable Training of Artificial Neural Networks with Adaptive Sparse Connectivity Inspired by Network Science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Later among the works it cites.
Scaling Neural Machine Translation
Ott, M., Edunov, S., Grangier, D., and Auli, M · 2018
Later among the works it cites.
Faster gaze prediction with dense networks and Fisher pruning
Theis, L., Korshunova, I., Tejani, A., and Huszár, F · 2018
Later among the works it cites.
Lpcnet: Improving Neural Speech Synthesis Through Linear Prediction
Valin, J. and Skoglund, J · 2018
Later among the works it cites.