Fetching the paper…
Reading the bibliography…
Sparse neural networks have been widely applied to reduce the computational demands of training and deploying over-parameterized deep neural networks.
A distance measure between attributed relational graphs for pattern recognition
Sanfeliu, A. and Fu, K.-S · 1983
Earlier work this paper cites.
Pruning versus clipping in neural networks
Janowsky, S. A · 1989
Earlier work this paper cites.
Using relevance to reduce network size automatically
Mozer, M. C. and Smolensky, P · 1989
Earlier work this paper cites.
Finding structure in time
Elman, J. L · 1990
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcus, M., Santorini, B., and Marcinkiewicz, M. A · 1993
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Deep read: A reading comprehension system
Hirschman, L., Light, M., Breck, E., and Burger, J. D · 1999
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Černockỳ, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Recurrent continuous translation models
Kalchbrenner, N. and Blunsom, P · 2013
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J., Vinyals, O., and Saxe, A. M · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Jaderberg, M., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Earlier work this paper cites.
Training very deep networks
Srivastava, R. K., Greff, K., and Schmidhuber, J · 2015
Earlier work this paper cites.
A theoretically grounded application of dropout in recurrent neural networks
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J. E · 2016
Earlier work this paper cites.
A topological insight into restricted boltzmann machines
Mocanu, D. C., Mocanu, E., Nguyen, P. H., Gibescu, M., and Liotta, A · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
Grave, E., Joulin, A., and Usunier, N · 2017
Cited alongside, same era.
Gpu kernels for block-sparse weights
Gray, S., Radford, A., and Kingma, D. P · 2017
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Inan, H., Khosravi, K., and Socher, R · 2017
Cited alongside, same era.
Learning sparse neural networks through l _ 0 l\_0 regularization
Louizos, C., Welling, M., and Kingma, D. P · 2017
Cited alongside, same era.
Sparse networks from scratch: Faster training without losing performance
Dettmers, T. and Zettlemoyer, L · 2019
Later among the works it cites.
Large scale structure of neural network loss landscapes
Fort, S. and Jastrzebski, S · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2019
Later among the works it cites.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Later among the works it cites.
Radix-net: Structured sparse matrices for deep neural networks
Kepner, J. and Robinett, R · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Melis, G., Dyer, C., and Blunsom, P · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Molchanov, D., Ashukha, A., and Vetrov, D · 2017
Cited alongside, same era.
Machine comprehension using match-lstm and answer pointer
Wang, S. and Jiang, J · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2017
Cited alongside, same era.
Deep rewiring: Training very sparse deep networks
Bellec, G., Kappel, D., Maass, W., and Legenstein, R · 2018
Cited alongside, same era.
Fast training and model compression of gated rnns via singular value decomposition
Dai, R., Li, L., and Yu, W · 2018
Cited alongside, same era.
Lee, N., Ajanthan, T., and Torr, P · 2019
Later among the works it cites.
Intrinsically sparse long short-term memory networks
Liu, S., Mocanu, D. C., and Pechenizkiy, M · 2019
Later among the works it cites.
Transformed l1 regularization for learning sparse deep neural networks
Ma, R., Miao, J., Niu, L., and Zhang, P · 2019
Later among the works it cites.
Sparsemaps: convolutional networks with sparse feature maps for tiny image classification
Moradi, R., Berangi, R., and Minaei, B · 2019
Later among the works it cites.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Mostafa, H. and Wang, X · 2019
Later among the works it cites.
Ordered neurons: Integrating tree structures into recurrent neural networks
Shen, Y., Tan, S., Sordoni, A., and Courville, A · 2019
Later among the works it cites.
Autoprune: Automatic network pruning by regularizing auxiliary parameters
Xiao, X., Wang, Z., and Rajasekaran, S · 2019
Later among the works it cites.
Feed-forward neural network training using sparse representation
Yang, J. and Ma, J · 2019
Later among the works it cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H., Lan, J., Liu, R., and Yosinski, J · 2019
Later among the works it cites.
Progressive skeletonization: Trimming more fat from a network at initialization
de Jorge, P., Sanyal, A., Behl, H. S., Torr, P. H., Rogez, G., and Dokania, P. K · 2020
Later among the works it cites.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Later among the works it cites.
Top-kast: Top-k always sparse training
Jayakumar, S., Pascanu, R., Rae, J., Osindero, S., and Elsen, E · 2020
Later among the works it cites.
Soft threshold weight reparameterization for learnable sparsity
Kusupati, A., Ramanujan, V., Somani, R., Wortsman, M., Jain, P., Kakade, S., and Farhadi, A · 2020
Later among the works it cites.
A signal propagation perspective for pruning neural networks at initialization
Lee, N., Ajanthan, T., Gould, S., and Torr, P. H. S · 2020
Later among the works it cites.
Nvidia a100 tensor core gpu architecture
NVIDIA · 2020
Later among the works it cites.
Sparse weight activation training
Raihan, M. A. and Aamodt, T. M · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H., Kunin, D., Yamins, D. L., and Ganguli, S · 2020
Later among the works it cites.
Picking winning tickets before training by preserving gradient flow
Wang, C., Zhang, G., and Grosse, R · 2020
Later among the works it cites.