Fetching the paper…
Reading the bibliography…
Long Short-Term Memory (LSTM) has achieved state-of-the-art performances on a wide range of tasks.
Pruning recurrent neural networks for improved generalization performance
Giles, C. L. and Omlin, C. W · 1994
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Hybrid speech recognition with deep bidirectional lstm
Graves, A., Jaitly, N., and Mohamed, A.-r · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Learning compact recurrent neural networks
Lu, Z., Sindhwani, V., and Sainath, T. N · 2016
Earlier work this paper cites.
Hierarchical attention networks for document classification
Yang, Z., Yang, D., Dyer, C., He, X., Smola, A., and Hovy, E · 2016
Earlier work this paper cites.
Zen, H., Agiomyrgiannakis, Y., Egberts, N., Henderson, F., and Szczepaniak, P · 2016
Cited alongside, same era.
Deep rewiring: Training very sparse deep networks
Bellec, G., Kappel, D., Maass, W., and Legenstein, R · 2017
Cited alongside, same era.
Nest: a neural network synthesis tool based on a grow-and-prune paradigm
Dai, X., Yin, H., and Jha, N. K · 2017
Cited alongside, same era.
Ese: Efficient speech recognition engine with sparse lstm on fpga
Han, S., Kang, J., Mao, H., Hu, Y., Li, X., Li, Y., Xie, D., Luo, H., Yao, S., Wang, Y., et al · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Deep lstm for large vocabulary continuous speech recognition
Tian, X., Zhang, J., Ma, Z., He, Y., Wei, J., Wu, P., Situ, W., Li, S., and Zhang, Y · 2017
Later among the works it cites.
Learning intrinsic sparse structures within long short-term memory
Wen, W., He, Y., Rajbhandari, S., Zhang, M., Wang, W., Liu, F., Hu, B., Chen, Y., and Li, H · 2017
Later among the works it cites.
The loss landscape of overparameterized neural networks
Cooper, Y · 2018
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Later among the works it cites.
Snip: Single-shot network pruning based on connection sensitivity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Factorization tricks for lstm networks
Kuchaiev, O. and Ginsburg, B · 2017
Cited alongside, same era.
Bayesian sparsification of recurrent neural networks
Lobacheva, E., Chirkova, N., and Vetrov, D · 2017
Cited alongside, same era.
Exploring sparsity in recurrent neural networks
Narang, S., Elsen, E., Diamos, G., and Sengupta, S · 2017
Cited alongside, same era.
Lee, N., Ajanthan, T., and Torr, P. H · 2018
Later among the works it cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Later among the works it cites.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization, 2019
Mostafa, H. and Wang, X · 2019
Closest in time.