Fetching the paper…
Reading the bibliography…
Recurrent Neural Network (RNN) are a popular choice for modeling temporal and sequential tasks and achieve many state-of-the-art performance on various complex problems.
J. L. Elman, “Finding structure in time,” Cognitive science , vol. 14, no. 2, pp. 179–211, 1990
1990
Earlier work this paper cites.
G. E. Hinton and D. Van Camp, “Keeping the neural networks simple by minimizing the description length of the weights,” in Proceedings of the sixth annual conference on Computational learning theory . ACM, 1993, pp. 5–13
1993
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” Neural Networks, IEEE Transactions on , vol. 5, no. 2, pp. 157–166, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber, “Gradient flow in recurrent nets: the difficulty of learning long-term dependencies,” 2001
2001
Earlier work this paper cites.
M. Bay, A. F. Ehmann, and J. S. Downie, “Evaluation of multiple-f0 estimation and tracking systems.” in 2009 International Society for Music Information Retrieval Conference (ISMIR) , 2009, pp. 315–320
2009
Earlier work this paper cites.
M. Schuster, “Speech recognition for mobile devices at Google,” in Pacific Rim International Conference on Artificial Intelligence . Springer, 2010, pp. 8–10
2010
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in In Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS’10). Society for Artificial Intelligence and Statistics , 2010
2010
Earlier work this paper cites.
I. V. Oseledets, “Tensor-train decomposition,” SIAM Journal on Scientific Computing , vol. 33, no. 5, pp. 2295–2317, 2011. [Online]. Available: http://dx.doi.org/10.1137/090752286
2011
Earlier work this paper cites.
J. Martens and I. Sutskever, “Learning recurrent neural networks with Hessian-free optimization,” in Proceedings of the 28th International Conference on Machine Learning (ICML-11) , 2011, pp. 1033–1040
2011
Earlier work this paper cites.
A. Graves, “Practical variational inference for neural networks,” in Advances in Neural Information Processing Systems , 2011, pp. 2348–2356
2011
Earlier work this paper cites.
A. Graves et al. , Supervised sequence labelling with recurrent neural networks . Springer, 2012, vol. 385
2012
Earlier work this paper cites.
N. Boulanger-lewandowski, Y. Bengio, and P. Vincent, “Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription,” in Proceedings of the 29th International Conference on Machine Learning (ICML-12) , J. Langford and J. Pineau, Eds. New York, NY, USA: ACM, 2012, pp. 1159–1166. [Online]. Available: http://icml.cc/2012/papers/590.pdf
2012
Cited alongside, same era.
M. Denil, B. Shakibi, L. Dinh, N. de Freitas et al. , “Predicting parameters in deep learning,” in Advances in Neural Information Processing Systems , 2013, pp. 2148–2156
2013
Cited alongside, same era.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on . IEEE, 2013, pp. 6645–6649
2013
Cited alongside, same era.
A. Graves, N. Jaitly, and A.-r. Mohamed, “Hybrid speech recognition with deep bidirectional LSTM,” in Automatic Speech Recognition and Understanding (ASRU), 2013 IEEE Workshop on . IEEE, 2013, pp. 273–278
2014
Later among the works it cites.
2015
Later among the works it cites.
2015
Later among the works it cites.
A. Novikov, D. Podoprikhin, A. Osokin, and D. P. Vetrov, “Tensorizing neural networks,” in Advances in Neural Information Processing Systems , 2015, pp. 442–450
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
T. N. Sainath, B. Kingsbury, V. Sindhwani, E. Arisoy, and B. Ramabhadran, “Low-rank matrix factorization for deep neural network training with high-dimensional output targets,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 6655–6659
2013
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in neural information processing systems , 2014, pp. 3104–3112
2014
Cited alongside, same era.
J. Ba and R. Caruana, “Do deep nets really need to be deep?” in Advances in neural information processing systems , 2014, pp. 2654–2662
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2015
Later among the works it cites.
S. Gupta, A. Agrawal, K. Gopalakrishnan, and P. Narayanan, “Deep learning with limited numerical precision,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 , 2015, pp. 1737–1746
2015
Later among the works it cites.
M. Courbariaux, Y. Bengio, and J.-P. David, “BinaryConnect: Training deep neural networks with binary weights during propagations,” in Advances in Neural Information Processing Systems , 2015, pp. 3123–3131
2015
Later among the works it cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: http://tensorflow.org/
2015
Later among the works it cites.
2016
Later among the works it cites.
Z. Tang, D. Wang, and Z. Zhang, “Recurrent neural network training with dark knowledge transfer,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 5900–5904
2016
Later among the works it cites.
2016
Later among the works it cites.