Fetching the paper…
Reading the bibliography…
Sequence-to-sequence models, such as attention-based models in automatic speech recognition (ASR), are typically trained to optimize the cross-entropy criterion which corresponds to improving the log-likelihood of the data.
“Simple statistical gradient-following algorithms for connectionist reinforcement learning,”
R. J. Williams, · 1992
Earlier work this paper cites.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
M. Schuster and K. K. Paliwal, · 1997
Earlier work this paper cites.
“BLEU: A method for automatic evaluation of machine translation,”
K. Papineni, S. Roukos, T. Ward, and W-. J. Zhu, · 2002
Earlier work this paper cites.
Discriminative Training for Large Vocabulary Speech Recognition
D. Povey, · 2003
Earlier work this paper cites.
“Connectionist temporal classificatio: Labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Lattice-based optimization of sequence classification criteria for neural-network acoustic modeling,”
B. Kingsbury, · 2009
Earlier work this paper cites.
“Hogwild: A lock-free approach to parallelizing stochastic gradient descent,”
B. Recht, C. Re, S. Wright, and F. Niu, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Japanese and korean voice search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Speech recognition with deep neural networks,”
A. Graves, A-. R. Mohamed, and G. Hinton, · 2013
Earlier work this paper cites.
“Sequence-discriminative training of deep neural networks.,”
K. Veselỳ, A. Ghoshal, L. Burget, and D. Povey, · 2013
Earlier work this paper cites.
“Error back propagation for sequence training of context-dependent deep networks for conversational speech transcription,”
H. Su, G. Li, D. Yu, and F. Seide, · 2013
Cited alongside, same era.
“Towards end-to-end speech recognition with recurrent neural networks,”
A. Graves and N. Jaitly, · 2014
Cited alongside, same era.
“Sequence to sequence learning with neural networks,”
I. Sutskever, O. Vinyals, and Q. V. Le, · 2014
Cited alongside, same era.
“Acoustic modelling with cd-ctc-smbr lstm rnns,”
A. Senior, H. Sak, F. de Chaumont Quitry, T. N. Sainath, and K. Rao, · 2015
Cited alongside, same era.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2015
Cited alongside, same era.
“Scheduled sampling for sequence prediction with recurrent neural networks,”
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, · 2015
“Sequence level training with recurrent neural networks,”
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba, · 2016
Later among the works it cites.
“Lower frame rate neural network acoustic models,”
G. Pundak and T. N. Sainath, · 2016
Later among the works it cites.
“Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping,”
H. Sak, M. Shannon, K. Rao, and F. Beaufays, · 2017
Closest in time.
“Neural speech recognizer: Acoustic-to-word lstm model for large vocabulary speech recognition,”
H. Soltau, H. Liao, and H. Sak, · 2017
Closest in time.
“Exploring architectures, data and units for streaming end-to-end speech recognition with rnn-transducer,”
K. Rao, H. Sak, and R. Prabhavalkar, · 2017
Closest in time.
“A comparison of sequence-to-sequence models for speech recognition,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
H. Sak, A. Senior, K. Rao, and F. Beaufays, · 2015
Cited alongside, same era.
“Tensorflow: Large-scale machine learning on heterogeneous distributed systems,” 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2016
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, · 2016
Cited alongside, same era.
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, · 2017
Closest in time.
“Optimizing expected word error rate via sampling for speech recognition,”
M. Shannon, · 2017
Closest in time.
“An actor-critic algorithm for structured prediction,”
D. Bahadanau, P. Brakel, R. Lowe, J. Pineau, K. Xu, A. Goyal, A. Courville, and Y. Bengio, · 2017
Closest in time.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, L. Jones, J. Uszkoreit, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, · 2017
Closest in time.
“An analysis of “attention” in sequence-to-sequence models,”
R. Prabhavalkar, T. N. Sainath, B. Li, K. Rao, and N. Jaitly, · 2017
Closest in time.
“Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home,”
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. N. Sainath, and M. Bacchiani, · 2017
Closest in time.