Fetching the paper…
Reading the bibliography…
In conventional speech recognition, phoneme-based models outperform grapheme-based models for non-phonetic languages such as English.
1902
Earlier work this paper cites.
P. F. Brown, P. V. Desouza, R. L. Mercer, V. J. D. Pietra, and J. C. Lai, “Class-based n-gram models of natural language,”
1992
Earlier work this paper cites.
H. A. Bourlard and N. Morgan,
1993
Earlier work this paper cites.
M. Schuster and K. K. Paliwal, “Bidirectional recurrent neural networks,”
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
S. Kanthak and H. Ney, “Context-dependent acoustic modeling using graphemes for large vocabulary speech recognition,” in
2002
Earlier work this paper cites.
M. Killer, S. Stüker, and T. Schultz, “Grapheme based speech recognition,” in
2003
Earlier work this paper cites.
M. Mohri, F. Pereira, and M. Riley, “Speech recognition with weighted finite-state transducers,” in
2008
Earlier work this paper cites.
Y. Sung, T. Hughes, F. Beaufays, and B. Strope, “Revisiting graphemes with increasing amounts of data,” in
2009
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and korean voice search,” in
2012
Earlier work this paper cites.
M. Sundermeyer, R. Schlüter, and H. Ney, “LSTM neural networks for language modeling.” in
2012
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” in
2012
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in
2014
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Cited alongside, same era.
2016
Cited alongside, same era.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in
2016
Cited alongside, same era.
C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, E. Gonina, N. Jaitly, B. Li, J. Chorowski, and M. Bacchiani, “State-of-the-art speech recognition with sequence-to-sequence models,” in
2018
Later among the works it cites.
T. N. Sainath, R. Prabhavalkar, S. Kumar, S. Lee, A. Kannan, D. Rybach, V. Schogol, P. Nguyen, B. Li, and Y. Wu, “No need for a lexicon? evaluating the value of the pronunciation lexica in end-to-end models,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Zeyer, K. Irie, R. Schlüter, and H. Ney, “Improved training of end-to-end attention models for speech recognition,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
R. Prabhavalkar, K. Rao, T. Sainath, B. Li, L. Johnson, and N. Jaitly, “A comparison of sequence-to-sequence models for speech recognition,” in
2017
Cited alongside, same era.
S. Kim, T. Hori, and S. Watanabe, “Joint ctc-attention based end-to-end speech recognition using multi-task learning,” in
2017
Cited alongside, same era.
E. Battenberg, J. Chen, R. Child, A. Coates, Y. G. Y. Li, H. Liu, S. Satheesh, A. Sriram, and Z. Zhu, “Exploring neural transducers for end-to-end speech recognition,” in
2017
Cited alongside, same era.
R. J. Weiss, J. Chorowski, N. Jaitly, Y. Wu, and Z. Chen, “Sequence-to-sequence models can directly translate foreign speech,” in
2017
Cited alongside, same era.
Y. Zhang, W. Chan, and N. Jaitly, “Very deep convolutional networks for end-to-end speech recognition,” in
2017
Cited alongside, same era.
J. Chorowski and N. Jaitly, “Towards better decoding and language model integration in sequence to sequence models,” in
2017
Cited alongside, same era.
K. J. Han, A. Chandrashekaran, J. Kim, and I. Lane, “The CAPIO 2017 conversational speech recognition system,”
2018
Later among the works it cites.
R. Prabhavalkar, T. N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C. Chiu, and A. Kannan, “Minimum word error rate training for attention-based sequence-to-sequence models,” in
2018
Later among the works it cites.
S. Toshniwal, A. Kannan, C.-C. Chiu, Y. Wu, T. N. Sainath, and K. Livescu, “A comparison of techniques for language model integration in encoder-decoder speech recognition,” in
2018
Later among the works it cites.
2019
Closest in time.
A. Bruguier, R. Prabhavalkar, G. Pundak, and T. N. Sainath, “Phoebe: Pronunciation-aware contextualization for end-to-end speech recognition,” in
2019
Closest in time.
J. Shen, P. Nguyen, Y. Wu, Z. Chen
2019
Closest in time.
2019
Closest in time.
S. Sabour, W. Chan, and M. Norouzi, “Optimal completion distillation for sequence learning,” in
2019
Closest in time.