Fetching the paper…
Reading the bibliography…
For decades, context-dependent phonemes have been the dominant sub-word unit for conventional acoustic modeling systems.
“Long Short-Term Memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Multilingual acoustic modeling using graphemes,”
S. Kanthak and H. Ney, · 2003
Earlier work this paper cites.
“Revisiting graphemes with increasing amounts of data,”
Y. Sung, T. Hughes, F. Beaufays, and B. Strope, · 2008
Earlier work this paper cites.
“Speech recognition with weighted finite-state transducers,”
Mehryar Mohri, Fernando Pereira, and Michael Riley, · 2008
Earlier work this paper cites.
“From speech to letters using a novel neural network architecture for grapheme based ASR,”
F. Eyben, M. Wollmer, B. Schuller, and A. Graves, · 2009
Earlier work this paper cites.
“Large Scale Distributed Deep Networks,”
J. Dean, G.S. Corrado, R. Monga, K. Chen, M. Devin, Q.V. Le, M.Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A.Y. Ng, · 2012
Earlier work this paper cites.
“Learning lexicons from speech using a pronunciation mixture model,”
I. McGraw, I. Badr, and J. R. Glass, · 2013
Earlier work this paper cites.
“Acoustic data-driven pronunciation lexicon for large vocabulary speech recognition,”
L. Lu, A. Ghoshal, and S. Renals, · 2013
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2014
Earlier work this paper cites.
“Attention-Based Models for Speech Recognition,”
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Cited alongside, same era.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2015
Cited alongside, same era.
“Deep speech 2: End-to-end speech recognition in english and mandarin,”
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, E. Elsen, J. Engel, L. Fan, C. Fougner, T. Han, A. Hannun, B. Jun, P. LeGresley, L. Lin, S. Narang, A. Ng, S. Ozair, R. Prenger, J. Raiman, S. Satheesh, D. Seetapun, S. Sengupta, Y. Wang, Z. Wang, C. Wang, B. Xiao, D. Yogatama, J. Zhan, and Z. Zhu, · 2015
Cited alongside, same era.
“Fast and Accurate Recurrent Neural Network Acoustic Models for Speech Recognition,”
H. Sak, A. Senior, K. Rao, and F. Beaufays, · 2015
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
“A Comparison of Sequence-to-sequence Models for Speech Recognition,”
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, · 2017
Closest in time.
“Streaming Small-footprint Keyword Spotting Using Sequence-to-Sequence Models,”
Y. He, R. Prabhavalkar, K. Rao, W. Li, A. Bakhtin, and I. McGraw, · 2017
Closest in time.
“An Analysis of ”Attention” in Sequence-to-Sequence Models,” in Proc. Interspeech,”
R. Prabhavalkar, T. N. Sainath, B. Li, K. Rao, and N. Jaitly, · 2017
Closest in time.
“Towards Better Decoding and Language Model Integration in Sequence to Sequence Models,”
J. Chorowski and N. Jaitly, · 2017
Closest in time.
“Generated of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home,”
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. N. Sainath, and M. Bacchiani, · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems,” Available online: http://download.tensorflow.org/paper/whitepaper2015.pdf, 2015
M. Abadi et al., · 2015
Cited alongside, same era.
“Lower Frame Rate Neural Network Acoustic Models,”
G. Pundak and T. N. Sainath, · 2016
Cited alongside, same era.
“Multi-accent Speech Recognition with Hierarchical Grapheme Based Models,”
K. Rao and H. Sak, · 2017
Cited alongside, same era.
“Exploring Architectures, Data and Units for Streaming End-to-End Speech Recognition with RNN-Transducer,”
K. Rao, R. Prabhavalkar, and H. Sak, · 2017
Cited alongside, same era.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
A. Kannan, Y. Wu, P. Nguyen, T. N. Sainath, Z. Chen, and R. Prabhavalkar, · 2018
Closest in time.
“Multi-dialect speech recognition with a single sequence-to-sequence model,”
B. Li, T. N. Sainath, K. Sim, M. Bacchiani, E. Weinstein, P. Nguyen, Z. Chen, Y. Wu, and K. Rao, · 2018
Closest in time.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C. Chen, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, N. Jaitly, B. Li, and J. Chorowski, · 2018
Closest in time.