Fetching the paper…
Reading the bibliography…
Phones and their context-dependent variants have been the standard modeling units for conventional speech recognition systems, while characters and subwords have demonstrated their effectiveness for end-to-end recognition systems.
S. J. Young, J. J. Odell, and P. C. Woodland, “Tree-based state tying for high accuracy acoustic modelling,” in
1994
Earlier work this paper cites.
S. J. Young, D. Kernshaw, J. Odell, D. Ollason, V. Valtchev, and P. Woodland, “The HTK book version 2.2,” Tech. Rep., 1999
1999
Earlier work this paper cites.
M. Bisani and H. Ney, “Joint-sequence models for grapheme-to-phoneme conversion,”
2008
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi speech recognition toolkit,” in
2011
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and korean voice search,” in
2012
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in
2014
Earlier work this paper cites.
Y. Miao, M. Gowayyed, and F. Metze, “EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,” in
2015
Earlier work this paper cites.
H. Sak, A. Senior, K. Rao, and F. Beaufays, “Fast and accurate recurrent neural network acoustic models for speech recognition,” in
2015
Earlier work this paper cites.
K. Rao, F. Peng, H. Sak, and F. Beaufays, “Grapheme-to-phoneme conversion using long short-term memory recurrent neural networks,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in
2016
Earlier work this paper cites.
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, E. Elsen, J. Engel, L. Fan, C. Fougner, A. Hannun, B. Jun, T. Han, P. LeGresley, X. Li, L. Lin, S. Narang, A. Ng, S. Ozair, R. Prenger, S. Qian, J. Raiman, S. Satheesh, D. Seetapun, S. Sengupta, C. Wang, Y. Wang, Z. Wang, B. Xiao, Y. Xie, D. Yogatama, J. Zhan, and Z. Zhu, “Deep speech 2: End-to-end speech recognition in english and mandarin,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Cited alongside, same era.
Y. Miao, M. Gowayyed, X. Na, T. Ko, F. Metze, and A. Waibel, “An empirical exploration of CTC acoustic models,” in
2016
Cited alongside, same era.
S. Toshniwal and K. Livescu, “Jointly learning to align and convert graphemes to phonemes with neural attention models,” in
2016
Cited alongside, same era.
G. Zweig, C. Yu, J. Droppo, and A. Stolcke, “Advances in all-neural speech recognition,” in
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Kim, J. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,”
2017
Cited alongside, same era.
T. Kudo, “Subword regularization: Improving neural network translation models with multiple subword candidates,” in
2018
Later among the works it cites.
C. Yu, C. Zhang, C. Weng, J. Cui, and D. Yu, “A multistage training framework for acoustic-to-word model,” in
2018
Later among the works it cites.
J. Cui, C. Weng, G. Wang, J. Wang, P. Wang, C. Yu, D. Su, and D. Yu, “Improving attention-based end-to-end ASR systems with sequence-based loss functions,” in
2018
Later among the works it cites.
A. Zeyer, K. Irie, R. Schluter, and H. Ney, “A comprehensive analysis on attention models,” in
2018
Later among the works it cites.
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, Q. Liang, D. Bhatia, Y. Shangguan, B. Li, G. Pundak, K. C. Sim, T. Bagby, S. Chang, K. Rao, and A. Gruenstein, “Streaming end-to-end speech recognition for mobile devices,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Hori, S. Watanabe, and J. R. Hershey, “Multilevel language modeling and decoding for open vocabulary end-to-end speech recognition,” in
2017
Cited alongside, same era.
H. Sak and K. Rao, “Multi-accent speech recognition with hierarchical grapheme based models,” in
2017
Cited alongside, same era.
S. Toshniwal, H. Tang, L. Lu, and K. Livescu, “Multitask learning with low-level auxiliary tasks for encoder-decoder based speech recognition,” in
2017
Cited alongside, same era.
K. Rao, H. Sak, and R. Prabhavalkar, “Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,” in
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in
2018
Cited alongside, same era.
2019
Later among the works it cites.
Y. Wang, T. Chen, H. Xu, S. Ding, H. Lv, Y. Shao, N. Peng, L. Xie, S. Watanabe, and S. Khudanpur, “Espresso: A fast end-to-end neural speech recognition toolkit,” in
2019
Later among the works it cites.
S. Yolchuyeva, G. Németh, and B. Gyires-Tóth, “Transformer based grapheme-to-phoneme conversion,” in
2019
Later among the works it cites.
H. Xu, S. Ding, and S. Watanabe, “Improving end-to-end speech recognition with pronunciation-assisted sub-word modeling,” in
2019
Later among the works it cites.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang, S. Watanabe, T. Yoshimura, and W. Zhang, “A comparative study on transformer vs rnn in speech applications,” in
2019
Later among the works it cites.
N.-Q. Pham, T.-S. Nguyen, J. Niehues, M. Müller, S. Stüker, and A. Waibel, “Very deep self-attention networks for end-to-end speech recognition,” in
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” in
2019
Later among the works it cites.
M. K. Baskar, L. Burget, S. Watanabe, M. Karafiat, T. Hori, and J. H. Cernocky, “Promising accurate prefix boosting for sequence-to-sequence ASR,” in
2019
Later among the works it cites.
2019
Later among the works it cites.