Fetching the paper…
Reading the bibliography…
Following the rationale of end-to-end modeling, CTC, RNN-T or encoder-decoder-attention models for automatic speech recognition (ASR) use graphemes or grapheme-based subword units based on e.g.
Classification and regression trees
L. Breiman, J. Friedman, C. Stone, and R. Olshen, · 1984
Earlier work this paper cites.
“The general use of tying in phoneme-based HMM speech recognisers,”
S. J. Young, · 1992
Earlier work this paper cites.
“Switchboard: Telephone speech corpus for research and development,”
J. J. Godfrey, E. C. Holliman, and J. McDaniel, · 1992
Earlier work this paper cites.
“A new algorithm for data compression,”
P. Gage, · 1994
Earlier work this paper cites.
Connectionist speech recognition: a hybrid approach
H. Bourlard and N. Morgan, · 1994
Earlier work this paper cites.
“An application of recurrent nets to phone probability estimation,”
A. J. Robinson, · 1994
Earlier work this paper cites.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Context-dependent acoustic modeling using graphemes for large vocabulary speech recognition,”
S. Kanthak and H. Ney, · 2002
Earlier work this paper cites.
“Grapheme based speech recognition,”
M. Killer, S. Stuker, and T. Schultz, · 2003
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Revisiting graphemes with increasing amounts of data,”
Y.-H. Sung, T. Hughes, F. Beaufays, and B. Strope, · 2009
Earlier work this paper cites.
“RASR - the RWTH Aachen University open source speech recognition toolkit,”
D. Rybach, S. Hahn, P. Lehnen, D. Nolden, M. Sundermeyer, Z. Tüske, S. Wiesler, R. Schlüter, and H. Ney, · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,” Preprint arXiv:1211.3711, 2012
A. Graves, · 2012
Earlier work this paper cites.
“Japanese and korean voice search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
M.-T. Luong, H. Pham, and C. D. Manning, · 2015
Earlier work this paper cites.
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,”
Y. Miao, M. Gowayyed, and F. Metze, · 2015
Earlier work this paper cites.
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, et al., · 2015
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,” Preprint arXiv:1508.07909, 2015
R. Sennrich, B. Haddow, and A. Birch, · 2015
Earlier work this paper cites.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
H. Sak, A. Senior, K. Rao, and F. Beaufays, · 2015
Earlier work this paper cites.
“Attention-based models for speech recognition,”
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Cited alongside, same era.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al., · 2016
Cited alongside, same era.
“Latent sequence decompositions,” Preprint arXiv:1610.03035, 2016
W. Chan, Y. Zhang, Q. Le, and N. Jaitly, · 2016
Cited alongside, same era.
“An empirical exploration of CTC acoustic models,”
Y. Miao, M. Gowayyed, X. Na, T. Ko, F. Metze, and A. Waibel, · 2016
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Later among the works it cites.
“A comparison of Transformer and LSTM encoder decoder models for ASR,”
A. Zeyer, P. Bahar, K. Irie, R. Schlüter, and H. Ney, · 2019
Later among the works it cites.
“Improving end-to-end speech recognition with pronunciation-assisted sub-word modeling,”
H. Xu, S. Ding, and S. Watanabe, · 2019
Later among the works it cites.
“Advancing acoustic-to-word CTC model with attention and mixed-units,”
A. Das, J. Li, G. Ye, R. Zhao, and Y. Gong, · 2019
Later among the works it cites.
“Phoebe: Pronunciation-aware contextualization for end-to-end speech recognition,”
A. Bruguier, R. Prabhavalkar, G. Pundak, and T. N. Sainath, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Wav2letter: an end-to-end convnet-based speech recognition system,” Preprint arXiv:1609.03193, 2016
R. Collobert, C. Puhrsch, and G. Synnaeve, · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2016
Cited alongside, same era.
“Advances in all-neural speech recognition,”
G. Zweig, C. Yu, J. Droppo, and A. Stolcke, · 2017
Cited alongside, same era.
“Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition,”
H. Soltau, H. Liao, and H. Sak, · 2017
Cited alongside, same era.
“Direct acoustics-to-word models for english conversational speech recognition,”
K. Audhkhasi, B. Ramabhadran, G. Saon, M. Picheny, and D. Nahamoo, · 2017
Cited alongside, same era.
“A comprehensive study of deep bidirectional LSTM RNNs for acoustic modeling in speech recognition,”
A. Zeyer, P. Doetsch, P. Voigtlaender, R. Schlüter, and H. Ney, · 2017
Cited alongside, same era.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, · 2017
Cited alongside, same era.
D. Le, X. Zhang, W. Zheng, C. Fügen, G. Zweig, and M. L. Seltzer, · 2019
Later among the works it cites.
“On the choice of modeling unit for sequence-to-sequence speech recognition,”
K. Irie, R. Prabhavalkar, A. Kannan, A. Bruguier, D. Rybach, and P. Nguyen, · 2019
Later among the works it cites.
K. Hu, A. Bruguier, T. N. Sainath, R. Prabhavalkar, and G. Pundak, · 2019
Later among the works it cites.
T. Nguyen, S. Stueker, J. Niehues, and A. Waibel, · 2019
Later among the works it cites.
“A comparative study on Transformer vs RNN in speech applications,”
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. Soplin, R. Yamamoto, X. Wang, et al., · 2019
Later among the works it cites.
Z. Tüske, G. Saon, K. Audhkhasi, and B. Kingsbury, · 2020
Closest in time.
“A new training pipeline for an improved neural transducer,”
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, · 2020
Closest in time.
“Byte pair encoding is suboptimal for language model pretraining,” Preprint arXiv:2004.03720, 2020
K. Bostrom and G. Durrett, · 2020
Closest in time.
“Learning a subword inventory jointly with end-to-end automatic speech recognition,”
J. Drexler and J. Glass, · 2020
Closest in time.
“Layer-normalized LSTM for hybrid-HMM and end-to-end ASR,”
M. Zeineldeen, A. Zeyer, R. Schlüter, and H. Ney, · 2020
Closest in time.
W. Wang, Y. Zhou, C. Xiong, and R. Socher, · 2020
Closest in time.
“Joint phoneme-grapheme model for end-to-end speech recognition,”
Y. Kubo and M. Bacchiani, · 2020
Closest in time.
“Hybrid autoregressive transducer (HAT),”
E. Variani, D. Rybach, C. Allauzen, and M. Riley, · 2020
Closest in time.
“Phoneme based neural transducer for large vocabulary speech recognition,”
W. Zhou, S. Berger, R. Schlüter, and H. Ney, · 2021
Closest in time.