Fetching the paper…
Reading the bibliography…
The acoustic-to-word model based on the connectionist temporal classification (CTC) criterion was shown as a natural end-to-end (E2E) model directly targeting words as output units.
“Long Short-Term Memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
Modelling out-of-vocabulary words for robust speech recognition
Issam Bazzi, · 2002
Earlier work this paper cites.
“Transcription of out-of-vocabulary words in large vocabulary speech recognition based on phoneme-to-grapheme conversion,”
Bart Decadt, Jacques Duchateau, Walter Daelemans, and Patrick Wambacq, · 2002
Earlier work this paper cites.
“Hybrid language models for out of vocabulary word detection in large vocabulary conversational speech recognition,”
Ali Yazgan and Murat Saraclar, · 2004
Earlier work this paper cites.
“Open vocabulary speech recognition with flat hybrid models.,”
Maximilian Bisani and Hermann Ney, · 2005
Earlier work this paper cites.
“Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Japanese and korean voice search,”
Mike Schuster and Kaisuke Nakajima, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
A. Graves, A. Mohamed, and G. Hinton, · 2013
Earlier work this paper cites.
“Towards End-to-End Speech Recognition with Recurrent Neural Networks,”
A. Graves and N. Jaitley, · 2014
Earlier work this paper cites.
“Deep Speech: Scaling up End-to-End Speech Recognition,”
A. Y. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, and A. Y. Ng, · 2014
Earlier work this paper cites.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling.,”
H. Sak, A. Senior, and F. Beaufays, · 2014
Earlier work this paper cites.
“Learning Acoustic Frame Labeling for Speech Recognition with Recurrent Neural Networks,”
H. Sak, A. Senior, K. Rao, O. Irsoy, A. Graves, F. Beaufays, and J. Schalkwyk, · 2015
Cited alongside, same era.
“Fast and Accurate Recurrent Neural Network Acoustic Models for Speech Recognition,”
H. Sak, A. Senior, K. Rao, and F. Beaufays, · 2015
Cited alongside, same era.
“EESEN: End-to-End Speech Recognition using Deep RNN Models and WFST-based Decoding,”
Y. Miao, M. Gowayyed, and F. Metze, · 2015
Cited alongside, same era.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2015
Cited alongside, same era.
“Maximum a posteriori Based Decoding for CTC Acoustic Models.,”
Naoyuki Kanda, Xugang Lu, and Hisashi Kawai, · 2016
Cited alongside, same era.
“Direct Acoustics-to-Word Models for English Conversational Speech Recognition,”
K. Audhkhasi, B. Ramabhadran, G. Saon, M. Picheny, and D. Nahamoo, · 2017
Later among the works it cites.
“Acoustic-to-Word Model Without OOV,”
J. Li, G. Ye, R. Zhao, J. Droppo, and Y. Gong, · 2017
Later among the works it cites.
“Recent Progresses in Deep Learning Based Acoustic Models,”
D. Yu and J. Li, · 2017
Later among the works it cites.
“CTC training of multi-phone acoustic models for speech recognition,”
Olivier Siohan, · 2017
Later among the works it cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Later among the works it cites.
“Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Soltau, H. Liao, and H. Sak, · 2016
Cited alongside, same era.
“On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition,”
Liang Lu, Xingxing Zhang, and Steve Renais, · 2016
Cited alongside, same era.
“Latent sequence decompositions,”
William Chan, Yu Zhang, Quoc Le, and Navdeep Jaitly, · 2016
Cited alongside, same era.
“Advances in All-Neural Speech Recognition,”
G. Zweig, C. Yu, J. Droppo, and A. Stolcke, · 2017
Cited alongside, same era.
“Gram-CTC: Automatic Unit Selection and Target Decomposition for Sequence Labelling,”
H. Liu, Z. Zhu, X. Li, and S. Satheesh, · 2017
Cited alongside, same era.
Hasim Sak, Matt Shannon, Kanishka Rao, and Françoise Beaufays, · 2017
Later among the works it cites.
“Developing far-field speaker system via teacher-student learning,”
J. Li, R. Zhao, et al., · 2018
Closest in time.
“Building competitive direct acoustics-to-word models for english conversational speech recognition,”
Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, and Michael Picheny, · 2018
Closest in time.
“Advancing connectionist temporal classification with attention modeling,”
A. Das, J. Li, R. Zhao, and Y. Gong, · 2018
Closest in time.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, et al., · 2018
Closest in time.