Fetching the paper…
Reading the bibliography…
Recent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or character sequences.
F. Jelinek,
1997
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
S. Fernandez, A. Graves, and J. Schmidhuber, “Sequence labelling in structured domains with hierarchical recurrent neural networks,” in
2007
Earlier work this paper cites.
S. F. Chen, “Shrinking exponential language models,” in
2009
Earlier work this paper cites.
T. Mikolov, A. Deoras, S. Kombrink, L. Burget, and J. Černockỳ, “Empirical evaluation and combination of advanced language modeling techniques,” in
2011
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
R. Collobert, K. Kavukcuoglu, and C. Farabet, “Torch7: A Matlab-like Environment for Machine Learning,” in
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson, “One billion word benchmark for measuring progress in statistical language modeling,” in
2014
Cited alongside, same era.
2014
Cited alongside, same era.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks.” in
2014
Cited alongside, same era.
2014
Cited alongside, same era.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global Vectors for Word Representation,” in
2014
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in
2016
Later among the works it cites.
Y. Miao, M. Gowayyed, X. Na, T. Ko, F. Metze, and A. Waibel, “An empirical exploration of CTC acoustic models,” in
2016
Later among the works it cites.
2016
Later among the works it cites.
K. Audhkhasi, A. Sethy, and B. Ramabhadran, “Semantic word embedding neural network language models for automatic speech recognition,” in
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2014
Cited alongside, same era.
A. L. Maas, Z. Xie, D. Jurafsky, and A. Y. Ng, “Lexicon-free conversational speech recognition with neural networks,” in
2015
Cited alongside, same era.
Y. Miao, M. Gowayyed, and F. Metze, “EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,” in
2015
Cited alongside, same era.
H. Sak, A. Senior, K. Rao, and F. Beaufays, “Fast and accurate recurrent neural network acoustic models for speech recognition,” in
2015
Cited alongside, same era.
“Warp CTC,” https://github.com/baidu-research/warp-ctc
Cited in the paper.
2017
Closest in time.
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, “The Microsoft 2016 conversational speech recognition system,” in
2017
Closest in time.
G. Zweig, C. Yu, J. Droppo, and A. Stolcke, “Advances in all-neural speech recognition,” in
2017
Closest in time.
K. Rao and H. Sak, “Multi-accent speech recognition with hierarchical grapheme based models,” in
2017
Closest in time.