Fetching the paper…
Reading the bibliography…
Sequence-to-sequence attention-based models on subword units allow simple open-vocabulary end-to-end speech recognition.
H. Bourlard and N. Morgan,
1994
Earlier work this paper cites.
A. J. Robinson, “An application of recurrent nets to phone probability estimation,”
1994
Earlier work this paper cites.
S. J. Young, J. J. Odell, and P. C. Woodland, “Tree-based state tying for high accuracy acoustic modelling,” in
1994
Earlier work this paper cites.
R. Kneser and H. Ney, “Improved backing-off for m-gram language modeling,” in
1995
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
S. Kanthak and H. Ney, “Context-dependent acoustic modeling using graphemes for large vocabulary speech recognition,” in
2002
Earlier work this paper cites.
A. Stolcke, “SRILM-an extensible language modeling toolkit.” in
2002
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
R. Schlüter, I. Bezrukov, H. Wagner, and H. Ney, “Gammatone features and feature combination for large vocabulary speech recognition,” in
2007
Earlier work this paper cites.
M. Sundermeyer, R. Schlüter, and H. Ney, “LSTM neural networks for language modeling.” in
2012
Earlier work this paper cites.
A. Senior, G. Heigold, M. Bacchiani, and H. Liao, “GMM-free DNN acoustic model training,” in
2014
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Ba, V. Mnih, and K. Kavukcuoglu, “Multiple object recognition with visual attention,”
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,”
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
S. Wiesler, A. Richard, P. Golik, R. Schlüter, and H. Ney, “RASR/NN: The RWTH neural network toolkit for speech recognition,” in
2014
Earlier work this paper cites.
H. Sak, A. Senior, K. Rao, and F. Beaufays, “Fast and accurate recurrent neural network acoustic models for speech recognition,” in
2015
Earlier work this paper cites.
Y. Miao, M. Gowayyed, and F. Metze, “Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
TensorFlow Development Team, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available:
2015
Cited alongside, same era.
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in
2015
Cited alongside, same era.
A. L. Maas, Z. Xie, D. Jurafsky, and A. Y. Ng, “Lexicon-free conversational speech recognition with neural networks,” in
2015
Cited alongside, same era.
H. Soltau, H. Liao, and H. Sak, “Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition,” in
2017
Later among the works it cites.
K. Audhkhasi, B. Ramabhadran, G. Saon, M. Picheny, and D. Nahamoo, “Direct acoustics-to-word models for english conversational speech recognition,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
E. Battenberg, J. Chen, R. Child, A. Coates, Y. Gaur, Y. Li, H. Liu, S. Satheesh, A. Sriram, and Z. Zhu, “Exploring neural transducers for end-to-end speech recognition,” in
2017
Later among the works it cites.
S. Toshniwal, H. Tang, L. Lu, and K. Livescu, “Multitask learning with low-level auxiliary tasks for encoder-decoder based speech recognition,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Cited alongside, same era.
P. Doetsch, A. Zeyer, and H. Ney, “Bidirectional decoder networks for attention-based end-to-end offline handwriting recognition,” in
2016
Cited alongside, same era.
N. Jaitly, Q. V. Le, O. Vinyals, I. Sutskever, D. Sussillo, and S. Bengio, “An online sequence-to-sequence model using partial conditioning,” in
2016
Cited alongside, same era.
R. Aharoni and Y. Goldberg, “Morphological inflection generation with hard monotonic attention,”
2016
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
C.-C. Chiu and C. Raffel, “Monotonic chunkwise attention,”
2017
Later among the works it cites.
A. Tjandra, S. Sakti, and S. Nakamura, “Local monotonic attention mechanism for end-to-end speech and language processing,” in
2017
Later among the works it cites.
R. Prabhavalkar, T. N. Sainath, B. Li, K. Rao, and N. Jaitly, “An analysis of “attention” in sequence-to-sequence models,”,” in
2017
Later among the works it cites.
J. Hou, S. Zhang, and L. Dai, “Gaussian prediction based attention for online end-to-end speech recognition,” in
2017
Later among the works it cites.
P. Doetsch, M. Hannemann, R. Schlueter, and H. Ney, “Inverted alignments for end-to-end automatic speech recognition,”
2017
Later among the works it cites.
K. Rao, H. Sak, and R. Prabhavalkar, “Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, “A comparison of sequence-to-sequence models for speech recognition,” in
2017
Later among the works it cites.
H. Sak, M. Shannon, K. Rao, and F. Beaufays, “Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping,” in
2017
Later among the works it cites.
P. Doetsch, A. Zeyer, P. Voigtlaender, I. Kulikov, R. Schlüter, and H. Ney, “RETURNN: the RWTH extensible training framework for universal recurrent neural networks,” in
2017
Later among the works it cites.
P. Bahar, J. Rosendahl, N. Rossenbach, and H. Ney, “The RWTH Aachen machine translation systems for IWSLT 2017,” in
2017
Later among the works it cites.
librosa Development Team, “librosa 0.5.0,” Feb. 2017. [Online]. Available:
2017
Later among the works it cites.
T. Hori, S. Watanabe, Y. Zhang, and W. Chan, “Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
D. Nolden, “Progress in decoding for large vocabulary continuous speech recognition,” Ph.D. dissertation, RWTH Aachen University, Computer Science Department, RWTH Aachen University, Aachen, Germany, Apr. 2017
2017
Later among the works it cites.
“RETURNN as a generic flexible neural toolkit with application to translation and speech recognition,”
2018
Closest in time.