Fetching the paper…
Reading the bibliography…
In the recent literature, "end-to-end" speech systems often refer to letter-based acoustic models trained in a sequence-to-sequence manner, either via a recurrent model or via a structured output learning approach (such as CTC).
Maximum mutual information estimation of hidden Markov model parameters for speech recognition
L. R. Bahl, P. F. Brown, P. V. de Souza, and R. L. Mercer · 1986
Earlier work this paper cites.
Modular construction of time-delay neural networks for speech recognition
A. Waibel · 1989
Earlier work this paper cites.
Une Approche théorique de l’Apprentissage Connexionniste: Applications à la Reconnaissance de la Parole
L. Bottou · 1991
Earlier work this paper cites.
The HTK tied-state continuous speech recogniser
P. C. Woodland and S. J. Young · 1993
Earlier work this paper cites.
Improvements in beam search
V. Steinbiss, B.-H. Tran, and H. Ney · 1994
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Y. LeCun, Y. Bengio, et al · 1995
Earlier work this paper cites.
An rnn based speech recognition system with discriminative training
T. Lee, P. Ching, and L.-W. Chan · 1995
Earlier work this paper cites.
Mean and variance adaptation within the MLLR framework
M. J. F. Gales and P. C. Woodland · 1996
Earlier work this paper cites.
Hypothesis spaces for minimum Bayes risk training in large vocabulary speech recognition
M. Gibson and T. Hain · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
G. Hinton, L. Deng, D. Yu, G. Dahl, A. rahman Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury · 2012
Earlier work this paper cites.
Acoustic modeling using deep belief networks
A.-R. Mohamed, G. E. Dahl, and G. Hinton · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
A. Graves, A.-R. Mohamed, and G. Hinton · 2013
Earlier work this paper cites.
Scalable modified kneser-ney language model estimation
K. Heafield, I. Pouzyrevsky, J. H. Clark, and P. Koehn · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
Speaker adaptation of neural network acoustic models using I-Vectors
G. Saon, H. Soltau, D. Nahamoo, and M. Picheny · 2013
Cited alongside, same era.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Cited alongside, same era.
GMM-free DNN training
A. Senior, G. Heigold, M. Bacchiani, and H. Liao · 2014
Cited alongside, same era.
Joint training of convolutional and non-convolutional neural networks
H. Soltau, G. Saon, and T. N. Sainath · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Fast adaptation of deep neural network based on discriminant codes for speech recognition
S. Xue, O. Abdel-Hamid, H. Jiang, L. Dai, and Q. Liu · 2014
Towards better decoding and language model integration in sequence to sequence models
J. Chorowski and N. Jaitly · 2016
Later among the works it cites.
Wav2letter: an end-to-end convnet-based speech recognition system
R. Collobert, C. Puhrsch, and G. Synnaeve · 2016
Later among the works it cites.
Purely sequence-trained neural networks for ASR based on lattice-free MMI
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. P. Kingma · 2016
Later among the works it cites.
Very deep multilingual convolutional neural networks for LVCSR
T. Sercu, C. Puhrsch, B. Kingsbury, and Y. LeCun · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improving deep neural network acoustic models using generalized maxout networks
X. Zhang, J. Trmal, D. Povey, and S. Khudanpur · 2014
Cited alongside, same era.
Audio augmentation for speech recognition
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur · 2015
Cited alongside, same era.
Eesen: End-to-end speech recognition using deep RNN models and WFST-based decoding
Y. Miao, M. Gowayyed, and F. Metze · 2015
Cited alongside, same era.
Librispeech: an ASR corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Cited alongside, same era.
Acoustic modelling with cd-ctc-smbr lstm rnns
H. Sak, F. de Chaumont Quitry, T. Sainath, K. Rao, et al · 2015
Cited alongside, same era.
The IBM 2015 english conversational telephone speech recognition system
G. Saon, H.-K. J. Kuo, S. Rennie, and M. Picheny · 2015
Cited alongside, same era.
Language modeling with gated convolutional nets
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier · 2017
Closest in time.
Convolutional sequence to sequence learning
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Closest in time.
Multi-level language modeling and decoding for open vocabulary end-to-end speech recognition
T. Hori, S. Watanabe, and J. R. Hershey · 2017
Closest in time.
Gram-ctc: Automatic unit selection and target decomposition for sequence labelling
H. Liu, Z. Zhu, X. Li, and S. Satheesh · 2017
Closest in time.
English conversational telephone speech recognition by humans and machines
G. Saon, G. Kurata, T. Sercu, K. Audhkhasi, S. Thomas, D. Dimitriadis, X. Cui, B. Ramabhadran, M. Picheny, L.-L. Lim, et al · 2017
Closest in time.
End-to-end neural segmental models for speech recognition
H. Tang, L. Lu, L. Kong, K. Gimpel, K. Livescu, C. Dyer, N. A. Smith, and S. Renals · 2017
Closest in time.
The Microsoft 2016 conversational speech recognition system
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig · 2017
Closest in time.
Improved training of end-to-end attention models for speech recognition
A. Zeyer, K. Irie, R. Schlüter, and H. Ney · 2018
Closest in time.
Improving end-to-end speech recognition with policy learning
Y. Zhou, C. Xiong, and R. Socher · 2018
Closest in time.