Fetching the paper…
Reading the bibliography…
In this paper, we report state-of-the-art results on LibriSpeech among end-to-end speech recognition models without any external training data.
A. Waibel, T. Hanazawa, G. Hinton, K. Shirano, and K. Lang, “A time-delay neural network architecture for isolated word recognition,”
1989
Earlier work this paper cites.
Y. Bengio, R. De Mori, G. Flammia, and R. Kompe, “Global optimization of a neural network-hidden markov model hybrid,”
1992
Earlier work this paper cites.
A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional lstm and other neural network architectures,”
2005
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
K. Heafield, “Kenlm: Faster and smaller language model queries,” in
2011
Earlier work this paper cites.
G. Hinton
2012
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
K. Tom, P. Vijayaditya, P. Daniel, and K. Sanjeev, “Audio augmentation for speech recognition,”
2015
Earlier work this paper cites.
Y. Zhang
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,”
2016
Cited alongside, same era.
A. van den Oord
2016
Cited alongside, same era.
C. Laurent, G. Pereyra, P. Brakel, Y. Zhang, and Y. Bengio, “Batch normalized recurrent neural networks,” in
2016
Cited alongside, same era.
D. Amodei
H. Hadian, H. Sameti, D. Povey, and S. Khudanpur, “End-to-end speech recognition using lattice-free mmi,” in
2018
Later among the works it cites.
J. Tang, Y. Song, L. Dai, and I. McLoughlin, “Acoustic modeling with densely connected residual network for multichannel speech recognition,” in
2018
Later among the works it cites.
A. Zeyer, K. Irie, R. Schlüter, and H. Ney, “Improved training of end-to-end attention models for speech recognition,” in
2018
Later among the works it cites.
D. Povey
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2017
Cited alongside, same era.
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” in
2017
Cited alongside, same era.
Z. Zhang, L. Ma, Z. Li, and C. Wu, “Normalized direction-preserving adam,”
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Zhang, W. Chan, and N. Jaitly, “Very deep convolutional networks for end-to-end speech recognition,” in
2017
Cited alongside, same era.
E. Battenberg
2017
Cited alongside, same era.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
D. S. Park
2019
Closest in time.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in
2019
Closest in time.
B. Ginsburg
2019
Closest in time.