Fetching the paper…
Reading the bibliography…
In recent years, there has been a great deal of research in developing end-to-end speech recognition models, which enable simplifying the traditional pipeline and achieving promising results.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Direct acoustics-to-word models for english conversational speech recognition
K. Audhkhasi, B. Ramabhadran, G. Saon, and D. Picheny, M. Nahamoo · 2008
Earlier work this paper cites.
Kenlm: faster and smaller language model queries
K. Heafield · 2011
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
A. Graves · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2014
Earlier work this paper cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
Learning small-size dnn with output-distribution-based criteria
J. Li, R. Zhao, T. J. Huang, and Y. Gong · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Earlier work this paper cites.
Attention-based models for speech recognition
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Fitnets: hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2015
Earlier work this paper cites.
Acoustic modelling with cd-ctc-smbr lstm rnns
A. Senior, H. Sak, F. C. Quitry, T. Sainath, K. Rao, et al · 2015
Earlier work this paper cites.
Learning acoustic frame labeling for speech recognition with recurrent neural networks
H. Sak, A. Senior, K. Rao, O. Irsoy, A. Graves, F. Beaufays, and J. Schalkwyk · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Adam: a method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep speech 2: end-to-end speech recognition in english and mandarin
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, et al · 2016
Cited alongside, same era.
Wav2letter: an end-to-end convnet-based speech recognition system
R. Collobert, C. Puhrsch, and G. Synnaeve · 2016
Cited alongside, same era.
Distilling knowledge from ensembles of neural networks for speech recognition
Y. Chebotar and A. Waters · 2016
Cited alongside, same era.
Blending lstms into cnns
K. J. Geras et al · 2016
Cited alongside, same era.
Sequence student-teacher training of deep neural networks
J. H. M. Wong and M. J. F. Gales · 2016
Cited alongside, same era.
Listen, attend and spell: a neural network for large vocabulary conversational speech recognition
Knowledge transfer with jacobian matching
S. Srinivas and F. Fleuret · 2018
Later among the works it cites.
An investigation of a knowledge distillation method for ctc acoustic models
R. Takashima, S. Li, and H. Kawai · 2018
Later among the works it cites.
Improved knowledge distillation from bi-directional to uni-directional lstm ctc for end-to-end speech recognition
G. Kurata and K. Audhkhasi · 2018
Later among the works it cites.
Aishell-2: transforming mandarin asr research into industrial scale
J. Du, X. Nai, X. Liu, and H. Bu · 2018
Later among the works it cites.
Mixed-precision training for nlp and speech recognition with openseq2seq
O. Kuchaiev, B. Ginsburg, I. Gitman, V. Lavrukhin, J. Li, H. Nguyen, C. Case, and P. Micikevicius · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
R. Sennrich, B. Haddow, and A. Birch · 2016
Cited alongside, same era.
Attention is all you need
A. Vaswani et al · 2017
Cited alongside, same era.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
S. Zagoruyko and N. Komodakis · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J. Bae, and J. Kim · 2017
Cited alongside, same era.
Student-teacher network learning with enhanced features
S. Watanabe, T. Hori, J. L. Roux, and J. R. Hershey · 2017
Cited alongside, same era.
Knowledge distillation for small-footprint highway networks
L. Lu, M. Guo, and S. Renals · 2017
Cited alongside, same era.
Espnet: end-to-end speech processing toolkit
S. Watanabe et al · 2018
Later among the works it cites.
Jasper: An end-to-end convolutional neural acoustic model
J. Li, V. Lavrukhin, B. Ginsburg, R. Leary, O. Kuchaiev, J. M. Cohen, H. Nguyen, and R. T. Gadde · 2019
Later among the works it cites.
Quartznet: deep automatic speech recognition with 1d time-channel separable convolutions
S. Kriman, K. Beliaev, B. Ginsburg, J. Huang, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, and Y. Zhang · 2019
Later among the works it cites.
General sequence teacher–student learning
J. H. M. Wong, M. J. F. Gales, and Y. Wang · 2019
Later among the works it cites.
Investigation of sequence-level knowledge distillation methods for ctc acoustic models
R. Takashima, S. Li, and H. Kawai · 2019
Later among the works it cites.
Guiding ctc posterior spike timings for improved posterior fusion and knowledge distillation
G. Kurata and K. Audhkhasi · 2019
Later among the works it cites.
Nemo: a toolkit for building ai applications using neural modules
O. Kuchaiev et al · 2019
Later among the works it cites.
Stochastic gradient methods with layer-wise adaptive moments for training of deep networks
B. Ginsburg, P. Castonguay, O. Hrinchuk, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, H. Nguyen, and J. M. Cohen · 2019
Later among the works it cites.
Specaugment: a simple data augmentation method for automatic speech recognition
D. S. Park, W. Chan, C. Zhang, Y. Chiu, B. Zoph, E. D. Cubuk, and Le. Q. V · 2019
Later among the works it cites.
Improved multi-stage training of online attention-based encoder-decoder models
A. Garg, O. Gowda, A. Kumar, K. Kim, M. Kumar, and Kim C · 2019
Later among the works it cites.