Fetching the paper…
Reading the bibliography…
While deep learning based end-to-end automatic speech recognition (ASR) systems have greatly simplified modeling pipelines, they suffer from the data sparsity issue.
1904
Earlier work this paper cites.
1905
Earlier work this paper cites.
D. B. Paul and J. M. Baker, “The design for the wall street journal-based CSR corpus,” in
1992
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,”
1997
Earlier work this paper cites.
T. Kemp and A. Waibel, “Unsupervised training of a speech recognizer: recent experiments,” in
1999
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The Kaldi speech recognition toolkit,” in
2011
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” in
2012
Earlier work this paper cites.
Y. Huang, D. Yu, Y. Gong, and C. Liu, “Semi-supervised GMM and DNN acoustic model training with multi-system combination and confidence re-calibration,” in
2013
Earlier work this paper cites.
S. Thomas, M. L. Seltzer, K. Church, and H. Hermansky, “Deep neural network features and semi-supervised training for low resource speech recognition,” in
2013
Earlier work this paper cites.
N. Jaitly and G. Hinton, “Vocal tract length perturbation (VTLP) improves speech recognition,” in
2013
Earlier work this paper cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,”
2014
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in
2014
Cited alongside, same era.
Y. Miao, M. Gowayyed, and F. Metze, “EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,” in
2015
Cited alongside, same era.
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in
2015
Cited alongside, same era.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in
2015
Cited alongside, same era.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, and
2015
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
S. Karita, S. Watanabe, T. Iwata, A. Ogawa, and M. Delcroix, “Semi-supervised end-to-end speech recognition,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
S. H. K. Parthasarathi and N. Strom, “Lessons from building acoustic models with a million hours of speech,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
T. Hori, R. Astudillo, T. Hayashi, Y. Zhang, S. Watanabe, and J. L. Roux, “Cycle-consistency training for end-to-end speech recognition,” in
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Cited alongside, same era.
Y. Huang, Y. Wang, and Y. Gong, “Semi-supervised training in deep learning acoustic model,” in
2016
Cited alongside, same era.
R. Sennrich, B. Haddow, and A. Birch, “Improving neural machine translation models with monolingual data,” in
2016
Cited alongside, same era.
J. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in
2017
Cited alongside, same era.
A. Tjandra, S. Sakti, and S. Nakamura, “Listening while speaking: Speech chain by deep learning,” in
2017
Cited alongside, same era.
——, “Machine speech chain with one-shot speaker adaptation,” in
2018
Cited alongside, same era.
Later among the works it cites.
M.-K. Baskar, S. Watanabe, R. Astudillo, T. Hori, L. Burget, and J. Černocký, “Semi-supervised sequence-to-sequence ASR using unpaired speech and text,” in
2019
Later among the works it cites.
Y. Ren, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “Almost unsupervised text to speech and automatic speech recognition,” in
2019
Later among the works it cites.
A. H. Liu, H. Lee, and L. Lee, “Adversarial training of end-to-end speech recognition using a criticizing language model,” in
2019
Later among the works it cites.
G. Kurata and K. Audhkhasi, “Improved knowledge distillation from bi-directional to uni-directional LSTM CTC for end-to-end speech recognition,” in
2019
Later among the works it cites.
J. Kahn, A. Lee, and A. Hannun, “Self-training for end-to-end speech recognition,” in
2020
Closest in time.
2020
Closest in time.