Libri-light: A benchmark for asr with limited or no supervision
Original
Kahn, J., Rivière, M., Zheng, W., Kharitonov, E., Xu, Q., Mazaré, P.-E., Karadayi, J., Liptchinsky, V., Collobert, R., Fuegen, C., et al · 1912
Earlier work this paper cites.
Knowledge of language: Its nature, origin, and use
Chomsky, N · 1986
Earlier work this paper cites.
Towards phone segmentation for concatenative speech synthesis
Adell, J. and Bonafonte, A · 2004
Earlier work this paper cites.
Early language acquisition: cracking the speech code
Kuhl, P. K · 2004
Earlier work this paper cites.
Discriminative training for large vocabulary speech recognition
Povey, D · 2005
Earlier work this paper cites.
Posterior regularization for structured latent variable models
Ganchev, K., Gillenwater, J., Taskar, B., et al · 2010
Earlier work this paper cites.
Unsupervised speech segmentation: An analysis of the hypothesized phone boundaries
Scharenborg, O., Wan, V., and Ernestus, M · 2010
Earlier work this paper cites.
Towards unsupervised speech processing
Glass, J · 2012
Earlier work this paper cites.
Non-mainstream languages and speech recognition: Some challenges
Precoda, K · 2013
Earlier work this paper cites.
Sequence-discriminative training of deep neural networks
Veselỳ, K., Ghoshal, A., Burget, L., and Povey, D · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Deep Speech: Scaling up end-to-end speech recognition
Original
Hannun, A., Case, C., Casper, J., Catanzaro, B., Diamos, G., Elsen, E., Prenger, R., Satheesh, S., Sengupta, S., Coates, A., et al · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Librispeech: an ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Deep Speech 2: End-to-end speech recognition in English and Mandarin
Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Cheng, Q., Chen, G., et al · 2016
Earlier work this paper cites.
End-to-end attention-based large vocabulary speech recognition
Bahdanau, D., Chorowski, J., Serdyuk, D., Brakel, P., and Bengio, Y · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Nallapati, R., Zhou, B., Gulcehre, C., Xiang, B., et al · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.