Fetching the paper…
Reading the bibliography…
We study pseudo-labeling for the semi-supervised training of ResNet, Time-Depth Separable ConvNets, and Transformers for speech recognition, with either CTC or Seq2Seq loss functions.
Word-level speech recognition with a dynamic lexicon
Collobert, R., Hannun, A., and Synnaeve, G · 1906
Earlier work this paper cites.
Self-training for end-to-end speech recognition
Kahn, J., Lee, A., and Hannun, A · 1909
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence O (1/kˆ 2)
Nesterov, Y · 1983
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Kenlm: Faster and smaller language model queries
Heafield, K · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
Hinton, G., Deng, L., Yu, D., Dahl, G., rahman Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., and Kingsbury, B · 2012
Earlier work this paper cites.
Japanese and korean voice search
Schuster, M. and Nakajima, K · 2012
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Graves, A. and Jaitly, N · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Librispeech: an ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Amodei, D., Ananthanarayanan, S., Anubhai, R., et al · 2016
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, W., Jaitly, N., Le, Q. V., and Vinyals, O · 2016
Earlier work this paper cites.
Towards better decoding and language model integration in sequence to sequence models
Chorowski, J. and Jaitly, N · 2016
Earlier work this paper cites.
Wav2letter: an end-to-end convnet-based speech recognition system
Collobert, R., Puhrsch, C., and Synnaeve, G · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Neural speech recognizer: Acoustic-to-word lstm model for large vocabulary speech recognition
Soltau, H., Liao, H., and Sak, H · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Cited alongside, same era.
Knowledge distillation across ensembles of multilingual models for low-resource languages
Cui, J., Kingsbury, B., Ramabhadran, B., Saon, G., Sercu, T., Audhkhasi, K., Sethy, A., Nussbaum-Thom, M., and Rosenberg, A · 2017
Cited alongside, same era.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Cited alongside, same era.
The capio 2017 conversational speech recognition system, 2017
Han, K. J., Chandrashekaran, A., et al · 2017
wav2letter++: The fastest open-source speech recognition system
Pratap, V., Hannun, A., Xu, Q., et al · 2018
Later among the works it cites.
Fully convolutional speech recognition
Zeghidour, N., Xu, Q., Liptchinsky, V., Usunier, N., et al · 2018
Later among the works it cites.
Zhou, S., Dong, L., Xu, S., and Xu, B · 2018
Later among the works it cites.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2019
Closest in time.
Reducing transformer depth on demand with structured dropout
Fan, A., Grave, E., and Joulin, A · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A comparison of sequence-to-sequence models for speech recognition
Prabhavalkar, R., Rao, K., Sainath, T. N., et al · 2017
Cited alongside, same era.
English conversational telephone speech recognition by humans and machines
Saon, G., Kurata, G., Sercu, T., Audhkhasi, K., Thomas, S., Dimitriadis, D., Cui, X., Ramabhadran, B., Picheny, M., Lim, L.-L., Roomi, B., and Hall, P · 2017
Cited alongside, same era.
Cold fusion: Training seq2seq models together with language models
Sriram, A., Jun, H., Satheesh, S., and Coates, A · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., et al · 2017
Cited alongside, same era.
Semi-supervised DNN training with word selection for ASR
Veselỳ, K., Burget, L., and Cernockỳ, J · 2017
Cited alongside, same era.
Residual convolutional ctc networks for automatic speech recognition
Wang, Y., Deng, X., Pu, S., and Huang, Z · 2017
Cited alongside, same era.
The microsoft 2016 conversational speech recognition system
Xiong, W., Droppo, J., Huang, X., Seide, F., Seltzer, M., Stolcke, A., Yu, D., and G., Z · 2017
Cited alongside, same era.
Closest in time.
State-of-the-art speech recognition using multi-stream self-attention with dilated 1d convolutions, 2019
Han, K. J., Prieto, R., Wu, K., and Ma, T · 2019
Closest in time.
Sequence-to-sequence speech recognition with time-depth separable convolutions
Hannun, A., Lee, A., Xu, Q., and Collobert, R · 2019
Closest in time.
Streaming end-to-end speech recognition for mobile devices
He, Y., Sainath, T. N., Prabhavalkar, R., et al · 2019
Closest in time.
A comparative study on transformer vs rnn in speech applications, 2019
Karita, S., Chen, N., Hayashi, T., et al · 2019
Closest in time.
Semi-supervised training for end-to-end models via weak distillation
Li, B., Sainath, T. N., Pang, R., and Wu, Z · 2019
Closest in time.
Who needs words? lexicon-free speech recognition
Likhomanenko, T., Synnaeve, G., and Collobert, R · 2019
Closest in time.
Rwth asr systems for librispeech: Hybrid vs attention
Lüscher, C., Beck, E., Irie, K., et al · 2019
Closest in time.
Transformers with convolutional context for asr, 2019
Mohamed, A., Okhonko, D., and Zettlemoyer, L · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Closest in time.
Specaugment: A simple data augmentation method for automatic speech recognition
Park, D. S., Chan, W., Zhang, Y., et al · 2019
Closest in time.
Lessons from building acoustic models with a million hours of speech
Parthasarathi, S. H. K. and Strom, N · 2019
Closest in time.
Transformer-based acoustic modeling for hybrid speech recognition, 2019
Wang, Y., Mohamed, A., Le, D., Liu, C., Xiao, A., Mahadeokar, J., Huang, H., Tjandra, A., Zhang, X., Zhang, F., Fuegen, C., Zweig, G., and Seltzer, M. L · 2019
Closest in time.