Fetching the paper…
Reading the bibliography…
Self-training and unsupervised pre-training have emerged as effective approaches to improve speech recognition systems using unlabeled data.
Probability of error of some adaptive pattern-recognition machines
H. Scudder · 1965
Earlier work this paper cites.
Unsupervised word sense disambiguation rivaling supervised methods
D. Yarowsky · 1995
Earlier work this paper cites.
Automatically generating extraction patterns from untagged text
E. Riloff · 1996
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Product quantization for nearest neighbor search
H. Jegou, M. Douze, and C. Schmid · 2011
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Ethnologue: Languages of the world, nineteenth edition
M. P. Lewis, G. F. Simon, and C. D. Fennig · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, and et al · 2017
Earlier work this paper cites.
Adversarial training of end-to-end speech recognition using a criticizing language model
A. H. Liu, H.-Y. Lee, and L.-S. Lee · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
The challenge of realistic music generation: modelling raw audio at scale
S. Dieleman, A. v. d. Oord, and K. Simonyan · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
A. Baevski and M. Auli · 2018
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
T. Kudo and J. Richardson · 2018
Cited alongside, same era.
Specaugment: A simple data augmentation method for automatic speech recognition
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le · 2019
Cited alongside, same era.
End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
G. Synnaeve et al · 2019
Cited alongside, same era.
Wav2letter++: A fast open-source speech recognition system
V. Pratap et al · 2019
Later among the works it cites.
Contextnet: Improving convolutional neural networks for automatic speech recognition with global context
W. Han et al · 2020
Closest in time.
Conformer: Convolution-augmented transformer for speech recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, and et al · 2020
Closest in time.
Semi-supervised speech recognition via local prior matching
W.-N. Hsu, A. Lee, G. Synnaeve, and A. Hannun · 2020
Closest in time.
Self-training for end-to-end speech recognition
J. Kahn, A. Lee, and A. Hannun · 2020
Closest in time.
Iterative pseudo-labeling for speech recognition
Q. Xu, T. Likhomanenko, J. Kahn, A. Hannun, G. Synnaeve, and R. Collobert · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. K. Baskar, S. Watanabe, R. Astudillo, T. Hori, L. Burget, and J. Černocký · 2019
Cited alongside, same era.
Lessons from building acoustic models with a million hours of speech
S. H. K. Parthasarathi and N. Strom · 2019
Cited alongside, same era.
wav2vec: Unsupervised pre-training for speech recognition
S. Schneider, A. Baevski, R. Collobert, and M. Auli · 2019
Cited alongside, same era.
An unsupervised autoregressive model for speech representation learning
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. R. Glass · 2019
Cited alongside, same era.
Improving transformer-based speech recognition using unsupervised pre-training
D. Jiang, X. Lei, W. Li, N. Luo, Y. Hu, and et al · 2019
Cited alongside, same era.
Effectiveness of self-supervised pre-training for speech recognition
A. Baevski, M. Auli, and A. Mohamed · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
M. Ott et al · 2019
Cited alongside, same era.
Improved noisy student training for automatic speech recognition
D. S. Park, Y. Zhang, Y. Jia, W. Han, C.-C. Chiu, and et al · 2020
Closest in time.
Learning robust and multilingual speech representations
K. Kawakami, L. Wang, C. Dyer, P. Blunsom, and A. v. d. Oord · 2020
Closest in time.
Unsupervised pretraining transfers well across languages
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux · 2020
Closest in time.
Unsupervised pre-training of bidirectional speech encoders via masked reconstruction
W. Wang, Q. Tang, and K. Livescu · 2020
Closest in time.
Self-training improves pre-training for natural language understanding
J. Du, E. Grave, B. Gunel, V. Chaudhary, O. Celebi, and et al · 2020
Closest in time.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn and et al · 2020
Closest in time.