Fetching the paper…
Reading the bibliography…
Self-supervised learning via masked prediction pre-training (MPPT) has shown impressive performance on a range of speech-processing tasks.
Connectionist speech recognition: A hybrid approach
H. Bourlard, H.A. Bourlard, & N. Morgan, · 1994
Earlier work this paper cites.
“Utilizing untranscribed training data to improve performance,”
G. Zavaliagkos & T. Colthurst, · 1998
Earlier work this paper cites.
“Unsupervised training of a speech recognizer using TV broadcasts,”
T. Kemp & A. Waibel, · 1998
Earlier work this paper cites.
“Unsupervised training of acoustic models for large vocabulary continuous speech recognition,”
F. Wessel & H. Ney, · 2001
Earlier work this paper cites.
“Unsupervised acoustic model training,”
L. Lamel, J. Gauvain, & G. Adda, · 2002
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F.J. Gomez, & J. Schmidhuber, · 2006
Earlier work this paper cites.
“Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
G.E. Dahl, D. Yu, L. Deng, & A. Acero, · 2012
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, & S. Khudanpur, · 2015
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
R. Sennrich, B. Haddow, & A. Birch, · 2016
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, Ł. Kaiser, & I. Polosukhin, · 2017
Earlier work this paper cites.
“Understanding deep learning requires rethinking generalization,”
C. Zhang, S. Bengio, M. Hardt, B. Recht, & O. Vinyals, · 2017
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
A. van den Oord, Y. Li, & O. Vinyals, · 2018
Cited alongside, same era.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
T. Kudo & J. Richardson, · 2018
Cited alongside, same era.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, & Y. Bengio, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised pre-training for speech recognition,”
S. Schneider, A. Baevski, R. Collobert, & M. Auli, · 2019
Cited alongside, same era.
“Improved noisy student training for automatic speech recognition,”
D.S. Park, Y. Zhang, Y. Jia, W. Han, C.C. Chiu, B. Li, Y. Wu, & Q. Le, · 2020
Later among the works it cites.
“HuBERT: How much can a bad teacher benefit ASR pre-training?,”
W.N. Hsu, Y.H.H. Tsai, B. Bolte, R. Salakhutdinov, & A. Mohamed, · 2021
Later among the works it cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
W.N. Hsu, B. Bolte, Y.H.H. Tsai, K. Lakhotia, R. Salakhutdinov, & A. Mohamed, · 2021
Later among the works it cites.
“On the learning dynamics of semi-supervised training for asr,”
E. Wallington, B. Kershenbaum, O. Klejch, & P. Bell, · 2021
Later among the works it cites.
“Improving streaming automatic speech recognition with non-streaming model distillation on unsupervised data,”
T. Doutre, W. Han, M. Ma, Z. Lu, C.C. Chiu, R. Pang, A. Narayanan, A. Misra, Y. Zhang, & L. Cao, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“An unsupervised autoregressive model for speech representation learning,”
Y. Chung, W. Hsu, H. Tang, & J.R. Glass, · 2019
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D.S. Park, W. Chan, Y. Zhang, C.C. Chiu, B. Zoph, E.D. Cubuk, & Q. Le, · 2019
Cited alongside, same era.
“Multi-task self-supervised learning for robust speech recognition,”
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, & Y. Bengio, · 2020
Cited alongside, same era.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
A. Baevski, S. Schneider, & M. Auli, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, & M. Auli, · 2020
Cited alongside, same era.
“Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,”
Y. Shi, Y. Wang, C. Wu, C.F. Yeh, J. Chan, F. Zhang, D. Le, & M. Seltzer, · 2021
Later among the works it cites.
“Self-supervised speech representation learning: A review,”
A. Mohamed, H.y. Lee, L. Borgholt, J.D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.W. Li, K. Livescu, L. Maaløe, T.N. Sainath, & S. Watanabe, · 2022
Closest in time.
“Supervision-guided codebooks for masked prediction in speech pre-training,”
C. Wang, Y. Wang, Y. Wu, S. Chen, J. Li, S. Liu, & F. Wei, · 2022
Closest in time.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, & others, · 2022
Closest in time.
“Knowledge distillation for neural transducers from large self-supervised pre-trained models,”
X. Yang, Q. Li, & P.C. Woodland, · 2022
Closest in time.