Fetching the paper…
Reading the bibliography…
Recent progress in self-training, self-supervised pretraining and unsupervised learning enabled well performing speech recognition systems without any labeled data.
“Globalphone: a multilingual speech and text database developed at karlsruhe university,”
T. Schultz, · 2002
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Product quantization for nearest neighbor search,”
H. Jegou, M. Douze, and C. Schmid, · 2011
Earlier work this paper cites.
“Speech recognition and keyword spotting for low-resource languages: Babel project research at cued,”
M. Gales, K. M Knill, A. Ragni, and S. Rath, · 2014
Earlier work this paper cites.
“Categorical reparameterization with gumbel-softmax,”
Eric Jang, Shixiang Gu, and Ben Poole, · 2016
Earlier work this paper cites.
“Panphon: A resource for mapping IPA segments to articulatory feature vectors,”
David R. M., Patrick L., et al., · 2016
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, and et al., · 2017
Earlier work this paper cites.
“Sequence-based multi-lingual low resource speech recognition,”
S. Dalmia, R. Sanabria, F. Metze, and A. Black, · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aä. Oord, Y. Li, and O. Vinyals, · 2018
Earlier work this paper cites.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Y.-A. Chung and J. Glass, · 2018
Earlier work this paper cites.
“Completely unsupervised phoneme recognition by adversarially learning mapping relationships from audio embeddings,”
D. Liu, K.-Y. Chen, H.-Y. Lee, and L. s. Lee, · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, · 2018
Earlier work this paper cites.
“An unsupervised autoregressive model for speech representation learning,”
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, · 2019
Cited alongside, same era.
“End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures,”
G. Synnaeve, Q. Xu, et al., · 2019
Cited alongside, same era.
“Completely unsupervised speech recognition by a generative adversarial network harmonized with iteratively refined hidden markov models,”
K.-Y. Chen, C.-P. Tsai, D.-R. Liu, et al., · 2019
Cited alongside, same era.
“Common voice: A massively-multilingual speech corpus,”
R. Ardila et al., · 2019
Cited alongside, same era.
“fairseq: A fast, extensible toolkit for sequence modeling,”
M. Ott, S. Edunov, A. Baevski, et al., · 2019
“Towards zero-shot learning for automatic phonemic transcription,”
X. Li, S. Dalmia, D. Mortensen, et al., · 2020
Later among the works it cites.
“Adapt-and-adjust: Overcoming the long-tail problem of multilingual speech recognition,”
G. Winata, G. Wang, C. Xiong, and S. Hoi, · 2020
Later among the works it cites.
“Unsupervised cross-lingual representation learning for speech recognition,”
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, · 2020
Later among the works it cites.
“Mls: A large-scale multilingual dataset for speech research,”
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, · 2020
Later among the works it cites.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
A. Baevski, S. Schneider, and M. Auli, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Wav2letter++: A fast open-source speech recognition system,”
V. Pratap, A. Hannun, Q. Xu, et al., · 2019
Cited alongside, same era.
“Massively multilingual asr: 50 languages, 1 model, 1 billion parameters,”
V. Pratap, A. Sriram, et al., · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, · 2020
Cited alongside, same era.
“Iterative pseudo-labeling for speech recognition,”
Q. Xu, T. Likhomanenko, J. Kahn, A. Hannun, G. Synnaeve, and R. Collobert, · 2020
Cited alongside, same era.
“slimipl: Language-model-free iterative pseudo-labeling,”
T. Likhomanenko, Q. Xu, J. Kahn, G. Synnaeve, and R. Collobert, · 2020
Cited alongside, same era.
“Improved noisy student training for automatic speech recognition,”
D. Park, Y. Zhang, Y. Jia, et al., · 2020
Cited alongside, same era.
“Universal phone recognition with a multilingual allophone system,”
X. Li, S. Dalmia, J. Li, et al., · 2020
Cited alongside, same era.
Later among the works it cites.
“Grapheme-to-phoneme transduction for cross-language asr,”
M. Hasegawa-Johnson et al., · 2020
Later among the works it cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
W.-N. Hsu et al., · 2021
Closest in time.
“Self-training and pre-training are complementary for speech recognition,”
Qiantong X., Alexei B., et al., · 2021
Closest in time.
“Unsupervised speech recognition,”
A. Baevski, W.-N. Hsu, A. Conneau, and M. Auli, · 2021
Closest in time.
“Zero-shot cross-lingual phonetic recognition with external language embedding,”
H. Gao, J. Ni, Y. Zhang, K. Qian, et al., · 2021
Closest in time.
C. Jacobs and H. Kamper, · 2021
Closest in time.
“Differentiable allophone graphs for language-universal speech recognition,”
B. Yan, S. Dalmia, D. Mortensen, F. Metze, and S. Watanabe, · 2021
Closest in time.