Fetching the paper…
Reading the bibliography…
Pseudo-labeling is the most adopted method for pre-training automatic speech recognition (ASR) models.
“Utilizing untranscribed training data to improve performance,”
George Zavaliagkos and Thomas Colthurst, · 1998
Earlier work this paper cites.
“Unsupervised training of a speech recognizer: Recent experiments,”
Thomas Kemp and Alex Waibel, · 1999
Earlier work this paper cites.
“Unsupervised training of acoustic models for large vocabulary continuous speech recognition,”
F. Wessel and H. Ney, · 2005
Earlier work this paper cites.
“Exploiting large quantities of spontaneous speech for unsupervised training of acoustic models.,”
Bhuvana Ramabhadran, · 2005
Earlier work this paper cites.
“Unsupervised training on large amounts of broadcast news data,”
Jeff Ma, S. Matsoukas, O. Kimball, and Richard Schwartz, · 2006
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, and Faustino Gomez, · 2006
Earlier work this paper cites.
“Unsupervised training for mandarin broadcast news and conversation transcription,”
L. Wang, M. J. F. Gales, and P. C. Woodland, · 2007
Earlier work this paper cites.
“Large scale deep neural network acoustic modeling with semi-supervised training data for youtube video transcription,”
Hank Liao, Erik McDermott, and Andrew Senior, · 2013
Earlier work this paper cites.
“Do deep nets really need to be deep?,”
Jimmy Ba and Rich Caruana, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Audio augmentation for speech recognition.,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Accurate, large minibatch sgd: Training imagenet in 1 hour,”
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He, · 2017
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
Aäron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Cited alongside, same era.
“Mixed precision training,”
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu, · 2018
Cited alongside, same era.
“Iterative pseudo-labeling for speech recognition,”
Qiantong Xu, Tatiana Likhomanenko, Jacob Kahn, Awni Hannun, Gabriel Synnaeve, and Ronan Collobert, · 2018
Cited alongside, same era.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“Momentum contrast for unsupervised visual representation learning,”
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick, · 2020
Later among the works it cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Later among the works it cites.
“Supervised contrastive learning,”
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan, · 2020
Later among the works it cites.
“Large scale weakly and semi-supervised learning for low-resource video asr,”
Kritika Singh, Vimal Manohar, Alex Xiao, Sergey Edunov, Ross Girshick, Vitaliy Liptchinsky, Christian Fuegen, Yatharth Saraf, Geoffrey Zweig, and Abdelrahman Mohamed, · 2020
Later among the works it cites.
“Improved noisy student training for automatic speech recognition,”
Daniel S. Park, Yu Zhang, Ye Jia, Wei Han, Chung-Cheng Chiu, Bo Li, Yonghui Wu, and Quoc V. Le, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“End-to-end asr: from supervised to semi-supervised learning with modern architectures,” 2019
Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Tatiana Likhomanenko, et al., · 2019
Cited alongside, same era.
“Self-training for end-to-end speech recognition,”
Jacob Kahn, Ann Lee, and Awni Hannun, · 2019
Cited alongside, same era.
“Semi-supervised training for end-to-end models via weak distillation,”
Bo Li, Ruoming Pang, Tara Sainath, and Zelin Wu, · 2019
Cited alongside, same era.
“Effectiveness of self-supervised pre-training for speech recognition,”
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed, · 2019
Cited alongside, same era.
“From senones to chenones: Tied context-dependent graphemes for hybrid speech recognition,”
Duc Le, Xiaohui Zhang, Weiyi Zheng, Christian Fügen, Geoffrey Zweig, and Michael L Seltzer, · 2019
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Cited alongside, same era.
“Generative pre-training for speech with autoregressive predictive coding,”
Yu-An Chung and James Glass, · 2020
Later among the works it cites.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
Andy T Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee, · 2020
Later among the works it cites.
“Deep contextualized acoustic representations for semi-supervised speech recognition,”
Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff, · 2020
Later among the works it cites.
“Transformer-based acoustic modeling for hybrid speech recognition,”
Yongqiang Wang, Abdelrahman Mohamed, Duc Le, Chunxi Liu, Alex Xiao, et al., · 2020
Later among the works it cites.
“HuBERT: How much can a bad teacher benefit ASR pre-training?,”
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.