Fetching the paper…
Reading the bibliography…
In this paper, we introduce the Kaizen framework that uses a continuously improving teacher to generate pseudo-labels for semi-supervised speech recognition (ASR).
“On information and sufficiency,”
Solomon Kullback and Richard A Leibler, · 1951
Earlier work this paper cites.
“A time-delay neural network architecture for isolated word recognition,”
Kevin J Lang, Alex H Waibel, and Geoffrey E Hinton, · 1990
Earlier work this paper cites.
“A neural network for speaker-independent isolated word recognition,”
Kouichi Yamaguchi, Kenji Sakamoto, Toshio Akabane, and Yoshiji Fujimoto, · 1990
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Utilizing untranscribed training data to improve performance,”
George Zavaliagkos and Thomas Colthurst, · 1998
Earlier work this paper cites.
“Unsupervised training of acoustic models for large vocabulary continuous speech recognition,”
Frank Wessel and Hermann Ney, · 2004
Earlier work this paper cites.
“Model compression,”
Cristian Buciluundefined, Rich Caruana, and Alexandru Niculescu-Mizil, · 2006
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, and Faustino Gomez, · 2006
Earlier work this paper cites.
“Adaptive subgradient methods for online learning and stochastic optimization,”
John Duchi, Elad Hazan, and Yoram Singer, · 2011
Earlier work this paper cites.
“Deep neural network features and semi-supervised training for low resource speech recognition,”
Samuel Thomas, Michael L. Seltzer, Kenneth Church, and Hynek Hermansky, · 2013
Earlier work this paper cites.
“Do deep nets really need to be deep?,”
Lei Jimmy Ba and Rich Caruana, · 2014
Earlier work this paper cites.
“Dropout: a simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition.,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“A time delay neural network architecture for efficient modeling of long temporal contexts,”
Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Sequence student-teacher training of deep neural networks,”
Jeremy HM Wong and Mark Gales, · 2016
Cited alongside, same era.
“Temporal ensembling for semi-supervised learning,”
Samuli Laine and Timo Aila, · 2016
Cited alongside, same era.
“Wav2letter: an end-to-end convnet-based speech recognition system,”
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve, · 2016
Cited alongside, same era.
“Who needs words? lexicon-free speech recognition,”
Tatiana Likhomanenko, Gabriel Synnaeve, and Ronan Collobert, · 2019
Later among the works it cites.
“Self-training for end-to-end speech recognition,”
Jacob Kahn, Ann Lee, and Awni Hannun, · 2020
Later among the works it cites.
“Semi-supervised speech recognition via local prior matching,”
Wei-Ning Hsu, Ann Lee, Gabriel Synnaeve, and Awni Hannun, · 2020
Later among the works it cites.
“Improved noisy student training for automatic speech recognition,”
Daniel S Park, Yu Zhang, Ye Jia, Wei Han, Chung-Cheng Chiu, Bo Li, Yonghui Wu, and Quoc V Le, · 2020
Later among the works it cites.
“Iterative pseudo-labeling for speech recognition,”
Qiantong Xu, Tatiana Likhomanenko, Jacob Kahn, Awni Hannun, Gabriel Synnaeve, and Ronan Collobert, · 2020
Later among the works it cites.
“Semi-supervised ASR by End-to-end Self-training,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Purely sequence-trained neural networks for asr based on lattice-free mmi,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,”
Antti Tarvainen and Harri Valpola, · 2017
Cited alongside, same era.
“An exploration of dropout with lstms.,”
Gaofeng Cheng, Vijayaditya Peddinti, Daniel Povey, Vimal Manohar, Sanjeev Khudanpur, and Yonghong Yan, · 2017
Cited alongside, same era.
Low latency modeling of temporal contexts for speech recognition
Vijayaditya Peddinti et al., · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Mixed precision training,”
Sharan Narang, Gregory Diamos, Erich Elsen, Paulius Micikevicius, Jonah Alben, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al., · 2017
Cited alongside, same era.
“Semi-supervised training of acoustic models using lattice-free mmi,”
Vimal Manohar, Hossein Hadian, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
Yang Chen, Weiran Wang, and Chao Wang, · 2020
Later among the works it cites.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, et al., · 2020
Later among the works it cites.
“Bootstrap your own latent - a new approach to self-supervised learning,”
Jean-Bastien Grill et al., · 2020
Later among the works it cites.
“Large scale weakly and semi-supervised learning for low-resource video asr,”
Kritika Singh, Vimal Manohar, et al., · 2020
Later among the works it cites.
“End-to-end asr: from supervised to semi-supervised learning with modern architectures,”
Gabriel Synnaeve et al., · 2020
Later among the works it cites.
“Transformer-based acoustic modeling for hybrid speech recognition,”
Yongqiang Wang, Abdelrahman Mohamed, Duc Le, Chunxi Liu, Alex Xiao, et al., · 2020
Later among the works it cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Later among the works it cites.
“slimipl: Language-model-free iterative pseudo-labeling,”
Tatiana Likhomanenko, Qiantong Xu, Jacob Kahn, Gabriel Synnaeve, and Ronan Collobert, · 2021
Closest in time.
“Momentum Pseudo-Labeling for Semi-Supervised Speech Recognition,” 2021
Yosuke Higuchi, Niko Moritz, Jonathan Le Roux, and Takaaki Hori, · 2021
Closest in time.
“On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models,”
Xiaohui Zhang, Vimal Manohar, David Zhang, et al., · 2021
Closest in time.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.