Fetching the paper…
Reading the bibliography…
Self-supervised pre-training is an effective approach to leveraging a large amount of unlabelled data to reduce word error rates (WERs) of automatic speech recognition (ASR) systems.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, & J. Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Restructuring of deep neural network acoustic models with singular value decomposition,”
J. Xue, J. Li, & Y. Gong, · 2013
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
G. Hinton, O. Vinyals, & J. Dean, · 2014
Earlier work this paper cites.
“Learning small-size DNN with output-distribution-based criteria,”
J. Li, R. Zhao, J.T. Huang, & Y. Gong, · 2014
Earlier work this paper cites.
“LibriSpeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, & S. Khudanpur, · 2015
Earlier work this paper cites.
“On the compression of recurrent neural networks with an application to LVCSR acoustic modeling for embedded speech recognition,”
R. Prabhavalkar, O. Alsharif, A. Bruguier, & I. McGraw, · 2016
Earlier work this paper cites.
“Sequence student-teacher training of deep neural networks,”
J.H.M. Wong & M.J.F. Gales, · 2016
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, & Y. Bengio, · 2016
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,”
K. Rao, H. Sak, & R. Prabhavalkar, · 2017
Earlier work this paper cites.
“An investigation of a knowledge distillation method for CTC acoustic models,”
R. Takashima, S. Li, & H. Kawai, · 2018
Earlier work this paper cites.
“Improved knowledge distillation from bi-directional to uni-directional LSTM CTC for end-to-end speech recognition,”
G. Kurata & K. Audhkhasi, · 2018
Earlier work this paper cites.
“Compression of end-to-end models,”
R. Pang, T. Sainath, R. Prabhavalkar, S. Gupta, & C.C. Chiu, · 2018
Cited alongside, same era.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
A. Kannan, Y. Wu, P. Nguyen, T.N. Sainath, Z. Chen, & R. Prabhavalkar, · 2018
Cited alongside, same era.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
T. Kudo & J. Richardson, · 2018
Cited alongside, same era.
“ESPnet: End-to-end speech processing toolkit,”
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, et al · 2018
Cited alongside, same era.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M.W. Chang, K. Lee, & K. Toutanova, · 2019
Cited alongside, same era.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, & M. Auli, · 2020
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
A. Gulati, J. Qin, C.C. Chiu, N. Parmar, Y. Zhang, et al · 2020
Later among the works it cites.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Y. Zhang, J. Qin, D.S. Park, W. Han, C. Chiu, et al · 2020
Later among the works it cites.
“Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,”
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, et al · 2020
Later among the works it cites.
“Knowledge distillation from offline to streaming RNN transducer for end-to-end speech recognition,”
G. Kurata & G. Saon, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Baevski, S. Schneider, & M. Auli, · 2019
Cited alongside, same era.
“Investigation of sequence-level knowledge distillation methods for CTC acoustic models,”
R. Takashima, L. Sheng, & H. Kawai, · 2019
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Y. He, T.N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, et al · 2019
Cited alongside, same era.
“Improving transformer-based speech recognition using unsupervised pre-training,”
D. Jiang, X. Lei, W. Li, N. Luo, Y. Hu, et al · 2019
Cited alongside, same era.
“RNN-T for latency controlled ASR with improved beam search,”
M. Jain, K. Schubert, J. Mahadeokar, C.F. Yeh, K. Kalgaonkar, et al · 2019
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
D.S. Park, W. Chan, Y. Zhang, C.C. Chiu, B. Zoph, et al · 2019
Cited alongside, same era.
“Language models are few-shot learners,”
T. Brown, B. Mann, N. Ryder, M. Subbiah, J.D. Kaplan, & Dhariwal, · 2020
Cited alongside, same era.
Z. Peng, A. Budhkar, I. Tuil, J. Levy, P. Sobhani, R. Cohen, & J. Nassour, · 2021
Closest in time.
“Sparsification via compressed sensing for automatic speech recognition,”
K. Zhen, H.D. Nguyen, F.J. Chang, A. Mouchtaris, & A. Rastrow, · 2021
Closest in time.
“Improving streaming transformer based ASR under a framework of self-supervised learning,”
S. Cao, Y. Kang, Y. Fu, X. Xu, S. Sun, et al · 2021
Closest in time.
“Improving streaming automatic speech recognition with non-streaming model distillation on unsupervised data,”
T. Doutre, W. Han, M. Ma, Z. Lu, C.C. Chiu, et al · 2021
Closest in time.
“TutorNet: Towards flexible knowledge distillation for end-to-end speech recognition,”
J. Yoon, H. Lee, H.Y. Kim, W.I. Cho, & N.S. Kim, · 2021
Closest in time.
“Efficient knowledge distillation for RNN-transducer models,”
S. Panchapagesan, D.S. Park, C.C. Chiu, Y. Shangguan, Q. Liang, & A. Gruenstein, · 2021
Closest in time.
“CoDERT: Distilling encoder representations with co-learning for transducer-based speech recognition,”
R.V. Swaminathan, B. King, G.P. Strimel, J. Droppo, & A. Mouchtaris, · 2021
Closest in time.