Fetching the paper…
Reading the bibliography…
Self-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success.
“The ami meeting corpus: A pre-announcement,”
Jean Carletta, Simone Ashby, Sebastien Bourban, et al., · 2005
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Intelligent selection of language model training data,”
Robert C Moore and William Lewis, · 2010
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“On using monolingual corpora in neural machine translation,” 2015
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Dual learning for machine translation,”
Di He, Yingce Xia, Tao Qin, Liwei Wang, Nenghai Yu, Tie-Yan Liu, and Wei-Ying Ma, · 2016
Earlier work this paper cites.
“Listening while speaking: Speech chain by deep learning,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Earlier work this paper cites.
“Joint ctc-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Earlier work this paper cites.
“Cold fusion: Training seq2seq models together with language models,” 2017
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates, · 2017
Earlier work this paper cites.
“Effectively building tera scale maxent language models incorporating non-linguistic signals,”
Fadi Biadsy, Mohammadreza Ghodsi, and Diamantino Caseiro, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Back-translation-style data augmentation for end-to-end ASR,”
Tomoki Hayashi, Shinji Watanabe, Yu Zhang, Tomoki Toda, Takaaki Hori, Ramon Astudillo, and Kazuya Takeda, · 2018
Earlier work this paper cites.
“Machine speech chain with one-shot speaker adaptation,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2018
Earlier work this paper cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, et al., · 2018
Earlier work this paper cites.
“Hierarchical generative modeling for controllable speech synthesis,”
Wei-Ning Hsu et al., · 2018
Earlier work this paper cites.
“Training neural speech recognition systems with synthetic speech augmentation,”
Jason Li, Ravi Gadde, Boris Ginsburg, and Vitaly Lavrukhin, · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson, · 2018
Earlier work this paper cites.
“Adafactor: Adaptive learning rates with sublinear memory cost,”
Noam Shazeer and Mitchell Stern, · 2018
Cited alongside, same era.
“Cycle-consistency training for end-to-end speech recognition,”
Takaaki Hori, Ramon Astudillo, Tomoki Hayashi, Yu Zhang, Shinji Watanabe, and Jonathan Le Roux, · 2019
Cited alongside, same era.
“Speech recognition with augmented synthesized speech,”
Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia, Pedro Moreno, Yonghui Wu, and Zelin Wu, · 2019
Cited alongside, same era.
“Semi-supervised end-to-end speech recognition using text-to-speech and autoencoders,”
Shigeki Karita et al., · 2019
Cited alongside, same era.
“Almost unsupervised text to speech and automatic speech recognition,”
Yi Ren, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Cited alongside, same era.
“Self-training with noisy student improves ImageNet classification,”
Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le, · 2020
Later among the works it cites.
“Improved noisy student training for automatic speech recognition,”
Daniel S Park, Yu Zhang, Ye Jia, Wei Han, Chung-Cheng Chiu, Bo Li, Yonghui Wu, and Quoc V Le, · 2020
Later among the works it cites.
“FixMatch: Simplifying semi-supervised learning with consistency and confidence,”
Sohn Kihyuk et al., · 2020
Later among the works it cites.
“Semi-supervised learning with data augmentation for end-to-end ASR,”
Felix Weninger, Franco Mana, Roberto Gemello, Jesús Andrés-Ferrer, and Puming Zhan, · 2020
Later among the works it cites.
“Improving speech recognition using consistent predictions on synthesized speech,”
Gary Wang, Andrew Rosenberg, Zhehuai Chen, Yu Zhang, Bhuvana Ramabhadran, Yonghui Wu, and Pedro Moreno, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Adversarial training of end-to-end speech recognition using a criticizing language model,”
Alexander H Liu, Hung-yi Lee, and Lin-shan Lee, · 2019
Cited alongside, same era.
“Semi-supervised sequence-to-sequence ASR using unpaired speech and text,”
Murali Karthick Baskar et al., · 2019
Cited alongside, same era.
“Streaming end-to-end speech recognition with joint ctc-attention based models,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Cited alongside, same era.
“Component fusion: Learning replaceable language model component for end-to-end speech recognition system,”
Changhao Shan, Chao Weng, Guangsen Wang, Dan Su, Min Luo, Dong Yu, and Lei Xie, · 2019
Cited alongside, same era.
“Parrotron: An end-to-end speech-to-speech conversion model and its applications to hearing-impaired speech and speech separation,” 2019
Fadi Biadsy, Ron J. Weiss, Pedro J. Moreno, Dimitri Kanevsky, and Ye Jia, · 2019
Cited alongside, same era.
“Contrastive sequence-to-sequence data selector,” Nov. 14 2019,
Wei Wang, Bowen Liang, Macduff Hughes, Taro Watanabe, Tetsuji Nakagawa, and Alexander Rudnick, · 2019
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Cited alongside, same era.
Later among the works it cites.
“Improving speech recognition using GAN-based speech synthesis and contrastive unspoken text selection,”
Zhehuai Chen, Andrew Rosenberg, Yu Zhang, Gary Wang, Bhuvana Ramabhadran, and Pedro Moreno, · 2020
Later among the works it cites.
“Hybrid autoregressive transducer (hat),”
Ehsan Variani, David Rybach, Cyril Allauzen, and Michael Riley, · 2020
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, et al., · 2020
Later among the works it cites.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, et al., · 2020
Later among the works it cites.
“Improving tail performance of a deliberation e2e asr model using a largetext corpus,”
Cal Peyser, Sepand Mavandadi, Tara N Sainath, James Apfel, Ruoming Pang, and Shankar Kumar, · 2020
Later among the works it cites.
“A domain-specific supercomputer for training deep neural networks,”
Norman P Jouppi, Doe Hyun Yoon, George Kurian, Sheng Li, Nishant Patil, James Laudon, Cliff Young, and David Patterson, · 2020
Later among the works it cites.
“Self-supervised text-independent speaker verification using prototypical momentum contrastive learning,”
Wei Xia, Chunlei Zhang, Chao Weng, Meng Yu, and Dong Yu, · 2021
Closest in time.
“Joint masked cpc and ctc training for asr,”
Chaitanya Talnikar, Tatiana Likhomanenko, Ronan Collobert, and Gabriel Synnaeve, · 2021
Closest in time.
“Semi-supervision in asr: Sequential mixmatch and factorized tts-based augmentation,”
Zhehuai Chen et al., · 2021
Closest in time.
“Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,”
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, et al., · 2021
Closest in time.
“An efficient streaming non-recurrent on-device end-to-endmodel with improvements to rare-word modeling,”
Tara Sainath et al., · 2021
Closest in time.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.
“Speechstew: Simply mix all available speech recognition data to train one large neural network,”
William Chan, Daniel Park, Chris Lee, Yu Zhang, Quoc Le, and Mohammad Norouzi, · 2021
Closest in time.