Fetching the paper…
Reading the bibliography…
Self-supervised training has shown promising gains in pretraining models and facilitating the downstream finetuning for speech recognition, like multilingual ASR.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Categorical reparameterization with gumbel-softmax,”
Eric Jang, Shixiang Gu, and Ben Poole, · 2016
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Cross-lingual language model pretraining,”
Alexis Conneau and Guillaume Lample, · 2019
Earlier work this paper cites.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
Santiago Pascual, Mirco Ravanelli, Joan Serra, Antonio Bonafonte, and Yoshua Bengio, · 2019
Earlier work this paper cites.
“Unsupervised pre-training of bidirectional speech encoders via masked reconstruction,”
Weiran Wang, Qingming Tang, and Karen Livescu, · 2020
Earlier work this paper cites.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Yu Zhang, James Qin, Daniel S Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V Le, and Yonghui Wu, · 2020
Earlier work this paper cites.
“Representation learning for sequence data with deep autoencoding predictive components,”
Junwen Bai, Weiran Wang, Yingbo Zhou, and Caiming Xiong, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Recall and learn: Fine-tuning deep pretrained language models with less forgetting,”
Sanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che, Ting Liu, and Xiangzhan Yu, · 2020
Cited alongside, same era.
“Mls: A large-scale multilingual dataset for speech research,”
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert, · 2020
Cited alongside, same era.
“Unsupervised pretraining transfers well across languages,”
Morgane Riviere, Armand Joulin, Pierre-Emmanuel Mazaré, and Emmanuel Dupoux, · 2020
Cited alongside, same era.
“Injecting text in self-supervised speech pretraining,”
Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Gary Wang, and Pedro Moreno, · 2021
Closest in time.
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu, · 2021
Closest in time.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.
Samuel Kessler, Bethan Thomas, and Salah Karout, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Don’t stop pretraining: adapt language models to domains and tasks,”
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith, · 2020
Cited alongside, same era.
“Joint masked cpc and ctc training for asr,”
Chaitanya Talnikar, Tatiana Likhomanenko, Ronan Collobert, and Gabriel Synnaeve, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Cited alongside, same era.
Bo Li, Ruoming Pang, Tara N Sainath, Anmol Gulati, Yu Zhang, James Qin, Parisa Haghani, W Ronny Huang, Min Ma, and Junwen Bai, · 2021
Closest in time.
“Self-supervised and supervised joint training for resource-rich machine translation,”
Yong Cheng, Wei Wang, Lu Jiang, and Wolfgang Macherey, · 2021
Closest in time.
“Unispeech: Unified speech representation learning with labeled and unlabeled data,”
Chengyi Wang, Yu Wu, Yao Qian, Kenichi Kumatani, Shujie Liu, Furu Wei, Michael Zeng, and Xuedong Huang, · 2021
Closest in time.
“Hybrid unsupervised and supervised multitask learning for speech recognition in low resource languages,”
Srinivasa Raghavan and Kumar Shubham, · 2021
Closest in time.
“Large-scale asr domain adaptation using self- and semi-supervised learning,”
Dongseong Hwang, Ananya Misra, Zhouyuan Huo, Nikhil Siddhartha, Shefali Garg, David Qiu, Khe Chai Sim, Trevor Strohman, Françoise Beaufays, and Yanzhang He, · 2021
Closest in time.