Fetching the paper…
Reading the bibliography…
Recently, self-supervised pretraining has achieved impressive results in end-to-end (E2E) automatic speech recognition (ASR).
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al., · 2011
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Unsupervised visual representation learning by context prediction,”
Carl Doersch, Abhinav Gupta, and Alexei A. Efros, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Joint CTC/attention decoding for end-to-end speech recognition,”
Takaaki Hori, Shinji Watanabe, and John Hershey, · 2017
Earlier work this paper cites.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Earlier work this paper cites.
“Unsupervised pretraining for sequence to sequence learning,”
Prajit Ramachandran, Peter Liu, and Quoc Le, · 2017
Earlier work this paper cites.
“Accurate, large minibatch SGD: training imagenet in 1 hour,”
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He, · 2017
Earlier work this paper cites.
“Universal language model fine-tuning for text classification,”
Jeremy Howard and Sebastian Ruder, · 2018
Earlier work this paper cites.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
A. Kannan, Y. Wu, P. Nguyen, T. N. Sainath, Z. Chen, and R. Prabhavalkar, · 2018
Earlier work this paper cites.
“Cold fusion: Training seq2seq models together with language models,”
A. Sriram, H. Jun, S. Satheesh, and A. Coates, · 2018
Earlier work this paper cites.
“AISHELL-2: transforming mandarin ASR research into industrial scale,”
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu, · 2018
Cited alongside, same era.
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai, · 2018
Cited alongside, same era.
“The speechtransformer for large-scale mandarin chinese speech recognition,”
Yuanyuan zhao, Jie Li, Xiaorui Wang, and Yan Li, · 2019
Cited alongside, same era.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised Pre-Training for Speech Recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“Distilling the knowledge of bert for sequence-to-sequence asr,”
Hayato Futami, Hirofumi Inaguma, Sei Ueno, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara, · 2020
Later among the works it cites.
“Applying wav2vec2.0 to speech recognition in various low-resource languages,”
Cheng Yi, Jianzhong Wang, Ning Cheng, Shiyu Zhou, and Bo Xu, · 2020
Later among the works it cites.
“Leveraging unpaired text data for training end-to-end speech-to-intent systems,”
Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas, Zvi Kons, Kartik Audhkhasi, Brian Kingsbury, Ron Hoory, and Michael Picheny, · 2020
Later among the works it cites.
“Efficiently fusing pretrained acoustic and linguistic encoders for low-resource speech recognition,”
Cheng Yi, Shiyu Zhou, and Bo Xu, · 2021
Closest in time.
“Improving Accent Identification and Accented Speech Recognition Under a Framework of Self-Supervised Learning,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Language models are unsupervised multitask learners,”
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever, · 2019
Cited alongside, same era.
“Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter,”
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf, · 2019
Cited alongside, same era.
“fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Cited alongside, same era.
“Uer: An open-source toolkit for pre-training models,”
Zhe Zhao, Hui Chen, Jinbin Zhang, Xin Zhao, Tao Liu, Wei Lu, Xi Chen, Haotang Deng, Qi Ju, and Xiaoyong Du, · 2019
Cited alongside, same era.
“SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Transformers: State-of-the-art natural language processing,”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush, · 2020
Cited alongside, same era.
Keqi Deng, Songjun Cao, and Long Ma, · 2021
Closest in time.
“Improving Streaming Transformer Based ASR Under a Framework of Self-Supervised Learning,”
Songjun Cao, Yueteng Kang, Yanzhe Fu, Xiaoshuo Xu, Sining Sun, Yike Zhang, and Long Ma, · 2021
Closest in time.
“Pre-training transformer decoder for end-to-end asr model with unpaired text data,”
Changfeng Gao, Gaofeng Cheng, Runyan Yang, Han Zhu, Pengyuan Zhang, and Yonghong Yan, · 2021
Closest in time.
“History utterance embedding transformer lm for speech recognition,”
Keqi Deng, Gaofeng Cheng, Haoran Miao, Pengyuan Zhang, and Yonghong Yan, · 2021
Closest in time.
“Integrating knowledge into end-to-end speech recognition from external text-only data,”
Ye Bai, Jiangyan Yi, Jianhua Tao, Zhengqi Wen, Zhengkun Tian, and Shuai Zhang, · 2021
Closest in time.
“Non-autoregressive transformer-based end-to-end ASR using BERT,”
Fu-Hao Yu and Kuan-Yu Chen, · 2021
Closest in time.
“On scaling contrastive representations for low-resource speech recognition,”
Lasse Borgholt, Tycho M. S. Tax, Jakob D. Havtorn, Lars Maaløe, and Christian Igel, · 2021
Closest in time.
“Joint masked cpc and ctc training for asr,”
Chaitanya Talnikar, Tatiana Likhomanenko, Ronan Collobert, and Gabriel Synnaeve, · 2021
Closest in time.
“Self-training and pre-training are complementary for speech recognition,”
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli, · 2021
Closest in time.