Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel multi-modal multi-task encoder-decoder pre-training framework (MMSpeech) for Mandarin automatic speech recognition (ASR), which employs both unlabeled speech and text data.
“A new algorithm for data compression,”
Philip Gage, · 1994
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
William Chan, Navdeep Jaitly, Quoc V Le, and Oriol Vinyals, · 2015
Earlier work this paper cites.
The Homophone Effect in Mandarin Word Recognition
Wei Zhou, · 2015
Earlier work this paper cites.
“On using monolingual corpora in neural machine translation,”
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Earlier work this paper cites.
“A study on data augmentation of reverberant speech for robust speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, Michael L Seltzer, and Sanjeev Khudanpur, · 2017
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Earlier work this paper cites.
“Aishell-2: Transforming mandarin asr research into industrial scale,”
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu, · 2018
Earlier work this paper cites.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Earlier work this paper cites.
“Transformers with convolutional context for asr,”
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer, · 2019
Cited alongside, same era.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer, · 2019
Cited alongside, same era.
“Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,”
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee, · 2019
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
“Slam: A unified encoder for speech and language modeling via speech-text joint pre-training,”
Ankur Bapna, Yu-an Chung, Nan Wu, Anmol Gulati, Ye Jia, Jonathan H Clark, Melvin Johnson, Jason Riesa, Alexis Conneau, and Yu Zhang, · 2021
Later among the works it cites.
“Learning shared semantic space for speech-to-text translation,”
Chi Han, Mingxuan Wang, Heng Ji, and Lei Li, · 2021
Later among the works it cites.
“Improving speech translation by understanding and learning from the auxiliary text translation task,”
Yun Tang, Juan Pino, Xian Li, Changhan Wang, and Dmitriy Genzel, · 2021
Later among the works it cites.
“Data2vec: A general framework for self-supervised learning in speech, vision and language,”
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli, · 2022
Closest in time.
“Pre-training transformer decoder for end-to-end asr model with unpaired speech data,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“fairseq s2t: Fast speech-to-text modeling with fairseq,”
Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,”
Junyi Ao, Rui Wang, Long Zhou, Shujie Liu, Shuo Ren, Yu Wu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, et al., · 2021
Cited alongside, same era.
“M6: A chinese multimodal pretrainer,”
Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, et al., · 2021
Cited alongside, same era.
“A general multi-task learning framework to leverage text data for speech to text tasks,”
Yun Tang, J. Pino, Changhan Wang, Xutai Ma, and Dmitriy Genzel, · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Cited alongside, same era.
Junyi Ao, Ziqiang Zhang, Long Zhou, Shujie Liu, Haizhou Li, Tom Ko, Lirong Dai, Jinyu Li, Yao Qian, and Furu Wei, · 2022
Closest in time.
“Unified speech-text pre-training for speech translation and recognition,”
Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang, Wei-Ning Hsu, Jiatao Gu, Alexei Baevski, Xian Li, Abdelrahman Mohamed, Michael Auli, et al., · 2022
Closest in time.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Closest in time.
“Wav2seq: Pre-training speech-to-text encoder-decoder models using pseudo languages,”
Felix Wu, Kwangyoun Kim, Shinji Watanabe, Kyu Han, Ryan McDonald, Kilian Q Weinberger, and Yoav Artzi, · 2022
Closest in time.
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang, · 2022
Closest in time.
“mslam: Massively multilingual joint pre-training for speech and text,”
Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, and Alexis Conneau, · 2022
Closest in time.
“Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition,”
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al., · 2022
Closest in time.