Fetching the paper…
Reading the bibliography…
Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech.
“Multilingual acoustic models using distributed deep neural networks,”
Georg Heigold, Vincent Vanhoucke, Alan Senior, Patrick Nguyen, Marc’Aurelio Ranzato, Matthieu Devin, and Jeffrey Dean, · 2013
Earlier work this paper cites.
“Speech recognition and keyword spotting for low-resource languages: Babel project research at cued,”
Mark JF Gales, Kate M Knill, Anton Ragni, and Shakti P Rath, · 2014
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Earlier work this paper cites.
“Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling,”
Jaejin Cho, Murali Karthick Baskar, Ruizhi Li, Matthew Wiesner, Sri Harish Mallidi, Nelson Yalta, Martin Karafiat, Shinji Watanabe, and Takaaki Hori, · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Yu-An Chung and James Glass, · 2018
Earlier work this paper cites.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N Sainath, Zhijeng Chen, and Rohit Prabhavalkar, · 2018
Earlier work this paper cites.
Da-Rong Liu, Kuan-Yu Chen, Hung-yi Lee, and Lin-shan Lee, · 2018
Earlier work this paper cites.
“Unsupervised speech recognition via segmental empirical output distribution matching,”
Chih-Kuan Yeh, Jianshu Chen, Chengzhu Yu, and Dong Yu, · 2018
Earlier work this paper cites.
“Transliteration based approaches to improve code-switched speech recognition performance,”
Jesse Emond, Bhuvana Ramabhadran, Brian Roark, Pedro Moreno, and Min Ma, · 2018
Earlier work this paper cites.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Earlier work this paper cites.
Kuan-Yu Chen, Che-Ping Tsai, Da-Rong Liu, Hung-Yi Lee, and Lin-shan Lee, · 2019
Earlier work this paper cites.
“Common Voice: A massively-multilingual speech corpus,”
Rosana Ardila et al., · 2019
Earlier work this paper cites.
“Cross-lingual transfer learning for question answering,”
Chia-Hsuan Lee and Hung-Yi Lee, · 2019
Earlier work this paper cites.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Cited alongside, same era.
“Large-scale multilingual speech recognition with a streaming end-to-end model,”
Anjuli Kannan, Arindrima Datta, Tara N Sainath, Eugene Weinstein, Bhuvana Ramabhadran, Yonghui Wu, Ankur Bapna, Zhifeng Chen, and Seungji Lee, · 2019
Cited alongside, same era.
“Pay less attention with lightweight and dynamic convolutions,”
Felix Wu, Angela Fan, Alexei Baevski, Yann N Dauphin, and Michael Auli, · 2019
Cited alongside, same era.
“Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,”
Bo Li, Yu Zhang, Tara Sainath, Yonghui Wu, and William Chan, · 2019
Cited alongside, same era.
“MLS: A large-scale multilingual dataset for speech research,”
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux, · 2021
Later among the works it cites.
Yu-An Chung et al., · 2021
Later among the works it cites.
“Parallel Tacotron: Non-autoregressive and controllable TTS,”
Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang, Ye Jia, Ron J Weiss, and Yonghui Wu, · 2021
Later among the works it cites.
“Injecting text in self-supervised speech pretraining,”
Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Gary Wang, and Pedro Moreno, · 2021
Later among the works it cites.
“Large-scale asr domain adaptation using self- and semi-supervised learning,” 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski et al., · 2020
Cited alongside, same era.
“The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,”
Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière, Evgeny Kharitonov, Alexei Baevski, Ewan Dunbar, and Emmanuel Dupoux, · 2020
Cited alongside, same era.
“mT5: A massively multilingual pre-trained text-to-text transformer,”
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel, · 2020
Cited alongside, same era.
“Language id in the wild: Unexpected challenges on the path to a thousand-language web text corpus,”
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna, · 2020
Cited alongside, same era.
“Language-agnostic multilingual modeling,”
Arindrima Datta, Bhuvana Ramabhadran, Jesse Emond, Anjuli Kannan, and Brian Roark, · 2020
Cited alongside, same era.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Yu Zhang et al., · 2020
Cited alongside, same era.
“Semi-supervision in asr: Sequential mixmatch and factorized tts-based augmentation,”
Zhehuai Chen et al., · 2021
Cited alongside, same era.
Dongseong Hwang et al., · 2021
Later among the works it cites.
“Joint unsupervised and supervised training for multilingual ASR,”
Junwen Bai, Bo Li, Yu Zhang, Ankur Bapna, and Tara N. Sainath, · 2021
Later among the works it cites.
“Massively multilingual asr: A lifelong learning solution,”
Bo Li, Ruoming Pang, Yu Zhang, Tara N Sainath, Trevor Strohman, Parisa Haghani, Yun Zhu, Brian Farris, Neeraj Gaur, and Manasa Prasad, · 2022
Closest in time.
“mSLAM: Massively multilingual joint pre-training for speech and text,”
Ankur Bapna, Colin Cherry, Yu Zhang, Ye Jia, Melvin Johnson, Yong Cheng, Simran Khanuja, Jason Riesa, and Alexis Conneau, · 2022
Closest in time.
“Maestro: Matched speech text representations through modality matching,”
Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro Moreno, Ankur Bapna, and Heiga Zen, · 2022
Closest in time.
“Towards end-to-end unsupervised speech recognition,”
Alexander H Liu, Wei-Ning Hsu, Michael Auli, and Alexei Baevski, · 2022
Closest in time.
“Fleurs: Few-shot learning evaluation of universal representations of speech,”
Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna, · 2022
Closest in time.
“Quality at a glance: An audit of web-crawled multilingual datasets,”
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, et al., · 2022
Closest in time.
“Tts4pretrain 2.0: Advancing the use of text and speech in ASR pretraining with consistency and contrastive losses,”
Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro Moreno, and Gary Wang, · 2022
Closest in time.