Fetching the paper…
Reading the bibliography…
Much of text-to-speech research relies on human evaluation, which incurs heavy costs and slows down the development process.
“A multilingual text-to-speech system,”
H. Javkin, K. Hata, et al., · 1989
Earlier work this paper cites.
“Psychometric properties of the mean opinion scale,”
James R Lewis, · 2001
Earlier work this paper cites.
“BLEU: A method for automatic evaluation of machine translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“A study on multilingual acoustic modeling for large vocabulary ASR,”
Hui Lin, Li Deng, Dong Yu, Yi-fan Gong, Alex Acero, and Chin-Hui Lee, · 2009
Earlier work this paper cites.
“Unsupervised cross-lingual knowledge transfer in DNN-based LVCSR,”
Pawel Swietojanski, Arnab Ghoshal, and Steve Renals, · 2012
Earlier work this paper cites.
“Cross-language knowledge transfer using multilingual deep neural network with shared hidden layers,”
Jui-Ting Huang, Jinyu Li, Dong Yu, Li Deng, and Yifan Gong, · 2013
Earlier work this paper cites.
“Rating naturalness in speech synthesis: The effect of style and expectation,”
Rasmus Dall, Junichi Yamagishi, and Simon King, · 2014
Earlier work this paper cites.
“Are we using enough listeners? no! — an empirically-supported critique of interspeech 2014 TTS evaluations,”
Mirjam Wester, Cassia Valentini-Botinhao, and Gustav Eje Henter, · 2015
Earlier work this paper cites.
“AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech,”
Brian Patton, Yannis Agiomyrgiannakis, et al., · 2016
Earlier work this paper cites.
“Transfer learning for low-resource neural machine translation,”
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight, · 2016
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Skerry-Ryan RJ Wang, Yuxuan et al., · 2017
Earlier work this paper cites.
“Parallel WaveNet: Fast high-fidelity speech synthesis,”
Aaron van den Oord, Yazhe Li, et al., · 2018
Earlier work this paper cites.
“Multi-dialect speech recognition with a single sequence-to-sequence model,”
Bo Li, Tara N Sainath, et al., · 2018
Earlier work this paper cites.
“MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion,”
Chen-Chou Lo, Szu-Wei Fu, et al., · 2019
Earlier work this paper cites.
“Massively multilingual neural machine translation,”
Roee Aharoni, Melvin Johnson, and Orhan Firat, · 2019
Earlier work this paper cites.
“Massively multilingual neural machine translation in the wild: Findings and challenges,”
N. Arivazhagan, Ankur Bapna, Orhan Firat, et al., · 2019
Cited alongside, same era.
“Large-scale multilingual speech recognition with a streaming end-to-end model,”
Anjuli Kannan, Arindrima Datta, et al., · 2019
Cited alongside, same era.
“Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning,”
Yu Zhang, Ron J. Weiss, et al., · 2019
Cited alongside, same era.
“Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges,”
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham, · 2019
Cited alongside, same era.
“COMET: A neural framework for MT evaluation,”
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie, · 2020
Cited alongside, same era.
“BLEURT: Learning robust metrics for text generation,”
“mT5: A massively multilingual pre-trained text-to-text transformer,”
Linting Xue, Noah Constant, et al., · 2021
Later among the works it cites.
“How phonotactics affect multilingual and zero-shot ASR performance,”
Siyuan Feng, Piotr Żelasko, et al., · 2021
Later among the works it cites.
“Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,”
Ryandhimas E Zezario, Szu-Wei Fu, et al., · 2021
Later among the works it cites.
“VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”
Changhan Wang, Morgane Riviere, et al., · 2021
Later among the works it cites.
“Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,”
Wei-Ning Hsu, Anuroop Sriram, et al., · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thibault Sellam, Dipanjan Das, and Ankur Parikh, · 2020
Cited alongside, same era.
“Unsupervised cross-lingual representation learning at scale,”
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, et al., · 2020
Cited alongside, same era.
“Comparison of speech representations for automatic quality estimation in multi-speaker text-to-speech synthesis,”
Jennifer Williams, Joanna Rownicka, Pilar Oplustil, and Simon King, · 2020
Cited alongside, same era.
“Universal phone recognition with a multilingual allophone system,”
Xinjian Li, Siddharth Dalmia, et al., · 2020
Cited alongside, same era.
“That sounds familiar: an analysis of phonetic representations transfer across languages,”
Piotr Żelasko, Laureano Moro-Velázquez, et al., · 2020
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, et al., · 2021
Cited alongside, same era.
“Multilingual Byte2Speech models for scalable low-resource speech synthesis,”
Mutian He, Jingzhou Yang, Lei He, and Frank K Soong, · 2021
Cited alongside, same era.
Wen-Chin Huang, Erica Cooper, Junichi Yamagishi, and Tomoki Toda, · 2022
Closest in time.
“Back to the Future: Extending the Blizzard Challenge 2013,”
Sébastien Le Maguer, Simon King, and Naomi Harte, · 2022
Closest in time.
“Generalization ability of MOS prediction networks,”
Erica Cooper, Wen-Chin Huang, Tomoki Toda, and Junichi Yamagishi, · 2022
Closest in time.
“mSLAM: Massively multilingual joint pre-training for speech and text,”
Ankur Bapna, Colin Cherry, et al., · 2022
Closest in time.
“The VoiceMOS Challenge 2022,”
Wen-Chin Huang, Erica Cooper, Yu Tsao, Hsin-Min Wang, Tomoki Toda, and Junichi Yamagishi, · 2022
Closest in time.
“A comparison of deep learning mos predictors for speech synthesis quality,”
Alessandro Ragano, Emmanouil Benetos, et al., · 2022
Closest in time.
“A Transfer and Multi-Task Learning based Approach for MOS Prediction,”
Xiaohai Tian, Kaiqi Fu, Shaojun Gao, Yiwei Gu, Kai Wang, Wei Li, and Zejun Ma, · 2022
Closest in time.
“Unify and conquer: How phonetic feature representation affects polyglot text-to-speech (TTS),”
Ariadna Sanchez, Alessio Falai, et al., · 2022
Closest in time.
“UTMOS: Utokyo-sarulab system for voiceMOS Challenge 2022,”
Takaaki Saeki, Detai Xin, et al., · 2022
Closest in time.