Fetching the paper…
Reading the bibliography…
We aim to characterize how different speakers contribute to the perceived output quality of multi-speaker Text-to-Speech (TTS) synthesis.
“MOSNet: Deep Learning based Objective Assessment for Voice Conversion,”
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang, · 1904
Earlier work this paper cites.
“A Computer Method for Calculating Kendall’s tau With Ungrouped Data,”
William R Knight, · 1966
Earlier work this paper cites.
“Digital Selection and Analogue Amplification Coexist in a Cortex-Inspired Silicon Circuit,”
Richard HR Hahnloser, Rahul Sarpeshkar, Misha A Mahowald, Rodney J Douglas, and H Sebastian Seung, · 2000
Earlier work this paper cites.
“Aurora Working group: DSR Front End LVCSR evaluation AU/384/02,”
Naveen Parihar and Joseph Picone, · 2002
Earlier work this paper cites.
“Feature Selection, L1 vs. L2 Regularization, and Rotational Invariance,”
Andrew Y Ng, · 2004
Earlier work this paper cites.
“Recognition and Understanding of Meetings: The AMI and AMIDA projects,”
S. Renals, T. Hain, and H. Bourlard, · 2007
Earlier work this paper cites.
“Influence Functions of the Spearman and Kendall Correlation Measures,”
Christophe Croux and Catherine Dehon, · 2010
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Keras,”
François Chollet et al., · 2015
Earlier work this paper cites.
“AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech,”
Brian Patton, Yannis Agiomyrgiannakis, Michael Terry, Kevin Wilson, Rif A. Saurous, and D. Sculley, · 2016
Cited alongside, same era.
“Tensorflow: A System for Large-Scale Machine Learning,”
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al., · 2016
Cited alongside, same era.
“Deep Voice 2: Multi-Speaker Neural Text-to-Speech,”
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou, · 2017
Cited alongside, same era.
“An Image-based Deep Spectrum Feature Representation for the Recognition of Emotional Speech,”
Nicholas Cummins, Shahin Amiriparian, Gerhard Hagerer, Anton Batliner, Stefan Steidl, and Björn W. Schuller, · 2017
Cited alongside, same era.
“Snore Sound Classification Using Image-Based Deep Spectrum Features,”
Shahin Amiriparian, Maurice Gerczuk, Sandra Ottl, Nicholas Cummins, Michael Freitag, Sergey Pugachevskiy, Alice Baird, and Björn Schuller, · 2017
“webMUSHRA — A Comprehensive Framework for Web-Based Listening Tests,”
Michael Schoeffler, Sarah Bartoschek, Fabian-Robert Stöter, Marlene Roess, Susanne Westphal, Bernd Edler, and Jürgen Herre, · 2018
Later among the works it cites.
“Libritts: A Corpus Derived from LibriSpeech for Text-to-Speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Later among the works it cites.
“Speech Synthesis Evaluation — State-of-the-Art Assessment and Suggestion for a Novel Research Program,”
Petra Wagner, Jonas Beskow, Simon Betz, Jens Edlund, Joakim Gustafson, Gustav Eje Henter, Sébastien Le Maguer, Zofia Malisz, Éva Székely, Christina Tånnander, and Jana Voße, · 2019
Later among the works it cites.
“Quality Degradation Diagnosis for Voice Networks — Estimating the Perceived Noisiness, Coloration, and Discontinuity of Transmitted Speech,”
Gabriel Mittag and Sebastian Möller, · 2019
Later among the works it cites.
“Speech Replay Detection with x-Vector Attack Embeddings and Spectral Features,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model Based on BLSTM,”
Szu-wei Fu, Yu Tsao, Hsin-Te Hwang, and Hsin-Min Wang, · 2018
Cited alongside, same era.
“Analyzing deep CNN-based utterance embeddings for acoustic model adaptation,”
Joanna Rownicka, Peter Bell, and Steve Renals, · 2018
Cited alongside, same era.
“X-Vectors: Robust DNN Embeddings for Speaker Recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention,”
Hideyuki Tachibana, Katsuya Uenoyama, and Shunsuke Aihara, · 2018
Cited alongside, same era.
“Ophelia,”
Oliver Watts,
Cited in the paper.
Jennifer Williams and Joanna Rownicka, · 2019
Later among the works it cites.
“Automatic Speaker Verification Spoofing and Countermeasures Challenge 2019,”
Junichi Yamagishi, Massimiliano Todisco, Md Sahidullah, Hector Delgado, Xin Wang, Nicholas Evans, Tomi Kinnunen, Kong Aik Lee, Ville Vestman, and Andreas Nautsch, · 2019
Later among the works it cites.
Xin Wang, Junichi Yamagishi, Massimiliano Todisco, Hector Delgado, Andreas Nautsch, Nicholas Evans, Md Sahidullah, Ville Vestman, Tomi Kinnunen, Kong Aik Lee, et al., · 2019
Later among the works it cites.
“Embeddings for DNN Speaker Adaptive Training,” 2019
Joanna Rownicka, Peter Bell, and Steve Renals, · 2019
Later among the works it cites.
“Where Do the Improvements Come From in Sequence-to-Sequence Neural TTS?,”
Oliver Watts, Gustav Eje Henter, Jason Fong, and Cassia Valentini-Botinhao, · 2019
Later among the works it cites.