Fetching the paper…
Reading the bibliography…
While speaker adaptation for end-to-end speech synthesis using speaker embeddings can produce good speaker similarity for speakers seen during training, there remains a gap for zero-shot adaptation to unseen speakers.
J. Park, K. Zhao, K. Peng, and W. Ping, “Multi-speaker end-to-end speech synthesis,”
1907
Earlier work this paper cites.
D. A. Reynolds, T. F. Quatieri, and R. B. Dunn, “Speaker verification using adapted gaussian mixture models,”
2000
Earlier work this paper cites.
A. W. Black and K. A. Lenzo, “Flite: a small fast run-time synthesis engine,” in
2001
Earlier work this paper cites.
International Telecommunication Union, Recommendation G.191: Software Tools and Audio Coding Standardization, Nov 11 2005
2005
Earlier work this paper cites.
N. Brümmer and J. du Preez, “Application-independent evaluation of speaker detection,”
2006
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2010
Earlier work this paper cites.
S. King and V. Karaiskos, “The blizzard challenge 2011,” in
2011
Earlier work this paper cites.
P. Kenny, T. Stafylakis, P. Ouellet, J. Alam, and P. Dumouchel, “PLDA for speaker verification with utterances of arbitrary duration,” in
2013
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” in
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, Z. Y. Jaitly, Y. Xiao, Z. Chen, S. Bengio, Q. Le
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “VoxCeleb: a large-scale speaker identification dataset,”
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” University of Edinburgh, The Centre for Speech Technology Research (CSTR), 2017
2017
Earlier work this paper cites.
W. Ping, K. Peng, and J. Chen, “Clarinet: Parallel wave generation in end-to-end text-to-speech,”
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan
2018
Cited alongside, same era.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,”
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Arık, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in
2018
Cited alongside, same era.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. L. Moreno, Y. Wu
2018
Cited alongside, same era.
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie
2019
Closest in time.
B. Bollepalli, L. Juvela, and P. Alku, “Lombard speech synthesis using transfer learning in a Tacotron text-to-speech system,”
2019
Closest in time.
Q. Hu, E. Marchi, D. Winarsky, Y. Stylianou, D. Naik, and S. Kajarekar, “Neural text-to-speech adaptation from low quality public recordings,”
2019
Closest in time.
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, “Learning problem-agnostic speech representations from multiple self-supervised tasks,”
2019
Closest in time.
M. Chen, M. Chen, S. Liang, J. Ma, L. Chen, S. Wang, and J. Xiao, “Cross-lingual, multi-speaker text-to-speech synthesis using neural speaker embedding,”
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,”
2018
Cited alongside, same era.
W. Cai, J. Chen, and M. Li, “Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,”
2018
Cited alongside, same era.
Z. Huang, S. Wang, and K. Yu, “Angular softmax for short-duration text-independent speaker verification,” in
2018
Cited alongside, same era.
M. Hajibabaei and D. Dai, “Unified hypersphere embedding for speaker recognition,”
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “VoxCeleb2: Deep speaker recognition,”
2018
Cited alongside, same era.
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai, “Forward attention in sequence-to-sequence acoustic modeling for speech synthesis,” in
2018
Cited alongside, same era.
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, “The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,” in
2018
Cited alongside, same era.
J. Villalba, N. Chen, D. Snyder, D. Garcia-Romero, A. McCree, G. Sell, J. Borgstrom, F. Richardson, S. Shon, F. Grondin
2019
Closest in time.
N. Chen, J. Villalba, and N. Dehak, “Tied mixture of factor analyzers layer to combine frame level representations in neural speaker embeddings,”
2019
Closest in time.
W. Xie, A. Nagrani, J. S. Chung, and A. Zisserman, “Utterance-level aggregation for speaker recognition in the wild,”
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
Y. Yasuda, X. Wang, S. Takaki, and J. Yamagishi, “Investigation of enhanced Tacotron text-to-speech synthesis systems with self-attention for pitch accent language,”
2019
Closest in time.
G. Bhattacharya, J. Alam, and P. Kenny, “Deep speaker recognition: Modular or monolithic?”
2019
Closest in time.
J. Taylor and K. Richmond, “Analysis of pronunciation learning in end-to-end speech synthesis,”
2019
Closest in time.