Fetching the paper…
Reading the bibliography…
Adapting a neural text-to-speech (TTS) model to a target speaker typically involves fine-tuning most if not all of the parameters of a pretrained multi-speaker backbone model.
“The Aligner: Text to speech alignment using markov models and a pronunciation dictionary,”
D. Talkin and C. W. Wightman, · 1994
Earlier work this paper cites.
“A robust algorithm for pitch tracking (rapt),”
D. Talkin and B. Kleijn, · 1995
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez, · 2014
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
S. Ioffe and C. Szegedy, · 2015
Earlier work this paper cites.
“The Kestrel TTS text normalization system,”
P. Ebden and R. Sproat, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, · 2016
Earlier work this paper cites.
“End-to-end text-dependent speaker verification,”
G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, · 2016
Earlier work this paper cites.
“Learning multiple visual domains with residual adapters,”
S.-A. Rebuffi, H. Bilen, and A. Vedaldi, · 2017
Earlier work this paper cites.
“ImageNet classification with deep convolutional neural networks,”
A. Krizhevsky, I. Sutskever, and G. E. Hinton, · 2017
Earlier work this paper cites.
“SGDR: Stochastic gradient descent with warm restarts,”
I. Loshchilov and F. Hutter, · 2017
Earlier work this paper cites.
“Neural voice cloning with a few samples,”
S. O. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, · 2018
Earlier work this paper cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno, and Y. Wu, · 2018
Earlier work this paper cites.
“Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, RJ Skerrv-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, · 2018
Earlier work this paper cites.
“Efficient neural audio synthesis,”
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, · 2018
Cited alongside, same era.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
T. Kudo and J. Richardson, · 2018
Cited alongside, same era.
“Sample efficient adaptive text-to-speech,”
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie, C. Gulcehre, A. van den Oord, O. Vinyals, and N. de Freitas, · 2019
Cited alongside, same era.
“Parameter-efficient transfer learning for NLP,”
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, · 2019
Cited alongside, same era.
“Simple, scalable adaptation for neural machine translation,”
A. Bapna and O. Firat, · 2019
Cited alongside, same era.
“BOFFIN TTS: Few-shot speaker adaptation by Bayesian optimization,”
“Continual speaker adaptation for text-to-speech synthesis,”
H. Hemati and D. Borth, · 2021
Later among the works it cites.
“Adapting TTS models for new speakers using transfer learning,”
P. Neekhara, J. Li, and B. Ginsburg, · 2021
Later among the works it cites.
“AdaSpeech 2: Adaptive text to speech with untranscribed data,”
Y. Yan, X. Tan, B. Li, T. Qin, S. Zhao, Y. Shen, and T.-Y. Liu, · 2021
Later among the works it cites.
“Lightweight adapter tuning for multilingual speech translation,”
H. Le, J. Pino, C. Wang, J. Gu, D. Schwab, and L. Besacier, · 2021
Later among the works it cites.
“Residual adapters for parameter-efficient ASR adaptation to atypical and accented speech,”
K. Tomanek, V. Zayats, D. Padfield, K. Vaillancourt, and F. Biadsy, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. B. Moss, V. Aggarwal, N. Prateek, J. González, and R. Barra-Chicote, · 2020
Cited alongside, same era.
“AdaDurIAN: Few-shot adaptation for neural text-to-speech with DurIAN,”
Z. Zhang, Q. Tian, H. Lu, L.-H. Chen, and S. Liu, · 2020
Cited alongside, same era.
“Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,”
E. Cooper, C.-I Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, · 2020
Cited alongside, same era.
“A study of residual adapters for multi-domain neural machine translation,”
M. Q. Pham, J.-M. Crego, F. Yvon, and J. Senellart, · 2020
Cited alongside, same era.
“Continual learning in task-oriented dialogue systems,”
A. Madotto, Z. Lin, Z. Zhou, S. Moon, P. Crook, B. Liu, Z. Yu, E. Cho, and Z. Wang, · 2020
Cited alongside, same era.
J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, · 2020
Cited alongside, same era.
Later among the works it cites.
“PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,”
Y. Jia, H. Zen, J. Shen, Y. Zhang, and Y. Wu, · 2021
Later among the works it cites.
“DelightfulTTS: The Microsoft speech synthesis system for Blizzard Challenge 2021,”
Y. Liu, Z. Xu, G. Wang, K. Chen, B. Li, X. Tan, J. Li, L. He, and S. Zhao, · 2021
Later among the works it cites.
“FastSpeech 2: Fast and high-quality end-to-end text-to-speech,”
Y. Ren, C. Hu, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2021
Later among the works it cites.
“FastPitch: Parallel text-to-speech with pitch prediction,”
L. Adrian, · 2021
Later among the works it cites.
“A scalable model specialization framework for training and inference using Submodels and its application to speech model personalization,”
F. Biadsy, Y. Chen, X. Zhang, O. Rybakov, A. Rosenberg, and P. J. Moreno, · 2022
Closest in time.
“Voice Filter: Few-shot text-to-speech speaker adaptation using voice conversion as a post-processing module,”
A. Gabryś, G. Huybrechts, M. S. Ribeiro, C.-M. Chien, J. Roth, G. Comini, R. Barra-Chicote, B. Perz, and J. Lorenzo-Trueba, · 2022
Closest in time.
“YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,”
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, · 2022
Closest in time.
“AdaSpeech 4: Adaptive text to speech in zero-shot scenarios,”
Y. Wu, X. Tan, B. Li, L. He, S. Zhao, R. Song, T. Qin, and T.-Y. Liu, · 2022
Closest in time.