Fetching the paper…
Reading the bibliography…
Fine-tuning is a popular method for adapting text-to-speech (TTS) models to new speakers.
“pyin: A fundamental frequency estimator using probabilistic threshold distributions,”
Matthias Mauch and Simon Dixon, · 2014
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, et al., · 2017
Earlier work this paper cites.
“Deep voice 3: 2000-speaker neural text-to-speech.,”
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan Ömer Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller, · 2017
Earlier work this paper cites.
“The LJ speech dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017
Keith Ito and Linda Johnson, · 2017
Earlier work this paper cites.
“Neural voice cloning with a few samples,”
Sercan Arik, Jitong Chen, Kainan Peng, et al., · 2018
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, et al., · 2018
Earlier work this paper cites.
“FastSpeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Earlier work this paper cites.
“LibriTTS: A corpus derived from librispeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Earlier work this paper cites.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (ver. 0.92),” 2019
Junichi Yamagishi, Christophe Veaux, Kirsten MacDonald, et al., · 2019
Cited alongside, same era.
“Parameter-efficient transfer learning for nlp,”
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly, · 2019
Cited alongside, same era.
“FastSpeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Cited alongside, same era.
“Boffin TTS: Few-shot speaker adaptation by bayesian optimization,”
Henry B Moss, Vatsal Aggarwal, Nishant Prateek, et al., · 2020
Cited alongside, same era.
“AdaDurian: Few-shot adaptation for neural text-to-speech with durian,”
Zewang Zhang, Qiao Tian, Heng Lu, et al., · 2020
Cited alongside, same era.
“AdaSpeech: Adaptive text to speech for custom voice,”
Mingjian Chen, Xu Tan, Bohan Li, et al., · 2021
Later among the works it cites.
“Lora: Low-rank adaptation of large language models,”
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen, · 2021
Later among the works it cites.
“Prefix-tuning: Optimizing continuous prompts for generation,”
Xiang Lisa Li and Percy Liang, · 2021
Later among the works it cites.
“Towards a unified view of parameter-efficient transfer learning,”
Junxian He, Chunting Zhou, Xuezhe Ma, et al., · 2021
Later among the works it cites.
“BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“FastPitch: Parallel text-to-speech with pitch prediction,”
Adrian Lańcucki, · 2021
Cited alongside, same era.
“Hi-Fi multi-speaker English TTS dataset,”
Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg, and Yang Zhang, · 2021
Cited alongside, same era.
Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg, · 2021
Later among the works it cites.
“Voice filter: Few-shot text-to-speech speaker adaptation using voice conversion as a post-processing module,”
Adam Gabryś, Goeric Huybrechts, Manuel Sam Ribeiro, et al., · 2022
Closest in time.
“One TTS alignment to rule them all,”
Rohan Badlani, Adrian Lańcucki, Kevin J Shih, et al., · 2022
Closest in time.
“TitaNet: Neural model for speaker representation with 1D depth-wise separable convolutions and global context,”
Nithin Rao Koluguri, Taejin Park, and Boris Ginsburg, · 2022
Closest in time.