Fetching the paper…
Reading the bibliography…
Personalizing a speech synthesis system is a highly desired application, where the system can generate speech with the user's voice with rare enrolled recordings.
E. Variani, X. Lei, E. McDermott, I. L. Moreno, and J. Gonzalez-Dominguez, “Deep neural networks for small footprint text-dependent speaker verification,” in 2014 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2014, pp. 4052–4056
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
G. Heigold, I. Moreno, S. Bengio, and N. Shazeer, “End-to-end text-dependent speaker verification,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 5115–5119
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al. , “Tacotron: Towards end-to-end speech synthesis,” Proc. Interspeech 2017 , pp. 4006–4010, 2017
2017
Earlier work this paper cites.
A. Gibiansky, S. Ö. Arik, G. F. Diamos, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, “Deep voice 2: Multi-speaker neural text-to-speech.” in NIPS , 2017, pp. 2966–2974
2017
Earlier work this paper cites.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International Conference on Machine Learning . PMLR, 2017, pp. 1126–1135
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” Proc. Interspeech 2017 , pp. 2616–2620, 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2017
Earlier work this paper cites.
W. Ping, K. Peng, A. Gibiansky, S. Ömer Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep voice 3: Scaling text-to-speech with convolutional sequence learning.” in ICLR (Poster) , 2018. [Online]. Available: https://openreview.net/forum?id=HJtEm4p6Z
2018
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , 2018, pp. 4485–4495
2018
Earlier work this paper cites.
S. Ö. Arık, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , 2018, pp. 10 040–10 050
2018
Earlier work this paper cites.
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani, “VoiceLoop: Voice fitting and synthesis via a phonological loop,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=SkFAWax0-
2018
Cited alongside, same era.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 4879–4883
2018
Cited alongside, same era.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” Proc. Interspeech 2018 , pp. 1086–1090, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
T. Wang, J. Tao, R. Fu, J. Yi, Z. Wen, and C. Qiang, “Bi-level speaker supervision for one-shot speech synthesis,” Interspeech , pp. 3989–3993, 2020
2020
Later among the works it cites.
Z. Cai, C. Zhang, and M. Li, “From speaker verification to multispeaker speech synthesis, deep transfer with feedback constraint,” Interspeech , pp. 3974–3978, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
A. Raghu, M. Raghu, S. Bengio, and O. Vinyals, “Rapid learning or feature reuse? towards understanding the effectiveness of maml,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=rkgMkCEtPB
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Chen, Y. Assael, B. Shillingford, D. Budden, S. Reed, H. Zen, Q. Wang, L. C. Cobo, A. Trask, B. Laurie, C. Gulcehre, A. van den Oord, O. Vinyals, and N. de Freitas, “Sample efficient adaptive text-to-speech,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=rkzjUoAcFX
2019
Cited alongside, same era.
C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, “Mosnet: Deep learning-based objective assessment for voice conversion,” Proc. Interspeech 2019 , pp. 1541–1545, 2019
2019
Cited alongside, same era.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” Proc. Interspeech 2019 , pp. 1526–1530, 2019
2019
Cited alongside, same era.
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92), [sound],” 2019. [Online]. Available: https://doi.org/10.7488/ds/2645
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “MelGAN: Generative adversarial networks for conditional waveform synthesis,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, “Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 6184–6188
2020
Cited alongside, same era.
T. Wang, J. Tao, R. Fu, J. Yi, Z. Wen, and R. Zhong, “Spoken content and voice factorization for few-shot speaker adaptation,” Interspeech , pp. 796–800, 2020
2020
Cited alongside, same era.
S. Choi, S. Han, D. Kim, and S. Ha, “Attentron: Few-shot text-to-speech utilizing attention-based variable-length embedding,” Interspeech , pp. 2007–2011, 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Later among the works it cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=piLPYqxtWuA
2021
Closest in time.
M. Chen, X. Tan, B. Li, Y. Liu, T. Qin, sheng zhao, and T.-Y. Liu, “Adaspeech: Adaptive text to speech for custom voice,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=Drynvt7gg4L
2021
Closest in time.
W. Song, X. Yuan, Z. Zhang, C. Zhang, Y. Wu, X. He, and B. Zhou, “Dian: Duration informed auto-regressive network for voice cloning,” ICASSP , 2021
2021
Closest in time.
C.-M. Chien, J.-H. Lin, C. yu Huang, P. chun Hsu, and H. yi Lee, “Investigating on incorporating pretrained and learnable speaker representations for multi-speaker multi-style text-to-speech,” ICASSP , 2021
2021
Closest in time.
Y. Leng, X. Tan, S. Zhao, F. Soong, X.-Y. Li, and T. Qin, “Mbnet: Mos prediction for synthesized speech with mean-bias network,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 391–395
2021
Closest in time.
W.-C. Tseng, C. yu Huang, W.-T. Kao, Y. Y. Lin, and H. yi Lee, “Utilizing Self-Supervised Representations for MOS Prediction,” in Proc. Interspeech 2021 , 2021, pp. 2781–2785
2021
Closest in time.
A. T. Liu, S.-W. Li, and H.-y. Lee, “Tera: Self-supervised learning of transformer encoder representation for speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2351–2366, 2021
2021
Closest in time.