Fetching the paper…
Reading the bibliography…
The zero-shot scenario for speech generation aims at synthesizing a novel unseen voice with only one utterance of the target speaker.
L. Van der Maaten and G. Hinton, “Visualizing data using t-sne.” Journal of machine learning research , vol. 9, no. 11, pp. 2579–2605, 2008
2008
Earlier work this paper cites.
S. Choi, S. Han, D. Kim, and S. Ha, “Attentron: Few-shot text-to-speech utilizing attention-based variable-length embedding,” in The Annual Conference of the International Speech Communication Association (Interspeech) , 2020, pp. 2007–2011
2011
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in the International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A Generative Model for Raw Audio,” in Proc. 9th ISCA Workshop on Speech Synthesis Workshop (SSW 9) , 2016, p. 125
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” 2016
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang et al. , “Tacotron: Towards end-to-end speech synthesis,” in The Annual Conference of the International Speech Communication Association (Interspeech) , 2017, pp. 4006–4010
2017
Earlier work this paper cites.
S. Ö. Arık, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman et al. , “Deep voice: Real-time neural text-to-speech,” in the International Conference on Machine Learning (ICML) . PMLR, 2017, pp. 195–204
2017
Earlier work this paper cites.
A. Gibiansky, S. Arik, G. Diamos, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, “Deep voice 2: Multi-speaker neural text-to-speech,” Advances in neural information processing systems , vol. 30, pp. 2962–2970, 2017
2017
Earlier work this paper cites.
H.-T. Luong, S. Takaki, G. E. Henter, and J. Yamagishi, “Adapting and controlling dnn-based speech synthesis using input codes,” in the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2017, pp. 4905–4909
2017
Earlier work this paper cites.
S. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani, “Voiceloop: Voice fitting and synthesis via a phonological loop,” in the International Conference on Learning Representations (ICLR) , 2018
2018
Cited alongside, same era.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,” in the International Conference on Machine Learning (ICML) . PMLR, 2018, pp. 3683–3691
2018
Cited alongside, same era.
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in The International Conference on Neural Information Processing Systems (NeurIPS) , 2018, pp. 4485–4495
2018
Cited alongside, same era.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2018, pp. 5329–5333
O. Barbany and M. Cernak, “Fastvc: Fast voice conversion with non-parallel data,” Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020 , pp. 145––149, 2020
2020
Later among the works it cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Later among the works it cites.
E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, “Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,” in the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2020, pp. 6184–6188
2020
Later among the works it cites.
H.-T. Luong and J. Yamagishi, “Nautilus: a versatile voice cloning system,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2967–2981, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2018, pp. 4879–4883
2018
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: fast, robust and controllable text to speech,” in The International Conference on Neural Information Processing Systems (NeurIPS) , 2019, pp. 3171–3180
2019
Cited alongside, same era.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2019, pp. 3617–3621
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “Melgan: Generative adversarial networks for conditional waveform synthesis,” Advances in neural information processing systems , vol. 32, pp. 14 910–14 921, 2019
2019
Cited alongside, same era.
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, “Learning problem-agnostic speech representations from multiple self-supervised tasks,” in The Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 161–165
2019
Cited alongside, same era.
J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz, R. Barra-Chicote, A. Moinet, and V. Aggarwal, “Towards achieving robust universal neural vocoding,” in the Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 181–185
2019
Cited alongside, same era.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “Libritts: A corpus derived from librispeech for text-to-speech,” in The Annual Conference of the International Speech Communication Association (Interspeech) , 2019, pp. 1526–1530
2019
Cited alongside, same era.
2020
Later among the works it cites.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-tts: A generative flow for text-to-speech via monotonic alignment search,” Advances in Neural Information Processing Systems , vol. 33, pp. 8067–8077, 2020
2020
Later among the works it cites.
E. Casanova, C. Shulby, E. Gölge, N. M. Müller, F. S. de Oliveira, A. Candido Jr., A. da Silva Soares, S. M. Aluisio, and M. A. Ponti, “SC-GlowTTS: An Efficient Zero-Shot Multi-Speaker Text-To-Speech Model,” in The Annual Conference of the International Speech Communication Association (Interspeech) , 2021, pp. 3645–3649
2021
Later among the works it cites.
Y. Yan, X. Tan, B. Li, T. Qin, S. Zhao, Y. Shen, and T.-Y. Liu, “Adaspeech 2: Adaptive text to speech with untranscribed data,” in the International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2021, pp. 6613–6617
2021
Later among the works it cites.
M. Chen, X. Tan, B. Li, Y. Liu, and T. Qin, “sheng zhao, and tie-yan liu. adaspeech: Adaptive text to speech for custom voice,” in the International Conference on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in the International Conference on Machine Learning (ICML) . PMLR, 2021, pp. 5530–5540
2021
Later among the works it cites.
J. Cong, S. Yang, L. Xie, and D. Su, “Glow-WaveGAN: Learning Speech Representations from GAN-Based Variational Auto-Encoder for High Fidelity Flow-Based Speech Synthesis,” in The Annual Conference of the International Speech Communication Association (Interspeech) , 2021, pp. 2182–2186
2021
Later among the works it cites.