Fetching the paper…
Reading the bibliography…
We introduce OpenVoice, a versatile voice cloning approach that requires only a short audio clip from the reference speaker to replicate their voice and generate speech in multiple languages.
Handbook of the International Phonetic Association: A guide to the use of the International Phonetic Alphabet
I. P. Association · 1999
Earlier work this paper cites.
Dynamic time warping
M. Müller · 2007
Earlier work this paper cites.
Dynamic time warping algorithm review
P. Senin · 2008
Earlier work this paper cites.
Variational inference with normalizing flows
D. Rezende and S. Mohamed · 2015
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
J. Kim, S. Kim, J. Kong, and S. Yoon · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
J. Kong, J. Kim, and J. Bae · 2020
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed · 2021
Cited alongside, same era.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
J. Kim, J. Kong, and J. Son · 2021
Cited alongside, same era.
Speech resynthesis from discrete disentangled self-supervised representations
A. Polyak, Y. Adi, J. Copet, E. Kharitonov, K. Lakhotia, W.-N. Hsu, A. Mohamed, and E. Dupoux · 2021
Cited alongside, same era.
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti · 2022
Cited alongside, same era.
Xtts taking text-to-speech to the next level
CoquiAI · 2023
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, et al · 2023
Closest in time.
Freevc: Towards high-quality text-free one-shot voice conversion
J. Li, W. Tu, and L. Xiao · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, et al · 2023
Closest in time.
Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt
D. Yang, S. Liu, R. Huang, G. Lei, C. Weng, H. Meng, and D. Yu · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A comparison of discrete and soft speech units for improved voice conversion
B. van Niekerk, M.-A. Carbonneau, J. Zaïdi, M. Baas, H. Seuté, and H. Kamper · 2022
Cited alongside, same era.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Z. Zhang, L. Zhou, C. Wang, S. Chen, Y. Wu, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, et al · 2023
Closest in time.