Fetching the paper…
Reading the bibliography…
Large language models (LLM)-based speech synthesis has been widely adopted in zero-shot speech synthesis.
Yet another algorithm for pitch tracking
K. Kasi and S. A. Zahorian · 2002
Earlier work this paper cites.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Earlier work this paper cites.
Superseded-CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit
C. Veaux, J. Yamagishi, K. MacDonald, et al · 2017
Earlier work this paper cites.
Tacotron: Towards End-to-End Speech Synthesis
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous · 2017
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
J. S. Chung, A. Nagrani, and A. Zisserman · 2018
Earlier work this paper cites.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu, et al · 2018
Earlier work this paper cites.
Towards end-to-end prosody transfer for expressive speech synthesis with tacotron
R. Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous · 2018
Earlier work this paper cites.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous · 2018
Earlier work this paper cites.
Robust and fine-grained prosody control of end-to-end speech synthesis
Y. Lee and T. Kim · 2019
Earlier work this paper cites.
Neural speech synthesis with transformer network
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
AutoVC: Zero-shot voice style transfer with only autoencoder loss
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson · 2019
Earlier work this paper cites.
Fastspeech: Fast, robust and controllable text to speech
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu · 2019
Earlier work this paper cites.
Zero-shot learning—a comprehensive evaluation of the good, the bad and the ugly
Y. Xian, C. H. Lampert, B. Schiele, and Z. Akata · 2019
Earlier work this paper cites.
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Earlier work this paper cites.
In Defence of Metric Learning for Speaker Recognition
J. S. Chung, J. Huh, S. Mun, M. Lee, H.-S. Heo, S. Choe, C. Ham, S. Jung, B.-J. Lee, and I. Han · 2020
Earlier work this paper cites.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux · 2020
Earlier work this paper cites.
Glow-TTS: A generative flow for text-to-speech via monotonic alignment search
J. Kim, S. Kim, J. Kong, and S. Yoon · 2020
Earlier work this paper cites.
HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
J. Kong, J. Kim, and J. Bae · 2020
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu · 2020
Earlier work this paper cites.
Phonemizer: Text to phones transcription for multiple languages in python
M. Bernard and H. Titeux · 2021
Earlier work this paper cites.
Adaspeech: Adaptive text to speech for custom voice
M. Chen, X. Tan, B. Li, Y. Liu, T. Qin, S. Zhao, and T.-Y. Liu · 2021
Earlier work this paper cites.
Reinforce-Aligner: Reinforcement Alignment Search for Robust End-to-End Text-to-Speech
H. Chung, S.-H. Lee, and S.-W. Lee · 2021
Earlier work this paper cites.
Many-to-many voice transformer network
H. Kameoka, W.-C. Huang, K. Tanaka, T. Kaneko, N. Hojo, and T. Toda · 2021
Earlier work this paper cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
J. Kim, J. Kong, and J. Son · 2021
Earlier work this paper cites.
Fre-GAN: Adversarial Frequency-Consistent Audio Synthesis
J.-H. Kim, S.-H. Lee, J.-H. Lee, and S.-W. Lee · 2021
Earlier work this paper cites.
The ins and outs of speaker recognition: lessons from VoxSRC 2020
Y. Kwon, H. S. Heo, B.-J. Lee, and J. S. Chung · 2021
Earlier work this paper cites.
VoiceMixer: Adversarial voice style mixup
S.-H. Lee, J.-H. Kim, H. Chung, and S.-W. Lee · 2021
Earlier work this paper cites.
Multi-spectrogan: High-diversity and high-fidelity spectrogram generation with adversarial style combination for speech synthesis
S.-H. Lee, H.-W. Yoon, H.-R. Noh, J.-H. Kim, and S.-W. Lee · 2021
Earlier work this paper cites.
Any-to-many voice conversion with location-relative sequence-to-sequence modeling
S. Liu, Y. Cao, D. Wang, X. Wu, X. Liu, and H. Meng · 2021
Cited alongside, same era.
Meta-stylespeech: Multi-speaker adaptive text-to-speech generation
D. Min, D. B. Lee, E. Yang, and S. J. Hwang · 2021
Cited alongside, same era.
Grad-TTS: A diffusion probabilistic model for text-to-speech
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi · 2021
Cited alongside, same era.
XLS-R: Self-supervised cross-lingual speech representation learning at scale
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli · 2022
Cited alongside, same era.
YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone
Prompttts: Controllable text-to-speech with text descriptions
Z. Guo, Y. Leng, Y. Wu, S. Zhao, and X. Tan · 2023
Closest in time.
A survey on vision transformer
K. Han, Y. Wang, H. Chen, X. Chen, J. Guo, Z. Liu, Y. Tang, A. Xiao, C. Xu, Y. Xu, Z. Yang, Y. Zhang, and D. Tao · 2023
Closest in time.
Make-a-voice: Unified voice synthesis with discrete representation
R. Huang, C. Zhang, Y. Wang, D. Yang, L. Liu, Z. Ye, Z. Jiang, C. Weng, Z. Zhao, and D. Yu · 2023
Closest in time.
J.-S. Hwang, S.-H. Lee, and S.-W. Lee · 2023
Closest in time.
Mega-tts 2: Zero-shot text-to-speech with arbitrary length speech prompts
Z. Jiang, J. Liu, Y. Ren, J. He, C. Zhang, Z. Ye, P. Wei, C. Wang, X. Yin, Z. Ma, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti · 2022
Cited alongside, same era.
Nansy++: Unified voice synthesis with neural analysis and synthesis
H.-S. Choi, J. Yang, J. Lee, and H. Kim · 2022
Cited alongside, same era.
High fidelity neural audio compression
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi · 2022
Cited alongside, same era.
Vqtts: High-fidelity text-to-speech synthesis with self-supervised vq acoustic feature
C. Du, Y. Guo, X. Chen, and K. Yu · 2022
Cited alongside, same era.
NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates
S. Han and J. Lee · 2022
Cited alongside, same era.
Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech
R. Huang, Y. Ren, J. Liu, C. Cui, and Z. Zhao · 2022
Cited alongside, same era.
Guided-TTS: A diffusion model for text-to-speech via classifier guidance
H. Kim, S. Kim, and S. Yoon · 2022
Cited alongside, same era.
Closest in time.
Grad-stylespeech: Any-speaker adaptive text-to-speech synthesis with diffusion models
M. Kang, D. Min, and S. J. Hwang · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour · 2023
Closest in time.
Unitspeech: Speaker-adaptive speech synthesis with untranscribed data
H. Kim, S. Kim, J. Yeom, and S. Yoon · 2023
Closest in time.
P-flow: A fast and data-efficient zero-shot TTS through speech prompting
S. Kim, K. J. Shih, R. Badlani, J. F. Santos, E. Bakhturina, M. T. Desta, R. Valle, S. Yoon, and B. Catanzaro · 2023
Closest in time.
High-fidelity audio compression with improved RVQGAN
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar · 2023
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, and W.-N. Hsu · 2023
Closest in time.
HierVST: Hierarchical Adaptive Zero-shot Voice Style Transfer
S.-H. Lee, H.-Y. Choi, H.-S. Oh, and S.-W. Lee · 2023
Closest in time.
Prompttts 2: Describing and generating voices with text prompt
Y. Leng, Z. Guo, K. Shen, X. Tan, Z. Ju, Y. Liu, Y. Liu, D. Yang, L. Zhang, K. Song, et al · 2023
Closest in time.
Y. A. Li, C. Han, V. S. Raghavan, G. Mischler, and N. Mesgarani · 2023
Closest in time.
Audiosr: Versatile audio super-resolution at scale
H. Liu, K. Chen, Q. Tian, W. Wang, and M. D. Plumbley · 2023
Closest in time.
MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra
Y.-X. Lu, Y. Ai, and Z.-H. Ling · 2023
Closest in time.
Expresso: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis
T. A. Nguyen, W.-N. Hsu, A. D’Avirro, B. Shi, I. Gat, M. Fazel-Zarani, T. Remez, J. Copet, G. Synnaeve, M. Hassid, F. Kreuk, Y. Adi, and E. Dupoux · 2023
Closest in time.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Closest in time.
Scaling speech technology to 1,000+ languages
V. Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi, et al · 2023
Closest in time.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
K. Shen, Z. Ju, X. Tan, Y. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian · 2023
Closest in time.
Random cycle loss and its application to voice conversion
H. Sun, D. Wang, L. Li, C. Chen, and T. F. Zheng · 2023
Closest in time.
Dawn of the transformer era in speech emotion recognition: Closing the valence gap
J. Wagner, A. Triantafyllopoulos, H. Wierstorf, M. Schmitt, F. Burkhardt, F. Eyben, and B. W. Schuller · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, et al · 2023
Closest in time.
A survey on non-autoregressive generation for neural machine translation and beyond
Y. Xiao, L. Wu, J. Guo, J. Li, M. Zhang, T. Qin, and T.-Y. Liu · 2023
Closest in time.
H. Xue, S. Guo, P. Zhu, and M. Bi · 2023
Closest in time.
Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt
D. Yang, S. Liu, R. Huang, G. Lei, C. Weng, H. Meng, and D. Yu · 2023
Closest in time.
Hifi-codec: Group-residual vector quantization for high fidelity audio codec
D. Yang, S. Liu, R. Huang, J. Tian, C. Weng, and Y. Zou · 2023
Closest in time.
Uniaudio: An audio foundation model toward universal audio generation
D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, S. Zhao, J. Bian, X. Wu, et al · 2023
Closest in time.
Comospeech: One-step speech and singing voice synthesis via consistency model
Z. Ye, W. Xue, X. Tan, J. Chen, Q. Liu, and Y. Guo · 2023
Closest in time.
Conditioning and sampling in variational diffusion models for speech super-resolution
C.-Y. Yu, S.-L. Yeh, G. Fazekas, and H. Tang · 2023
Closest in time.