Fetching the paper…
Reading the bibliography…
We propose PromptTTS++, a prompt-based text-to-speech (TTS) synthesis system that allows control over speaker identity using natural language descriptions.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Mixture density networks”
Christopher Bishop · 1994
Earlier work this paper cites.
“Mixture density networks”, 1994
Christopher. Bishop · 1994
Earlier work this paper cites.
“Extraction of everyday expression associated with voice quality of normal utterance”
Hiroshi Kido and Kasuya Hideki · 1999
Earlier work this paper cites.
“Continuous F0 modeling for HMM based statistical parametric speech synthesis”
Kai Yu and Steve Young · 2010
Earlier work this paper cites.
“Rectified linear units improve restricted boltzmann machines”
Vinod Nair and Geoffrey Hinton · 2010
Earlier work this paper cites.
“Text-to-speech technology to control speaker indivisuality with intuitive expressions”
Ohtani Yamato and Koichiro Mori · 2016
Earlier work this paper cites.
“WORLD: a vocoder-based high-quality speech synthesis system for real-time applications”
Masanori Morise, Fumiya Yokomori and Kenji Ozawa · 2016
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using kaldi.”
Michael McAuliffe et al · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions”
Jonathan Shen et al · 2018
Earlier work this paper cites.
“Style tokens: unsupervised style modeling, control and transfer in end-to-end speech synthesis”
Yuxuan Wang et al · 2018
Cited alongside, same era.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
J. Devlin et al · 2019
Cited alongside, same era.
“LibriTTS: A corpus derived from LibriSpeech for text-to-speech”
Heiga Zen et al · 2019
Cited alongside, same era.
“Decoupled weight decay regularization”
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
“Neural source-filter waveform models for statistical parametric speech synthesis”
Xin Wang, Shinji Takaki and Junichi Yamagishi · 2019
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition”
Anmol Gulati et al · 2020
Cited alongside, same era.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Later among the works it cites.
“Speaker generation”
Daisy Stanton et al · 2022
Later among the works it cites.
“DiffSinger: Singing voice synthesis via shallow diffusion mechanism”
Jinglin Liu et al · 2022
Later among the works it cites.
“PromptTTS: Controllable text-to-speech with text descriptions”
Zhifang Guo et al · 2023
Closest in time.
“InstructTTS: Modelling expressive tts in discrete latent space with natural language style prompt”
Dongchao Yang et al · 2023
Closest in time.
“PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Denoising diffusion probabilistic models”
Jonathan Ho, Ajay Jain and Pieter Abbeel · 2020
Cited alongside, same era.
“A survey on neural speech synthesis”
Xu Tan et al · 2021
Cited alongside, same era.
“FastSpeech 2: Fast and High-Quality End-to-End Text-to-Speech”
Yi Ren et al · 2021
Cited alongside, same era.
“Phone-level prosody modelling with GMM-based MDN for diverse and controllable speech synthesis”
Chenpeng Du and Kai Yu · 2021
Cited alongside, same era.
“Naturalspeech: End-to-end text to speech synthesis with human-level quality”
Xu Tan et al · 2022
Cited alongside, same era.
Guanghou Liu et al · 2023
Closest in time.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron et al · 2023
Closest in time.
“LibriTTS-R: A restored multi-speaker text-to-speech corpus”
Yuma Koizumi et al · 2023
Closest in time.
Yuma Koizumi et al · 2023
Closest in time.
“BigVGAN: A Universal Neural Vocoder with Large-Scale Training”
Sang-gil Lee et al · 2023
Closest in time.