Fetching the paper…
Reading the bibliography…
Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dynamic prosody information in the reference audio is difficult to extract, especially, when the reference audio contains background noise.
“Deep unsupervised learning using nonequilibrium thermodynamics,”
J. Sohl, E. Weiss, et al., · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, et al., · 2015
Earlier work this paper cites.
“Neural discrete representation learning,”
A. Van Den Oord, O. Vinyals, et al., · 2017
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, et al., · 2018
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Y. Wang, D. Stanton, Y. Zhang, et al., · 2018
Earlier work this paper cites.
“Additive margin softmax for face verification,”
F. Wang, J. Cheng, W. Liu, and H. Liu, · 2018
Earlier work this paper cites.
“A multi-device dataset for urban acoustic scene classification,”
A. Mesaros, T. Heittola, and T. Virtanen, · 2018
Earlier work this paper cites.
“Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factorization,”
W. Hsu, Y. Zhang, et al., · 2019
Earlier work this paper cites.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, et al., · 2020
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
J. Ho, A. Jain, and P. Abbeel, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, et al., · 2020
Cited alongside, same era.
“Unsupervised speech decomposition via triple information bottleneck,”
K. Qian, Y. Zhang, S. Chang, et al., · 2020
Cited alongside, same era.
“Diffwave: A versatile diffusion model for audio synthesis,”
Z. Kong, W. Ping, J. Huang, et al., · 2020
Cited alongside, same era.
“Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
J. Kong, J. Kim, and J. Bae, · 2020
Cited alongside, same era.
“Neural analysis and synthesis: Reconstructing speech from self-supervised representations,”
H. Choi, J. Lee, et al., · 2021
Later among the works it cites.
“Diffusion models beat gans on image synthesis,”
P. Dhariwal and A. Nichol, · 2021
Later among the works it cites.
“Generspeech: Towards style transfer for generalizable out-of-domain text-to-speech synthesis,”
R. Huang, Y. Ren, J. Liu, C. Cui, and Z. Zhao, · 2022
Closest in time.
“Hifidenoise: High-fidelity denoising text to speech with adversarial networks,”
L. Zhang, Y. Ren, et al., · 2022
Closest in time.
“SATTS: Speaker attractor text to speech, learning to speak by learning to separate,”
N. Goswami and T. Harada, · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Li, C. Song, J. Li, Z. Wu, J. Jia, and H. Meng, · 2021
Cited alongside, same era.
“Meta-stylespeech: Multi-speaker adaptive text-to-speech generation,”
D. Min, D. Lee, E. Yang, and S. Hwang, · 2021
Cited alongside, same era.
K. Lee, K. Park, and D. Kim, · 2021
Cited alongside, same era.
K. Nikitaras, G. Vamvoukakis, N. Ellinas, et al., · 2022
Closest in time.
“Conditional diffusion probabilistic model for speech enhancement,”
Y. Lu, Z. Wang, et al., · 2022
Closest in time.
“Diffsound: Discrete diffusion model for text-to-sound generation,”
D. Yang, J. Yu, et al., · 2022
Closest in time.