Fetching the paper…
Reading the bibliography…
We present EdiTTS, an off-the-shelf speech editing methodology based on score-based generative modeling for text-to-speech synthesis.
B. D. Anderson, “Reverse-time diffusion equation models,” Stochastic Processes and their Applications , vol. 12, no. 3, pp. 313–326, 1982
1982
Earlier work this paper cites.
É. Moulines and F. Charpentier, “Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,” in Speech Commun. , 1989
1989
Earlier work this paper cites.
S. Bird, E. Klein, and E. Loper, Natural Language Processing with Python , 1st ed. O’Reilly Media, Inc., 2009
2009
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Transactions on Information and Systems , vol. E99.D, no. 7, 2016
2016
Earlier work this paper cites.
Z. Jin, G. J. Mysore, S. Diverdi, J. Lu, and A. Finkelstein, “VoCo: Text-based insertion and replacement in audio narration,” ACM Transactions on Graphics , vol. 36, no. 4, 2017. [Online]. Available: https://doi.org/10.1145/3072959.3073702
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJ speech dataset,” 2017
2017
Earlier work this paper cites.
J. W. Kim, J. Salamon, P. Li, and J. P. Bello, “CREPE: A convolutional representation for pitch estimation,” in ICASSP , 2018
2018
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in NeurIPS , 2019
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in NeurIPS , 2020
2020
Earlier work this paper cites.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,” in NeurIPS , 2020
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in NeurIPS , 2020
2020
Cited alongside, same era.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” 2021
2021
Cited alongside, same era.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in ICLR , 2021
2021
Cited alongside, same era.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” in ICLR , 2021
2021
Cited alongside, same era.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, M. Kudinov, and J. Wei, “Diffusion-based voice conversion with fast maximum likelihood sampling scheme,” 2021
2021
Closest in time.
C. Meng, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon, “SDEdit: Image synthesis and editing with stochastic differential equations,” 2021
2021
Closest in time.
J. Choi, S. Kim, Y. Jeong, Y. Gwon, and S. Yoon, “ILVR: Conditioning method for denoising diffusion probabilistic models,” in ICCV , 2021
2021
Closest in time.
M. Morrison, L. Rencker, Z. Jin, N. J. Bryan, J.-P. Caceres, and B. Pardo, “Context-aware prosody correction for text-based speech editing,” in ICASSP , 2021
2021
Closest in time.
D. Tan, L. Deng, Y. T. Yeung, X. Jiang, X. Chen, and T. Lee, “EditSpeech: A text based speech editing system using partial inference and bidirectional fusion,” in ASRU , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” in ICLR , 2021
2021
Cited alongside, same era.
M. Jeong, H. Kim, S. J. Cheon, B. J. Choi, and N. S. Kim, “Diff-TTS: A denoising diffusion model for text-to-speech,” in INTERSPEECH , 2021
2021
Cited alongside, same era.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-TTS: A diffusion probabilistic model for text-to-speech,” in ICML , 2021
2021
Cited alongside, same era.
“Google cloud speech-to-text.” [Online]. Available: https://cloud.google.com/speech-to-text/
Cited in the paper.
2021
Closest in time.
A. Lańcucki, “FastPitch: Parallel text-to-speech with pitch prediction,” in ICASSP , 2021
2021
Closest in time.
A. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in ICML , 2021
2021
Closest in time.
R. Badlani, A. Łancucki, K. J. Shih, R. Valle, W. Ping, and B. Catanzaro, “One tts alignment to rule them all,” 2021
2021
Closest in time.