Fetching the paper…
Reading the bibliography…
In text-to-speech (TTS) synthesis, diffusion models have achieved promising generation quality.
Sur la théorie relativiste de l’électron et l’interprétation de la mécanique quantique
Erwin Schrödinger · 1932
Earlier work this paper cites.
On a formula concerning stochastic differentials
Kiyosi Itô · 1951
Earlier work this paper cites.
A class of explicit multistep exponential integrators for semilinear problems
Mari Paz Calvo and César Palencia · 2006
Earlier work this paper cites.
Exponential rosenbrock-type methods
Marlis Hochbruck, Alexander Ostermann, and Julia Schweitzer · 2009
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Pascal Vincent · 2011
Earlier work this paper cites.
Some properties of path measures
Christian Léonard · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D.P. Kingma and J. Ba · 2015
Earlier work this paper cites.
The LJ speech dataset
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Earlier work this paper cites.
Wavegrad: Estimating gradients for waveform generation
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, and William Chan · 2021
Earlier work this paper cites.
Diffusion schrödinger bridge with applications to score-based generative modeling
Valentin De Bortoli, James Thornton, Jeremy Heng, and Arnaud Doucet · 2021
Earlier work this paper cites.
Diff-tts: A denoising diffusion model for text-to-speech
Myeonghun Jeong, Hyeongju Kim, Sung Jun Cheon, Byoung Jin Choi, and Nam Soo Kim · 2021
Earlier work this paper cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Earlier work this paper cites.
Variational diffusion models
Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho · 2021
Cited alongside, same era.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Cited alongside, same era.
Grad-tts: A diffusion probabilistic model for text-to-speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail A. Kudinov · 2021
Cited alongside, same era.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2021
Cited alongside, same era.
A survey on neural speech synthesis
Xu Tan, Tao Qin, Frank K. Soong, and Tie-Yan Liu · 2021
Cited alongside, same era.
Solving schrödinger bridges via maximum likelihood
Francisco Vargas, Pierre Thodoroff, Austen Lamacraft, and Neil Lawrence · 2021
Lightgrad: Lightweight diffusion probabilistic model for text-to-speech
Jie Chen, Xingchen Song, Zhendong Peng, Binbin Zhang, Fuping Pan, and Zhiyong Wu · 2023
Closest in time.
Seeds: Exponential sde solvers for fast high-quality sampling from diffusion models
Martin Gonzalez, Nelson Fernandez, Thuy Tran, Elies Gherbi, Hatem Hajri, and Nader Masmoudif · 2023
Closest in time.
Reflow-tts: A rectified flow model for high-fidelity text-to-speech
Wenhao Guan, Qi Su, Haodong Zhou, Shiyu Miao, Xingjia Xie, Lin Li, and Qingyang Hong · 2023
Closest in time.
Voiceflow: Efficient text-to-speech with rectified flow matching
Yiwei Guo, Chenpeng Du, Ziyang Ma, Xie Chen, and Kai Yu · 2023
Closest in time.
Simple diffusion: end-to-end diffusion for high resolution images
Emiel Hoogeboom, Jonathan Heek, and Tim Salimans · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep generative learning via schrödinger bridge
Gefei Wang, Yuling Jiao, Qian Xu, Yang Wang, and Can Yang · 2021
Cited alongside, same era.
Prodiff: Progressive fast diffusion model for high-quality text-to-speech
Rongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu, Chenye Cui, and Yi Ren · 2022
Cited alongside, same era.
Priorgrad: Improving conditional denoising diffusion models with data-driven adaptive prior
Sang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan, Chang Liu, Qi Meng, Tao Qin, Wei Chen, Sungroh Yoon, and Tie-Yan Liu · 2022
Cited alongside, same era.
Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis
Yichong Leng, Zehua Chen, Junliang Guo, Haohe Liu, Jiawei Chen, Xu Tan, Danilo P. Mandic, Lei He, Xiang-Yang Li, Tao Qin, Sheng Zhao, and Tie-Yan Liu · 2022
Cited alongside, same era.
Diffusion-based voice conversion with fast maximum likelihood sampling scheme
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, Mikhail Kudinov, and Jiansheng Wei · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Cited alongside, same era.
Closest in time.
Myeongjin Ko and Yong-Hoon Choi · 2023
Closest in time.
Flow matching for generative modeling
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le · 2023
Closest in time.
Matcha-tts: A fast tts architecture with conditional flow matching
Shivam Mehta, Ruibo Tu, Jonas Beskow, Éva Székely, and Gustav Eje Henter · 2023
Closest in time.
Diffusion bridge mixture transports, schrödinger bridge problems and generative modeling
Stefano Peluchetti · 2023
Closest in time.
Se-bridge: Speech enhancement with consistent brownian bridge
Zhibin Qiu, Mengfan Fu, Fuchun Sun, Gulila Altenbek, and Hao Huang · 2023
Closest in time.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian · 2023
Closest in time.
Diffusion schrödinger bridge matching
Yuyang Shi, Valentin De Bortoli, Andrew Campbell, and Arnaud Doucet · 2023
Closest in time.
Aligned diffusion schrödinger bridges
Vignesh Ram Somnath, Matteo Pariset, Ya-Ping Hsieh, Maria Rodriguez Martinez, Andreas Krause, and Charlotte Bunne · 2023
Closest in time.
Consistency models
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever · 2023
Closest in time.
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu · 2023
Closest in time.
Comospeech: One-step speech and singing voice synthesis via consistency model
Zhen Ye, Wei Xue, Xu Tan, Jie Chen, Qifeng Liu, and Yike Guo · 2023
Closest in time.