Fetching the paper…
Reading the bibliography…
This paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT).
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019 · 1904
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. 2019 · 1912
Earlier work this paper cites.
Logistic-normal distributions: Some properties and uses
Jhon Atchison and Sheng M Shen. 1980 · 1980
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2020 · 2011
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
The LJ speech dataset
Keith Ito and Linda Johnson. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I Loshchilov. 2017 · 2017
Earlier work this paper cites.
torchdiffeq
Ricky T. Q. Chen. 2018 · 2018
Earlier work this paper cites.
Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, et al. 2018 · 2018
Earlier work this paper cites.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Libri-light: A benchmark for ASR with limited or no supervision
Jacob Kahn, Morgane Riviere, Weiyi Zheng, Evgeny Kharitonov, et al. 2020 · 2020
Earlier work this paper cites.
Glow-TTS: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon. 2020 · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. 2021 · 2021
Earlier work this paper cites.
Didispeech: A large scale Mandarin speech corpus
Tingwei Guo, Cheng Wen, Dongwei Jiang, Ne Luo, et al. 2021 · 2021
Earlier work this paper cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Earlier work this paper cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son. 2021 · 2021
Earlier work this paper cites.
Grad-tts: A diffusion probabilistic model for text-to-speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov. 2021 · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis. 2021 · 2021
Earlier work this paper cites.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2021 · 2021
Earlier work this paper cites.
A3t: Alignment-aware acoustic and text pretraining for speech synthesis and editing
He Bai, Renjie Zheng, Junkun Chen, Mingbo Ma, Xintong Li, and Liang Huang. 2022 · 2022
Earlier work this paper cites.
Large-scale self-supervised speech representation learning for automatic speaker verification
Zhengyang Chen, Sanyuan Chen, Yu Wu, Yao Qian, Chengyi Wang, Shujie Liu, Yanmin Qian, and Michael Zeng. 2022 · 2022
Cited alongside, same era.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2022 · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Cited alongside, same era.
Bigvgan: A universal neural vocoder with large-scale training
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon. 2022 · 2022
Cited alongside, same era.
VALL-E R: Robust and efficient zero-shot text-to-speech synthesis via monotonic alignment
Bing Han, Long Zhou, Shujie Liu, Sanyuan Chen, Lingwei Meng, Yanming Qian, Yanqing Liu, Sheng Zhao, Jinyu Li, and Furu Wei. 2024 · 2024
Closest in time.
Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, et al. 2024 · 2024
Closest in time.
Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, et al. 2024 · 2024
Closest in time.
Libriheavy: a 50,000 hours asr corpus with punctuation casing and context
Wei Kang, Xiaoyu Yang, Zengwei Yao, Fangjun Kuang, Yifan Yang, Liyong Guo, Long Lin, and Daniel Povey. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022 · 2022
Cited alongside, same era.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022 · 2022
Cited alongside, same era.
Utmos: Utokyo-sarulab system for voicemos challenge 2022
Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2022 · 2022
Cited alongside, same era.
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho. 2022 · 2022
Cited alongside, same era.
LibriSpeech-PC: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end ASR models
Aleksandr Meister, Matvei Novikov, Nikolay Karpov, Evelina Bakhturina, Vitaly Lavrukhin, and Boris Ginsburg. 2023 · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
William Peebles and Saining Xie. 2023 · 2023
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Cited alongside, same era.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, et al. 2023 · 2023
Cited alongside, same era.
Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. 2024 · 2024
Closest in time.
Understanding diffusion objectives as the ELBO with simple data augmentation
Diederik Kingma and Ruiqi Gao. 2024 · 2024
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, et al. 2024 · 2024
Closest in time.
DiTTo-TTS: Efficient and scalable zero-shot text-to-speech with diffusion transformer
Keon Lee, Dong Won Kim, Jaehyeon Kim, and Jaewoong Cho. 2024 · 2024
Closest in time.
Autoregressive image generation without vector quantization
Tianhong Li, Yonglong Tian, He Li, Mingyang Deng, and Kaiming He. 2024 · 2024
Closest in time.
WenetSpeech4TTS: A 12,800-hour Mandarin TTS corpus for large speech generation model benchmark
Linhan Ma, Dake Guo, Kun Song, Yuepeng Jiang, Shuai Wang, Liumeng Xue, Weiming Xu, Huan Zhao, Binbin Zhang, and Lei Xie. 2024 · 2024
Closest in time.
Matcha-TTS: A fast TTS architecture with conditional flow matching
Shivam Mehta, Ruibo Tu, Jonas Beskow, Éva Székely, and Gustav Eje Henter. 2024 · 2024
Closest in time.
Autoregressive speech synthesis without vector quantization
Lingwei Meng, Long Zhou, Shujie Liu, Sanyuan Chen, et al. 2024 · 2024
Closest in time.
NDVQ: Robust neural audio codec with normal distribution-based vector quantization
Zhikang Niu, Sanyuan Chen, Long Zhou, Ziyang Ma, Xie Chen, and Shujie Liu. 2024 · 2024
Closest in time.
Convnext-TTS and Convnext-VC: Convnext-based fast end-to-end sequence-to-sequence text-to-speech and voice conversion
Takuma Okamoto, Yamato Ohtani, Tomoki Toda, and Hisashi Kawai. 2024 · 2024
Closest in time.
Voicecraft: Zero-shot speech editing and text-to-speech in the wild
Puyuan Peng, Po-Yao Huang, Daniel Li, Abdelrahman Mohamed, and David Harwath. 2024 · 2024
Closest in time.
ELLA-V: Stable neural codec language modeling with alignment-guided sequence reordering
Yakun Song, Zhuo Chen, Xiaofei Wang, Ziyang Ma, and Xie Chen. 2024 · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2024 · 2024
Closest in time.
Naturalspeech: End-to-end text-to-speech synthesis with human-level quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, et al. 2024 · 2024
Closest in time.
MaskGCT: Zero-shot text-to-speech with masked generative codec transformer
Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng, Haotian Guo, Jiachen Zheng, Qiang Zhang, Shunsi Zhang, and Zhizheng Wu. 2024 · 2024
Closest in time.
Rall-e: Robust codec language modeling with chain-of-thought prompting for text-to-speech synthesis
Detai Xin, Xu Tan, Kai Shen, Zeqian Ju, et al. 2024 · 2024
Closest in time.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, et al. 2023b · 2024
Closest in time.