Fetching the paper…
Reading the bibliography…
This paper introduces Embarrassingly Easy Text-to-Speech (E2 TTS), a fully non-autoregressive zero-shot text-to-speech system that offers human-level naturalness and state-of-the-art speaker similarity and intelligibility.
“U-net: Convolutional networks for biomedical image segmentation,”
Olaf Ronneberger, Philipp Fischer, and Thomas Brox, · 2015
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using kaldi.,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Earlier work this paper cites.
“Neural voice cloning with a few samples,”
Sercan Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou, · 2018
Earlier work this paper cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, Yonghui Wu, et al., · 2018
Earlier work this paper cites.
“Neural ordinary differential equations,”
Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud, · 2018
Earlier work this paper cites.
“FastSpeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Earlier work this paper cites.
“Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,”
Bo Li, Yu Zhang, Tara Sainath, Yonghui Wu, and William Chan, · 2019
Earlier work this paper cites.
“Glow-TTS: A generative flow for text-to-speech via monotonic alignment search,”
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon, · 2020
Earlier work this paper cites.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
Jonathan Ho, Ajay Jain, and Pieter Abbeel, · 2020
Earlier work this paper cites.
“Score-based generative modeling through stochastic differential equations,”
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole, · 2020
Earlier work this paper cites.
“Libri-light: A benchmark for ASR with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al., · 2020
Cited alongside, same era.
“FastSpeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2021
Cited alongside, same era.
“Grad-TTS: A diffusion probabilistic model for text-to-speech,”
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov, · 2021
Cited alongside, same era.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Flow matching for generative modeling,”
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le, · 2022
Cited alongside, same era.
“E3 TTS: Easy end-to-end diffusion-based text to speech,”
Yuan Gao, Nobuyuki Morioka, Yu Zhang, and Nanxin Chen, · 2023
Later among the works it cites.
“LibriSpeech-PC: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models,”
Aleksandr Meister, Matvei Novikov, Nikolay Karpov, Evelina Bakhturina, Vitaly Lavrukhin, and Boris Ginsburg, · 2023
Later among the works it cites.
“Naturalspeech: End-to-end text-to-speech synthesis with human-level quality,”
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, et al., · 2024
Closest in time.
“Vall-t: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,”
Chenpeng Du, Yiwei Guo, Hankun Wang, Yifan Yang, Zhikang Niu, Shuai Wang, Hui Zhang, Xie Chen, and Kai Yu, · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“BigVGAN: A universal neural vocoder with large-scale training,”
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon, · 2022
Cited alongside, same era.
“Classifier-free diffusion guidance,”
Jonathan Ho and Tim Salimans, · 2022
Cited alongside, same era.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Cited alongside, same era.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Cited alongside, same era.
“Speechx: Neural codec language model as a versatile speech transformer,”
Xiaofei Wang, Manthan Thakker, Zhuo Chen, Naoyuki Kanda, Sefik Emre Eskimez, Sanyuan Chen, Min Tang, Shujie Liu, Jinyu Li, and Takuya Yoshioka, · 2023
Cited alongside, same era.
“Uniaudio: An audio foundation model toward universal audio generation,”
Dongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang, Songxiang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Xixin Wu, et al., · 2023
Cited alongside, same era.
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian, · 2023
Cited alongside, same era.
Detai Xin, Xu Tan, Kai Shen, Zeqian Ju, Dongchao Yang, Yuancheng Wang, Shinnosuke Takamichi, Hiroshi Saruwatari, Shujie Liu, Jinyu Li, et al., · 2024
Closest in time.
“VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers,”
Sanyuan Chen, Shujie Liu, Long Zhou, Yanqing Liu, Xu Tan, Jinyu Li, Sheng Zhao, Yao Qian, and Furu Wei, · 2024
Closest in time.
“Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,”
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, et al., · 2024
Closest in time.
“Voicebox: Text-guided multilingual universal speech generation at scale,”
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al., · 2024
Closest in time.
“Matcha-TTS: A fast TTS architecture with conditional flow matching,”
Shivam Mehta, Ruibo Tu, Jonas Beskow, Éva Székely, and Gustav Eje Henter, · 2024
Closest in time.
“Seed-TTS: A Family of High-Quality Versatile Speech Generation Models,”
Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, et al., · 2024
Closest in time.
“LibriHeavy: a 50,000 hours ASR corpus with punctuation casing and context,”
Wei Kang, Xiaoyu Yang, Zengwei Yao, Fangjun Kuang, Yifan Yang, Liyong Guo, Long Lin, and Daniel Povey, · 2024
Closest in time.
“An investigation of noise robustness for flow-matching-based zero-shot TTS,”
Xiaofei Wang, Sefik Emre Eskimez, Manthan Thakker, Hemin Yang, Zirun Zhu, Min Tang, Yufei Xia, Jinzhu Li, Sheng Zhao, Jinyu Li, and Naoyuki Kanda, · 2024
Closest in time.