Fetching the paper…
Reading the bibliography…
Recent advancements in large language models (LLMs) have driven significant progress in zero-shot text-to-speech (TTS) synthesis.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019 · 1904
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. 2019 · 1912
Earlier work this paper cites.
Ddsp: Differentiable digital signal processing
Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts. 2020 · 2001
Earlier work this paper cites.
Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra. 2001 · 2001
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008 · 2008
Earlier work this paper cites.
Design of speech corpus for mandarin text to speech
Jianhua Tao, Fangzhou Liu, Meng Zhang, and Huibin Jia. 2008 · 2008
Earlier work this paper cites.
Aishell-3: A multi-speaker mandarin tts corpus and the baselines
Yao Shi, Hui Bu, Xin Xu, Shaoji Zhang, and Ming Li. 2020 · 2010
Earlier work this paper cites.
Mls: A large-scale multilingual dataset for speech research
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert. 2020 · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013 · 2013
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma. 2014 · 2014
Earlier work this paper cites.
Surrey audio-visual expressed emotion (savee) database
Philip Jackson and SJUoSG Haq. 2014 · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
A non-intrusive short-time objective intelligibility measure
Asger Heidemann Andersen, Jan Mark de Haan, Zheng-Hua Tan, and Jesper Jensen. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al. 2017 · 2017
Earlier work this paper cites.
The emotional voices database: Towards controlling the emotion dimension in voice generation systems
Adaeze Adigwe, Noé Tits, Kevin El Haddad, Sarah Ostadabbas, and Thierry Dutoit. 2018 · 2018
Earlier work this paper cites.
An open source emotional speech corpus for human robot interaction applications
Jesin James, Li Tian, and Catherine Watson. 2018 · 2018
Earlier work this paper cites.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Steven R Livingstone and Frank A Russo. 2018 · 2018
Earlier work this paper cites.
Meld: A multimodal multi-party dataset for emotion recognition in conversations
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2018 · 2018
Earlier work this paper cites.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre De Brebisson, Yoshua Bengio, and Aaron C Courville. 2019 · 2019
Earlier work this paper cites.
Magicdata mandarin chinese read speech corpus
MagicData. 2019 · 2019
Earlier work this paper cites.
CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92)
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald. 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Earlier work this paper cites.
Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020 · 2020
Earlier work this paper cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020 · 2020
Earlier work this paper cites.
A spectral energy distance for parallel speech synthesis
Alexey Gritsenko, Tim Salimans, Rianne van den Berg, Jasper Snoek, and Nal Kalchbrenner. 2020 · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Earlier work this paper cites.
The msp-conversation corpus
Luz Martinez, Mohammed Abdelwahab, and Carlos Busso. 2020 · 2020
Earlier work this paper cites.
Mead: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. 2020 · 2020
Cited alongside, same era.
Hi-Fi Multi-Speaker English TTS Dataset
Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg, and Yang Zhang. 2021 · 2021
Cited alongside, same era.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, et al. 2021 · 2021
Cited alongside, same era.
Vector-quantized image modeling with improved vqgan
Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. 2021 · 2021
Cited alongside, same era.
Audiobox: Unified audio generation with natural language prompts
Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu, Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins, William Ngan, et al. 2023 · 2023
Later among the works it cites.
Wespeaker: A research and production oriented speaker embedding learning toolkit
Hongji Wang, Chengdong Liang, Shuai Wang, Zhengyang Chen, Binbin Zhang, Xu Xiang, Yanlei Deng, and Yanmin Qian. 2023b · 2023
Later among the works it cites.
Seed-tts: A family of high-quality versatile speech generation models
Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, et al. 2024 · 2024
Later among the works it cites.
Xtts: a massively multilingual zero-shot text-to-speech model
Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar, Iulian Gulea, Logan Hart, Aya Aljafari, Joshua Meyer, Reuben Morais, Samuel Olayemi, et al. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li. 2021 · 2021
Cited alongside, same era.
Large-scale self-supervised speech representation learning for automatic speaker verification
Zhengyang Chen, Sanyuan Chen, Yu Wu, Yao Qian, Chengyi Wang, Shujie Liu, Yanmin Qian, and Michael Zeng. 2022 · 2022
Cited alongside, same era.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2022 · 2022
Cited alongside, same era.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022 · 2022
Cited alongside, same era.
Utmos: Utokyo-sarulab system for voicemos challenge 2022
Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2022 · 2022
Cited alongside, same era.
M3ed: Multi-modal multi-scene multi-label emotional dialogue database
Jinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu, Qin Jin, Xinchao Wang, and Haizhou Li. 2022 · 2022
Cited alongside, same era.
Speech-based age and gender prediction with transformers
Felix Burkhardt, Johannes Wagner, Hagen Wierstorf, Florian Eyben, and Björn Schuller. 2023 · 2023
Cited alongside, same era.
Funasr: A fundamental end-to-end speech recognition toolkit
Zhifu Gao, Zerui Li, Jiaming Wang, Haoneng Luo, Xian Shi, Mengzhe Chen, Yabin Li, Lingyun Zuo, Zhihao Du, Zhangyu Xiao, et al. 2023 · 2023
Cited alongside, same era.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2024 · 2024
Later among the works it cites.
Moshi: a speech-text foundation model for real-time dialogue
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. 2024 · 2024
Later among the works it cites.
Fireredtts: A foundation text-to-speech framework for industry-level generative speech applications
Hao-Han Guo, Kun Liu, Fei-Yu Shen, Yi-Chen Wu, Feng-Long Xie, Kun Xie, and Kai-Tuo Xu. 2024 · 2024
Later among the works it cites.
Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, et al. 2024 · 2024
Later among the works it cites.
Textrolspeech: A text style control speech corpus with codec language text-to-speech models
Shengpeng Ji, Jialong Zuo, Minghui Fang, Ziyue Jiang, Feiyang Chen, Xinyu Duan, Baoxing Huai, and Zhou Zhao. 2024b · 2024
Later among the works it cites.
Speechcraft: A fine-grained expressive speech dataset with natural language description
Zeyu Jin, Jia Jia, Qixin Wang, Kehan Li, Shuoyi Zhou, Songtao Zhou, Xiaoyu Qin, and Zhiyong Wu. 2024 · 2024
Later among the works it cites.
High-fidelity audio compression with improved rvqgan
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. 2024 · 2024
Later among the works it cites.
Stack-and-delay: a new codebook pattern for music generation
Gael Le Lan, Varun Nagaraja, Ernie Chang, David Kant, Zhaoheng Ni, Yangyang Shi, Forrest Iandola, and Vikas Chandra. 2024 · 2024
Later among the works it cites.
Generative expressive conversational speech synthesis
Rui Liu, Yifan Hu, Yi Ren, Xiang Yin, and Haizhou Li. 2024 · 2024
Later among the works it cites.
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Dan Lyth and Simon King. 2024 · 2024
Later among the works it cites.
Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model benchmark
Linhan Ma, Dake Guo, Kun Song, Yuepeng Jiang, Shuai Wang, Liumeng Xue, Weiming Xu, Huan Zhao, Binbin Zhang, and Lei Xie. 2024 · 2024
Later among the works it cites.
Scaling transformers for low-bitrate high-quality speech coding
Julian D Parker, Anton Smirnov, Jordi Pons, CJ Carr, Zack Zukowski, Zach Evans, and Xubo Liu. 2024 · 2024
Later among the works it cites.
Fewer-token neural speech codec with time-invariant codes
Yong Ren, Tao Wang, Jiangyan Yi, Le Xu, Jianhua Tao, Chu Yuan Zhang, and Junzuo Zhou. 2024 · 2024
Later among the works it cites.
Toneunit: A speech discretization approach for tonal language speech synthesis
Dehua Tao, Daxin Tan, Yu Ting Yeung, Xiao Chen, and Tan Lee. 2024 · 2024
Later among the works it cites.
Ts3-codec: Transformer-based simple streaming single codec
Haibin Wu, Naoyuki Kanda, Sefik Emre Eskimez, and Jinyu Li. 2024 · 2024
Later among the works it cites.
Bigcodec: Pushing the limits of low-bitrate neural speech codec
Detai Xin, Xu Tan, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2024 · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. 2024 · 2024
Later among the works it cites.
Freecodec: A disentangled neural speech codec with fewer tokens
Youqiang Zheng, Weiping Tu, Yueteng Kang, Jie Chen, Yike Zhang, Li Xiao, Yuhong Yang, and Long Ma. 2024 · 2024
Later among the works it cites.
The codec language model-based zero-shot spontaneous style tts system for covoc challenge 2024
Shuoyi Zhou, Yixuan Zhou, Weiqing Li, Jun Chen, Runchuan Ye, Weihao Wu, Zijian Lin, Shun Lei, and Zhiyong Wu. 2024a · 2024
Later among the works it cites.
Unistyle: Unified style modeling for speaking style captioning and stylistic speech synthesis
Xinfa Zhu, Wenjie Tian, Xinsheng Wang, Lei He, Yujia Xiao, Xi Wang, Xu Tan, Sheng Zhao, and Lei Xie. 2024 · 2024
Later among the works it cites.
Masked audio generation using a single non-autoregressive transformer
Alon Ziv, Itai Gat, Gael Le Lan, Tal Remez, Felix Kreuk, Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2024 · 2024
Later among the works it cites.
Llasa: Scaling train-time and inference-time compute for llama-based speech synthesis
Zhen Ye, Xinfa Zhu, Chi-Min Chan, Xinsheng Wang, Xu Tan, Jiahe Lei, Yi Peng, Haohe Liu, Yizhu Jin, Zheqi DAI, et al. 2025 · 2025
Closest in time.