Fetching the paper…
Reading the bibliography…
We introduce VoiceCraft, a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on audiobooks, internet videos, and podcasts.
Mel-cepstral distance measure for objective speech quality assessment
Robert F. Kubichek. 1993 · 1993
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Pyin: A fundamental frequency estimator using probabilistic threshold distributions
Matthias Mauch and Simon Dixon. 2014 · 2014
Earlier work this paper cites.
librosa: Audio and music signal analysis in python
Brian McFee, Colin Raffel, Dawen Liang, Daniel P. W. Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. 2015 · 2015
Earlier work this paper cites.
The lj speech dataset
Keith Ito and Linda Johnson. 2017 · 2017
Earlier work this paper cites.
Voco: text-based insertion and replacement in audio narration
Zeyu Jin, Gautham J. Mysore, Stephen DiVerdi, Jingwan Lu, and Adam Finkelstein. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92)
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald. 2019 · 2019
Earlier work this paper cites.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet-Trung Dang, Robert A. J. Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Z. Chen, and Yonghui Wu. 2019 · 2019
Earlier work this paper cites.
100,000 podcasts: A spoken English document corpus
Ann Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu, Rezvaneh Rezapour, Hamed Bonab, Maria Eskevich, Gareth Jones, Jussi Karlgren, Ben Carterette, and Rosie Jones. 2020 · 2020
Earlier work this paper cites.
Enabling language models to fill in the blanks
Chris Donahue, Mina Lee, and Percy Liang. 2020 · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
One-class learning towards synthetic voice spoofing detection
You Zhang, Fei Jiang, and Zhiyao Duan. 2020 · 2020
Earlier work this paper cites.
Phonemizer: Text to phones transcription for multiple languages in python
Mathieu Bernard and Hadrien Titeux. 2021 · 2021
Earlier work this paper cites.
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone
Edresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior, Eren Gölge, and Moacir Antonelli Ponti. 2021 · 2021
Earlier work this paper cites.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio
Guoguo Chen, Shuzhou Chai, Guan-Bo Wang, Jiayu Du, Weiqiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, Mingjie Jin, Sanjeev Khudanpur, Shinji Watanabe, Shuaijiang Zhao, Wei Zou, Xiangang Li, Xuchen Yao, Yongqing Wang, Yujun Wang, Zhao You, and Zhiyong Yan. 2021a · 2021
Cited alongside, same era.
Text-free image-to-speech synthesis using learned segmental units
Wei-Ning Hsu, David Harwath, Tyler Miller, Christopher Song, and James R. Glass. 2021 · 2021
Cited alongside, same era.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son. 2021 · 2021
Cited alongside, same era.
On generative spoken language modeling from raw audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu Nguyen, Jade Copet, Alexei Baevski, Adel Ben Mohamed, and Emmanuel Dupoux. 2021 · 2021
Cited alongside, same era.
Context-aware prosody correction for text-based speech editing
Max Morrison, Lucas Rencker, Zeyu Jin, Nicholas J. Bryan, Juan Pablo Cáceres, and Bryan Pardo. 2021 · 2021
Soundstorm: Efficient parallel audio generation
Zalán Borsos, Matthew Sharifi, Damien Vincent, Eugene Kharitonov, Neil Zeghidour, and Marco Tagliasacchi. 2023 · 2023
Later among the works it cites.
Wavmark: Watermarking for audio generation
Guang Chen, Yu Wu, Shujie Liu, Tao Liu, Xiaoyong Du, and Furu Wei. 2023 · 2023
Later among the works it cites.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Defossez. 2023 · 2023
Later among the works it cites.
Singsong: Generating musical accompaniments from singing
Chris Donahue, Antoine Caillon, Adam Roberts, Ethan Manilow, Philippe Esling, Andrea Agostinelli, Mauro Verzetti, Ian Simon, Olivier Pietquin, Neil Zeghidour, and Jesse Engel. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Editspeech: A text based speech editing system using partial inference and bidirectional fusion
Daxin Tan, Liqun Deng, Yu Ting Yeung, Xin Jiang, Xiao Chen, and Tan Lee. 2021 · 2021
Cited alongside, same era.
Asvspoof 2021: accelerating progress in spoofed and deepfake speech detection
Junichi Yamagishi, Xin Wang, Massimiliano Todisco, Md. Sahidullah, Jose Patino, Andreas Nautsch, Xuechen Liu, Kong-Aik Lee, Tomi H. Kinnunen, Nicholas W. D. Evans, and Héctor Delgado. 2021 · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2021 · 2021
Cited alongside, same era.
Cm3: A causal masked multimodal model of the internet
Armen Aghajanyan, Po-Yao (Bernie) Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, and Luke Zettlemoyer. 2022 · 2022
Cited alongside, same era.
A3t: Alignment-aware acoustic and text pretraining for speech synthesis and editing
He Bai, Renjie Zheng, Junkun Chen, Xintong Li, Mingbo Ma, and Liang Huang. 2022 · 2022
Cited alongside, same era.
Efficient training of language models to fill in the middle
Mohammad Bavarian, Heewoo Jun, Nikolas A. Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. 2022 · 2022
Cited alongside, same era.
High fidelity neural audio compression
Alexandre Defossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2022 · 2022
Cited alongside, same era.
Hugo Flores Garcia, Prem Seetharaman, Rithesh Kumar, and Bryan Pardo. 2023 · 2023
Later among the works it cites.
Prompttts: Controllable text-to-speech with text descriptions
Zhifang Guo, Yichong Leng, Yihan Wu, Sheng Zhao, and Xuejiao Tan. 2022 · 2023
Later among the works it cites.
Textrolspeech: A text style control speech corpus with codec language text-to-speech models
Shengpeng Ji, Jia li Zuo, Minghui Fang, Ziyue Jiang, Feiyang Chen, Xinyu Duan, Baoxing Huai, and Zhou Zhao. 2023 · 2023
Later among the works it cites.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Eugene Kharitonov, Damien Vincent, Zalán Borsos, Raphaël Marinier, Sertan Girgin, Olivier Pietquin, Matthew Sharifi, Marco Tagliasacchi, and Neil Zeghidour. 2023 · 2023
Later among the works it cites.
Voicebox: Text-guided multilingual universal speech generation at scale
Matt Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, and Wei-Ning Hsu. 2023 · 2023
Later among the works it cites.
An unofficial pytorch implementation of vall-e
Feiteng Li. 2023 · 2023
Later among the works it cites.
Promptstyle: Controllable style transfer for text-to-speech with natural language descriptions
Guanghou Liu, Yongmao Zhang, Yinjiao Lei, Yunlin Chen, Rui Wang, Zhifei Li, and Linfu Xie. 2023 · 2023
Later among the works it cites.
Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt
Dongchao Yang, Songxiang Liu, Rongjie Huang, Guangzhi Lei, Chao Weng, Helen M. Meng, and Dong Yu. 2023 · 2023
Later among the works it cites.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Zi-Hua Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei. 2023 · 2023
Later among the works it cites.
Mega-TTS 2: Boosting prompting mechanisms for zero-shot speech synthesis
Ziyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He, Zhenhui Ye, Shengpeng Ji, Qian Yang, Chen Zhang, Pengfei Wei, Chunfeng Wang, Xiang Yin, Zejun MA, and Zhou Zhao. 2024 · 2024
Closest in time.
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Daniel Lyth and Simon King. 2024 · 2024
Closest in time.
Proactive detection of voice cloning with localized watermarking
Robin San Roman, Pierre Fernandez, Alexandre Defossez, Teddy Furon, Tuan Tran, and Hady ElSahar. 2024 · 2024
Closest in time.
Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering
Yakun Song, Zhuo Chen, Xiaofei Wang, Ziyang Ma, and Xie Chen. 2024 · 2024
Closest in time.
Zipformer: A faster and better encoder for automatic speech recognition
Zengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang, Fangjun Kuang, Yifan Yang, Zengrui Jin, Long Lin, and Daniel Povey. 2024 · 2024
Closest in time.