Fetching the paper…
Reading the bibliography…
The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation.
Simultaneous modeling of spectrum, pitch and duration in HMM-based speech synthesis
Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, et al. 1999 · 1999
Earlier work this paper cites.
Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, and Alan W Black. 2009 · 2009
Earlier work this paper cites.
TTS synthesis with bidirectional LSTM based recurrent neural networks
Yuchen Fan, Yao Qian, Feng-Long Xie, and Frank K Soong. 2014 · 2014
Earlier work this paper cites.
LibriSpeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis
Heiga Zen and Haşim Sak. 2015 · 2015
Earlier work this paper cites.
Investigating RNN-based speech enhancement methods for noise-robust text-to-speech
Cassia Valentini-Botinhao, Xin Wang, Shinji Takaki, and Junichi Yamagishi. 2016 · 2016
Earlier work this paper cites.
Deep Voice: Real-time neural text-to-speech
Sercan Ö Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, et al. 2017 · 2017
Earlier work this paper cites.
Montreal Forced Aligner: Trainable text-speech alignment using Kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, et al. 2017 · 2017
Earlier work this paper cites.
WaveNet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, et al. 2017 · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, et al. 2017 · 2017
Earlier work this paper cites.
Cluster-GCN: An efficient algorithm for training deep and large graph convolutional networks
Wei-Lin Chiang, Xuanqing Liu, Si Si, Yang Li, et al. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, et al. 2019 · 2019
Earlier work this paper cites.
Hierarchical generative modeling for controllable speech synthesis
Wei-Ning Hsu, Yu Zhang, Ron J Weiss, Heiga Zen, et al. 2019 · 2019
Earlier work this paper cites.
Comparison of diverse decoding methods from conditional language models
Daphne Ippolito, Reno Kriz, João Sedoc, Maria Kustikova, et al. 2019 · 2019
Earlier work this paper cites.
WaveGlow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro. 2019 · 2019
Earlier work this paper cites.
FastSpeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, et al. 2019 · 2019
Earlier work this paper cites.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, et al. 2019 · 2019
Earlier work this paper cites.
UniLMv2: Pseudo-masked language models for unified language model pre-training
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, et al. 2020 · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, et al. 2020 · 2020
Cited alongside, same era.
Conformer: Convolution-augmented Transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, et al. 2020 · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, et al. 2020 · 2020
Cited alongside, same era.
Glow-TTS: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon. 2020 · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, et al. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, et al. 2022 · 2022
Later among the works it cites.
LaMDA: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, et al. 2022 · 2022
Later among the works it cites.
Language models with image descriptors are strong few-shot video-language learners
Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, et al. 2022 · 2022
Later among the works it cites.
ItôWave: Itô stochastic differential equation is all you need for wave generation
Shoule Wu and Ziqiang Shi. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, et al. 2020 · 2020
Cited alongside, same era.
Flow-TTS: A non-autoregressive network for text to speech based on flow
Chenfeng Miao, Shuang Liang, Minchuan Chen, Jun Ma, et al. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, et al. 2020 · 2020
Cited alongside, same era.
Improved techniques for training score-based generative models
Yang Song and Stefano Ermon. 2020 · 2020
Cited alongside, same era.
Diff-TTS: A denoising diffusion model for text-to-speech
Myeonghun Jeong, Hyeongju Kim, Sung Jun Cheon, Byoung Jin Choi, et al. 2021 · 2021
Cited alongside, same era.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son. 2021 · 2021
Cited alongside, same era.
Grad-TTS: A diffusion probabilistic model for text-to-speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, et al. 2021 · 2021
Cited alongside, same era.
Zero-shot video question answering via frozen bidirectional language models
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, et al. 2022 · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich Text-to-Image generation
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, et al. 2022 · 2022
Later among the works it cites.
AudioLM: A language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, et al. 2023 · 2023
Later among the works it cites.
PaLM: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, et al. 2023 · 2023
Later among the works it cites.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2023 · 2023
Later among the works it cites.
Decoder-only or encoder-decoder? Interpreting language model as a regularized encoder-decoder
Zihao Fu, Wai Lam, Qian Yu, Anthony Man-Cho So, et al. 2023 · 2023
Later among the works it cites.
Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, et al. 2023 · 2023
Later among the works it cites.
Investigating the utility of surprisal from large language models for speech synthesis prosody
Sofoklis Kakouros, Juraj Šimko, Martti Vainio, and Antti Suni. 2023 · 2023
Later among the works it cites.
Speak, Read and Prompt: High-fidelity text-to-speech with minimal supervision
Eugene Kharitonov, Damien Vincent, Zalán Borsos, Raphaël Marinier, et al. 2023 · 2023
Later among the works it cites.
AudioGen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, et al. 2023 · 2023
Later among the works it cites.
Voicebox: Text-guided multilingual universal speech generation at scale
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, et al. 2023 · 2023
Later among the works it cites.
Flow matching for generative modeling
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, et al. 2023 · 2023
Later among the works it cites.
AudioPaLM: A large language model that can speak and listen
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, et al. 2023 · 2023
Later among the works it cites.
NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, et al. 2023 · 2023
Later among the works it cites.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, et al. 2023 · 2023
Later among the works it cites.