Fetching the paper…
Reading the bibliography…
Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts.
“Rank analysis of incomplete block designs: I. the method of paired comparisons,”
Ralph Allan Bradley and Milton E. Terry, · 1952
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov et al., · 2015
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang et al., · 2017
Earlier work this paper cites.
“Proximal policy optimization algorithms,”
John Schulman et al., · 2017
Earlier work this paper cites.
“Fastspeech: Fast, robust and controllable text to speech,”
Yi Ren et al., · 2019
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2019
Earlier work this paper cites.
“Rawnet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification,”
Jee weon Jung et al., · 2019
Earlier work this paper cites.
“CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit,” 2019
Junichi Yamagishi et al., · 2019
Earlier work this paper cites.
“MLS: A large-scale multilingual dataset for speech research,”
Vineel Pratap et al., · 2020
Earlier work this paper cites.
“Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,”
Brecht Desplanques et al., · 2020
Earlier work this paper cites.
“A survey on neural speech synthesis,”
Xu Tan et al., · 2021
Earlier work this paper cites.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim et al., · 2021
Earlier work this paper cites.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour et al., · 2021
Earlier work this paper cites.
“Lower perplexity is not always human-like,”
Tatsuki Kuribayashi et al., · 2021
Earlier work this paper cites.
“GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio,”
Guoguo Chen et al., · 2021
Earlier work this paper cites.
“Training language models to follow instructions with human feedback,”
Long Ouyang et al., · 2022
Earlier work this paper cites.
“Utmos: Utokyo-sarulab system for voicemos challenge 2022,”
Takaaki Saeki et al., · 2022
Earlier work this paper cites.
“Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,”
Edresson Casanova et al., · 2022
Earlier work this paper cites.
“High fidelity neural audio compression,”
Alexandre Défossez et al., · 2023
Cited alongside, same era.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang et al., · 2023
Cited alongside, same era.
“Speak, read and prompt: High-fidelity text-to-speech with minimal supervision,”
Eugene Kharitonov et al., · 2023
Cited alongside, same era.
“Speak foreign languages with your own voice: Cross-lingual neural codec language modeling,”
Ziqiang Zhang et al., · 2023
Cited alongside, same era.
“Audiolm: A language modeling approach to audio generation,”
Zalán Borsos et al., · 2023
Cited alongside, same era.
“Audiopalm: A large language model that can speak and listen,”
“Uniaudio: Towards universal audio generation with large language models,”
Dongchao Yang et al., · 2024
Closest in time.
“Seed-tts: A family of high-quality versatile speech generation models,”
Philip Anastassiou et al., · 2024
Closest in time.
“Voxtlm: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks,”
Soumi Maiti et al., · 2024
Closest in time.
“Towards audio language modeling-an overview,”
Haibin Wu et al., · 2024
Closest in time.
“Model alignment as prospect theoretic optimization,”
Kawin Ethayarajh et al., · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Paul K Rubenstein et al., · 2023
Cited alongside, same era.
“Soundstorm: Efficient parallel audio generation,”
Zalán Borsos et al., · 2023
Cited alongside, same era.
“A survey of reinforcement learning from human feedback,”
Timo Kaufmann et al., · 2023
Cited alongside, same era.
“Direct preference optimization: Your language model is secretly a reward model,”
Rafael Rafailov et al., · 2023
Cited alongside, same era.
Josh Achiam et al., · 2023
Cited alongside, same era.
“Fine-tuning language models with advantage-induced policy alignment,”
Banghua Zhu et al., · 2023
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford et al., · 2023
Cited alongside, same era.
Yu Meng et al., · 2024
Closest in time.
“A general theoretical paradigm to understand learning from human preferences,”
Mohammad Gheshlaghi Azar et al., · 2024
Closest in time.
“Reference-free monolithic preference optimization with odds ratio,”
Jiwoo Hong et al., · 2024
Closest in time.
Abhimanyu Dubey et al., · 2024
Closest in time.
“Speechalign: Aligning speech generation to human preferences,”
Dong Zhang et al., · 2024
Closest in time.
“Enhancing zero-shot text-to-speech synthesis with human feedback,”
Chen Chen et al., · 2024
Closest in time.
“Robust zero-shot text-to-speech synthesis with reverse inference optimization,”
Yuchen Hu et al., · 2024
Closest in time.
“Tango 2: Aligning diffusion-based text-to-audio generative models through direct preference optimization,”
Navonil Majumder et al., · 2024
Closest in time.
“Simple and controllable music generation,”
Jade Copet et al., · 2024
Closest in time.
“Autoprep: An automatic preprocessing framework for in-the-wild speech data,”
Jianwei Yu et al., · 2024
Closest in time.
“Espnet-spk: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models,”
Jee weon Jung et al., · 2024
Closest in time.
Detai Xin et al., · 2024
Closest in time.
“On the effects of heterogeneous data sources on speech-to-text foundation models,”
Jinchuan Tian et al., · 2024
Closest in time.