Fetching the paper…
Reading the bibliography…
While autoregressive speech token generation models produce speech with remarkable variety and naturalness, their inherent lack of controllability often results in issues such as hallucinations and undesired vocalizations that do not conform to conditioning inputs.
Multi-objective optimisation using evolutionary algorithms: an introduction
Deb, K · 2011
Earlier work this paper cites.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Zhang, Z., Zhou, L., Wang, C., Chen, S., Wu, Y., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al · 2011
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
LibriTTS: A corpus derived from librispeech for text-to-speech
Zen, H., Dang, V., Clark, R., Zhang, Y., Weiss, R. J., Jia, Y., Chen, Z., and Wu, Y · 2019
Earlier work this paper cites.
MLS: A Large-Scale Multilingual Dataset for Speech Research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2020
Earlier work this paper cites.
Hi-Fi Multi-Speaker English TTS Dataset
Bakhturina, E., Lavrukhin, V., Ginsburg, B., and Zhang, Y · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2021
Earlier work this paper cites.
SoundStream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Earlier work this paper cites.
One tts alignment to rule them all
Badlani, R., Łańcucki, A., Shih, K. J., Valle, R., Ping, W., and Catanzaro, B · 2022
Earlier work this paper cites.
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone
Casanova, E., Weber, J., Shulby, C. D., Junior, A. C., Gölge, E., and Ponti, M. A · 2022
Earlier work this paper cites.
Titanet: Neural model for speaker representation with 1d depth-wise separable convolutions and global context
Koluguri, N. R., Park, T., and Ginsburg, B · 2022
Earlier work this paper cites.
Titanet: Neural model for speaker representation with 1d depth-wise separable convolutions and global context
Koluguri, N. R., Park, T., and Ginsburg, B · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Robust speech recognition via large-scale weak supervision, 2022
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2022
Earlier work this paper cites.
Audiolm: a language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M., et al · 2023
Earlier work this paper cites.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2023
Cited alongside, same era.
Ace-vc: Adaptive and controllable voice conversion using explicitly disentangled self-supervised speech representations
Hussain, S., Neekhara, P., Huang, J., Li, J., and Ginsburg, B · 2023
Cited alongside, same era.
Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio
Kumar, A., Tan, K., Ni, Z., Manocha, P., Zhang, X., Henderson, E., and Xu, B · 2023
Cited alongside, same era.
CML-TTS: A multilingual dataset for speech synthesis in low-resource languages
Oliveira, F. S., Casanova, E., Junior, A. C., Soares, A. S., and Galvão Filho, A. R · 2023
Cited alongside, same era.
Stay on topic with classifier-free guidance
Sanchez, G., Fan, H., Spangher, A., Levi, E., Ammanamanchi, P. S., and Biderman, S · 2023
Cited alongside, same era.
High-fidelity audio compression with improved rvqgan
Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K · 2024
Later among the works it cites.
Spectral codecs: Spectrogram-based audio codecs for high quality speech synthesis
Langman, R., Jukić, A., Dhawan, K., Koluguri, N. R., and Ginsburg, B · 2024
Later among the works it cites.
Voicebox: Text-guided multilingual universal speech generation at scale
Le, M., Vyas, A., Shi, B., Karrer, B., Sari, L., Moritz, R., Williamson, M., Manohar, V., Adi, Y., Mahadeokar, J., et al · 2024
Later among the works it cites.
Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models
Li, Y. A., Han, C., Raghavan, V., Mischler, G., and Mesgarani, N · 2024
Later among the works it cites.
Finite scalar quantization: VQ-VAE made simple
Mentzer, F., Minnen, D., Agustsson, E., and Tschannen, M · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al · 2023
Cited alongside, same era.
Efficient sequence transduction by jointly predicting tokens and durations
Xu, H., Jia, F., Majumdar, S., Huang, H., Watanabe, S., and Ginsburg, B · 2023
Cited alongside, same era.
Nemotron-4 340b technical report
Adler, B., Agarwal, N., Aithal, A., Anh, D. H., Bhattacharya, P., Brundyn, A., Casper, J., Catanzaro, B., Clay, S., Cohen, J., et al · 2024
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Guo, Z. D., Piot, B., Munos, R., Rowland, M., Valko, M., and Calandriello, D · 2024
Cited alongside, same era.
Xtts: a massively multilingual zero-shot text-to-speech model
Casanova, E., Davis, K., Gölge, E., Göknar, G., Gulea, I., Hart, L., Aljafari, A., Meyer, J., Morais, R., Olayemi, S., et al · 2024
Cited alongside, same era.
Parakeet, 2024
Darefsky, J., Zhu, G., and Duan, Z · 2024
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
A comprehensive survey of hallucination in large language, image, video and audio foundation models
Sahoo, P., Meharia, P., Ghosh, A., Saha, S., Jain, V., and Chadha, A · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Later among the works it cites.
Classifier-free guidance in llms safety
Smirnov, R · 2024
Later among the works it cites.
Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering
Song, Y., Chen, Z., Wang, X., Ma, Z., and Chen, X · 2024
Later among the works it cites.
SpeechX: Neural codec language model as a versatile speech transformer
Wang, X., Thakker, M., Chen, Z., Kanda, N., Eskimez, S. E., Chen, S., Tang, M., Liu, S., Li, J., and Yoshioka, T · 2024
Later among the works it cites.
Uniaudio: Towards universal audio generation with large language models
Yang, D., Tian, J., Tan, X., Huang, R., Liu, S., Guo, H., Chang, X., Shi, J., Bian, J., Zhao, Z., et al · 2024
Later among the works it cites.
Speechalign: Aligning speech generation to human preferences
Zhang, D., Li, Z., Li, S., Zhang, X., Wang, P., Zhou, Y., and Qiu, X · 2024
Later among the works it cites.
Low frame-rate speech codec: a codec designed for fast high-quality speech llm training and inference
Casanova, E., Langman, R., Neekhara, P., Hussain, S., Li, J., Ghosh, S., Jukić, A., and Lee, S.-g · 2025
Closest in time.