Fetching the paper…
Reading the bibliography…
We present RALL-E, a robust language modeling method for text-to-speech (TTS) synthesis.
Sequence transduction with recurrent neural networks
A. Graves · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
World: a vocoder-based high-quality speech synthesis system for real-time applications
M. Morise, F. Yokomori, and K. Ozawa · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Forward attention in sequence-to-sequence acoustic modeling for speech synthesis
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai · 2018
Earlier work this paper cites.
Robust sequence-to-sequence acoustic modeling with stepwise monotonic attention for neural tts
M. He, Y. Deng, and L. He · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
Fastspeech: Fast, robust and controllable text to speech
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu · 2019
Earlier work this paper cites.
Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion
H. Sun, X. Tan, J.-W. Gan, H. Liu, S. Zhao, T. Qin, and T.-Y. Liu · 2019
Earlier work this paper cites.
Libritts: A corpus derived from librispeech for text-to-speech
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Multispeech: Multi-speaker text to speech with transformer
M. Chen, X. Tan, Y. Ren, J. Xu, H. Sun, S. Zhao, T. Qin, and T.-Y. Liu · 2020
Cited alongside, same era.
Conformer: Convolution-augmented Transformer for Speech Recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang · 2020
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux · 2020
Cited alongside, same era.
Utmos: Utokyo-sarulab system for voicemos challenge 2022
T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari · 2022
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al · 2023
Later among the works it cites.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour · 2023
Later among the works it cites.
Audiopalm: A large language model that can speak and listen
P. K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. d. C. Quitry, P. Chen, D. E. Badawy, W. Han, E. Kharitonov, et al · 2023
Later among the works it cites.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MLS: A Large-Scale Multilingual Dataset for Speech Research
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert · 2020
Cited alongside, same era.
J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu · 2020
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi · 2021
Cited alongside, same era.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, et al · 2022
Cited alongside, same era.
High fidelity neural audio compression
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi · 2022
Cited alongside, same era.
Neural codec language models are zero-shot text to speech synthesizers
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, et al
Cited in the paper.
K. Shen, Z. Ju, X. Tan, Y. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian · 2023
Later among the works it cites.
Uniaudio: An audio foundation model toward universal audio generation
D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, S. Zhao, J. Bian, X. Wu, et al · 2023
Later among the works it cites.
Vall-t: Decoder-only generative transducer for robust and decoding-controllable text-to-speech
C. Du, Y. Guo, H. Wang, Y. Yang, Z. Niu, S. Wang, H. Zhang, X. Chen, and K. Yu · 2024
Closest in time.
Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Z. Ju, Y. Wang, K. Shen, X. Tan, D. Xin, D. Yang, Y. Liu, Y. Leng, K. Song, S. Tang, et al · 2024
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar, et al · 2024
Closest in time.
Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering
Y. Song, Z. Chen, X. Wang, Z. Ma, and X. Chen · 2024
Closest in time.