Fetching the paper…
Reading the bibliography…
Recent advancements in text-to-speech (TTS) powered by language models have showcased remarkable capabilities in achieving naturalness and zero-shot voice cloning.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems (NIPS) , vol. 30, 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, “Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4779–4783
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” 2019
2019
Earlier work this paper cites.
N. Kitaev, L. Kaiser, and A. Levskaya, “Reformer: The Efficient Transformer,” in International Conference on Learning Representations (ICLR) , 2020
2020
Earlier work this paper cites.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-Light: A Benchmark for ASR with Limited or No Supervision,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7669–7673
2020
Earlier work this paper cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,” International Conference on Learning Representations (ICLR) , 2021
2021
Earlier work this paper cites.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in International Conference on Machine Learning (ICML) , 2021, pp. 5530–5540
2021
Earlier work this paper cites.
Z. Shen, M. Zhang, H. Zhao, S. Yi, and H. Li, “Efficient attention: Attention with linear complexities,” in IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , 2021, pp. 3531–3539
2021
Earlier work this paper cites.
S. Zhai, W. Talbott, N. Srivastava, C. Huang, H. Goh, R. Zhang, and J. Susskind, “An Attention Free Transformer,” in International Conference on Learning Representations (ICLR) , 2021
2021
Earlier work this paper cites.
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman, “MaskGIT: Masked Generative Image Transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022, pp. 11 315–11 325
2022
Earlier work this paper cites.
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness,” in Advances in Neural Information Processing Systems (NeurIPS) , 2022, pp. 16 344–16 359
2022
Earlier work this paper cites.
A. Gu, K. Goel, and C. Ré, “Efficiently Modeling Long Sequences with Structured State Spaces,” in International Conference on Learning Representations (ICLR) , 2022
2022
Earlier work this paper cites.
R. Badlani, A. Łańcucki, K. J. Shih, R. Valle, W. Ping, and B. Catanzaro, “One TTS Alignment to Rule Them All,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 6092–6096
2022
Cited alongside, same era.
J. Betker, “ocotillo - a fast, accurate and super simple speech recognition model,” https://github.com/neonbjb/ocotillo, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar et al. , “Voicebox: Text-guided Multilingual Universal Speech Generation at Scale,” Conference in Neural Information Processing Systems (NeurIPS) , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and Controllable Music Generation,” in Advances in Neural Information Processing Systems (NeurIPS) , 2023
2023
Later among the works it cites.
E. Casanova, J. Weber, C. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, “YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone,” 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
B. Peng, E. Alcaide, Q. G. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. Chung, M. Grella, G. Kranthikiran, X. He, H. Hou, P. Kazienko, J. Kocoń, J. Kong, B. Koptyra, H. Lau, K. S. I. Mantri, F. Mom, A. Saito, X. Tang, B. Wang, J. S. Wind, S. Wozniak, R. Zhang, Z. Zhang, Q. Zhao, P. Zhou, J. Zhu, and R. Zhu, “RWKV: Reinventing RNNs for the Transformer Era,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Betker, “Better Speech Synthesis through Scaling,” arXiv preprint arXiv:2305.07243 , 2023
2023
Cited alongside, same era.
Suno, “Bark,” https://github.com/suno-ai/bark, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Sun, L. Dong, S. Huang, S. Ma, Y. Xia, J. Xue, J. Wang, and F. Wei, “Retentive Network: A Successor to Transformer for Large Language Models,” in submitted to International Conference on Learning Representations (ICLR) , 2023
2023
Cited alongside, same era.
H. Siuzdak, “Vocos: Closing the Gap between Time-Domain and Fourier-based Neural Vocoders for High-Quality Audio Synthesis,” in International Conference on Learning Representations (ICLR) , 2023
2023
Later among the works it cites.
E. Georgiou, K. Kritsis, G. Paraskevopoulos, A. Katsamanis, V. Katsouros, and A. Potamianos, “Regotron: Regularizing the Tacotron2 Architecture Via Monotonic Alignment Loss,” in IEEE Spoken Language Technology Workshop (SLT) , 2023, pp. 977–983
2023
Later among the works it cites.
K. Shen, Z. Ju, X. Tan, Y. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian, “NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers,” in Submitted to International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
MetaVoice, “Metavoice-1b,” https://github.com/metavoiceio/metavoice-src, 2024
2024
Closest in time.
T. Parcollet, R. van Dalen, S. Zhang, and S. Bhattacharya, “SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition,” in submitted to Annual Conference on Neural Information Processing Systems (NeurIPS) , 2024
2024
Closest in time.
A. Gu and T. Dao, “Mamba: Linear-Time Sequence Modeling with Selective State Spaces,” in Submitted to International Conference on Learning Representations (ICLR) , 2024
2024
Closest in time.
J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu, “RoFormer: Enhanced Transformer with Rotary Position Embedding,” Neurocomputing , vol. 568, p. 127063, 2024
2024
Closest in time.