Fetching the paper…
Reading the bibliography…
Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation given a speech prompt.
2012
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in ACL , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar et al. , “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
Y. He, T. N. Sainath, R. Prabhavalkar et al. , “Streaming end-to-end speech recognition for mobile devices,” in IEEE ICASSP , 2019, pp. 6381–6385
2019
Earlier work this paper cites.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A corpus derived from librispeech for text-to-speech,” in ISCA Interspeech , 2019, pp. 1526–1530
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder et al. , “Language models are few-shot learners,” in NeurIPS , 2020
2020
Earlier work this paper cites.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. A. Kudinov, “Grad-TTS: A diffusion probabilistic model for text-to-speech,” in ICML , vol. 139, 2021, pp. 8599–8608
2021
Cited alongside, same era.
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed et al. , “On generative spoken language modeling from raw audio,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1336–1354, 2021
2021
Cited alongside, same era.
C. Du, Y. Guo, X. Chen, and K. Yu, “VQTTS: high-fidelity text-to-speech synthesis with self-supervised VQ acoustic feature,” in ISCA Interspeech , 2022, pp. 1596–1600
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Yao, L. Guo, X. Yang, W. Kang, F. Kuang, Y. Yang, Z. Jin, L. Lin, and D. Povey, “Zipformer: A faster and better encoder for automatic speech recognition,” ICLR , vol. 2310.11230, 2023
2023
Later among the works it cites.
M. Kim, M. Jeong, B. J. Choi, D. Lee, and N. S. Kim, “Transduce and speak: Neural transducer for text-to-speech with semantic token prediction,” IEEE ASRU , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Betker, “Better speech synthesis through scaling,” arXiv preprint arXiv:2305.07243 , 2023
2023
Cited alongside, same era.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in ICML , vol. 202, 2023, pp. 28 492–28 518
2023
Later among the works it cites.
C. Du, Y. Guo, F. Shen, Z. Liu, Z. Liang, X. Chen, S. Wang, H. Zhang, and K. Yu, “UniCATS: A unified context-aware text-to-speech framework with contextual vq-diffusion and vocoding,” AAAI , 2024
2024
Closest in time.