Fetching the paper…
Reading the bibliography…
Neural text-to-speech (TTS) has achieved human-like synthetic speech for single-speaker, single-language synthesis.
N. L. Technology, “NST Swedish speech synthesis,” https://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-18/ , 2003
2003
Earlier work this paper cites.
I. T. Union. Recommendation G.191: Software Tools and Audio Coding Standardization. (2005, Nov 11). [Online]. Available: https://www.itu.int/rec/T-REC-P.56/en
2005
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing data using t-SNE.” Journal of machine learning research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
T. Schultz, N. T. Vu, and T. Schlippe, “GlobalPhone: A multilingual text & speech database in 20 languages,” in Proc. ICASSP , 2013, pp. 8126–8130
2013
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. W. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” in ISCA, September 2016 . ISCA, 2016, p. 125
2016
Earlier work this paper cites.
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. V. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in Proc. Interspeech . ISCA, 2017, pp. 4006–4010
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJ speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions,” in Proc. ICASSP . IEEE, 2018, pp. 4779–4783
2018
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, z. Chen, P. Nguyen, R. Pang, I. Lopez Moreno, and Y. Wu, “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Proc. NIPS , vol. 31, 2018, pp. 4480–4490
2018
Earlier work this paper cites.
S. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in Proc. NIPS , vol. 31, 2018
2018
Earlier work this paper cites.
Y. Wang, D. Stanton, Y. Zhang, R.-S. Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, “Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,” in Proc. ICML . PMLR, 2018, pp. 5180–5189
2018
Earlier work this paper cites.
D. R. Mortensen, S. Dalmia, and P. Littell, “Epitran: Precision G2P for many languages,” in Proc. LREC , Paris, France, May 2018
2018
Earlier work this paper cites.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with Transformer network,” in Proc. AAAI , 2019, pp. 6706–6713
2019
Earlier work this paper cites.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
M. Chen, M. Chen, S. Liang, J. Ma, L. Chen, S. Wang, and J. Xiao, “Cross-lingual, multi-speaker text-to-speech synthesis using neural speaker embedding,” in Proc. Interspeech . ISCA, 2019, pp. 2105–2109
2019
Earlier work this paper cites.
B. Li, Y. Zhang, T. N. Sainath, Y. Wu, and W. Chan, “Bytes are all you need: End-to-end multilingual speech recognition and synthesis with bytes,” in Proc. ICASSP . IEEE, 2019, pp. 5621–5625
2019
Earlier work this paper cites.
K. Park and T. Mulc, “CSS10: A collection of single speaker speech datasets for 10 languages,” in Proc. Interspeech , 2019, pp. 1566–1570
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2019
Earlier work this paper cites.
T. Nekvinda and O. Dusek, “One model, many languages: Meta-learning for multilingual text-to-speech,” in Proc. Interspeech . ISCA, 2020, pp. 2972–2976
2020
Earlier work this paper cites.
E. Cooper, C.-I. Lai, Y. Yasuda, F. Fang, X. Wang, N. Chen, and J. Yamagishi, “Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,” in Proc. ICASSP , 2020, pp. 6184–6188
2020
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “Wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NIPS , 2020, pp. 12 449–12 460
2020
Earlier work this paper cites.
J. Yang and L. He, “Towards universal text-to-speech,” in Proc. Interspeech , 2020, pp. 3171–3175
2020
Earlier work this paper cites.
M. Staib, T. H. Teh, A. Torresquintero, D. S. R. Mohan, L. Foglianti, R. Lenain, and J. Gao, “Phonological Features for 0-Shot Multilingual Speech Synthesis,” in Proc. Interspeech 2020 , 2020, pp. 2942–2946
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. NIPS , vol. 33. Curran Associates, Inc., 2020, pp. 17 022–17 033
2020
Earlier work this paper cites.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A large-scale multilingual dataset for speech research,” in Proc. Interspeech , 2020, p. 2757–2761
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in Proc. ICLR , 2021
2021
Earlier work this paper cites.
Y. Jia, H. Zen, J. Shen, Y. Zhang, and Y. Wu, “PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS,” Proc. Interspeech , pp. 151–155, 2021
2021
Earlier work this paper cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
H.-S. Choi, J. Lee, W. Kim, J. Lee, H. Heo, and K. Lee, “Neural analysis and synthesis: Reconstructing speech from self-supervised representations,” in Proc. NIPS , vol. 34, 2021, pp. 16 251–16 265
2021
Cited alongside, same era.
A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” in Proc. ASRU , 2021, pp. 914–921
2021
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
D. Berrebbi, J. Shi, B. Yan, O. López-Francisco, J. Amith, and S. Watanabe, “Combining spectral and self-supervised features for low resource speech recognition and translation,” in Proc. Interspeech , 2022, pp. 3533–3537
2022
Later among the works it cites.
H. Guo, F. Xie, F. K. Soong, X. Wu, and H. Meng, “A multi-stage multi-codebook VQ-VAE approach to high-performance neural TTS,” in Proc. Interspeech . ISCA, 2022, pp. 1611–1615
2022
Later among the works it cites.
Data-baker. Chinese Standard Mandarin Speech Copus. (2022, Nov). [Online]. Available: https://www.data-baker.com/open_source.html
2022
Later among the works it cites.
P. Do, M. Coler, J. Dijkstra, and E. Klabbers, “Text-to-speech for under-resourced languages: Phoneme mapping and source language selection in transfer learning,” in Proc. ELRA/ISCA SIG on Under-Resourced Languages , 2022, pp. 16–22
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Casanova, C. Shulby, E. Gölge, N. M. Müller, F. S. de Oliveira, A. Candido Jr., A. da Silva Soares, S. M. Aluisio, and M. A. Ponti, “SC-GlowTTS: An efficient zero-shot multi-speaker text-to-speech model,” in Proc. Interspeech , 2021, pp. 3645–3649
2021
Cited alongside, same era.
D. Xin, T. Komatsu, S. Takamichi, and H. Saruwatari, “Disentangled speaker and language representations using mutual information minimization and domain adaptation for cross-lingual TTS,” in Proc. ICASSP . IEEE, 2021, pp. 6608–6612
2021
Cited alongside, same era.
A. Lancucki, “FastPitch: Parallel text-to-speech with pitch prediction,” in Proc. ICASSP . IEEE, 2021, pp. 6588–6592
2021
Cited alongside, same era.
K. J. Shih, R. Valle, R. Badlani, A. Lancucki, W. Ping, and B. Catanzaro, “RAD-TTS: Parallel flow-based TTS with robust alignment learning and diverse synthesis,” in Proc. ICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models , 2021
2021
Cited alongside, same era.
J. Kim, J. Kong, and J. Son, “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” in Proc. ICML . PMLR, 2021, pp. 5530–5540
2021
Cited alongside, same era.
D. Wells and K. Richmond, “Cross-lingual transfer of phonological features for low-resource speech synthesis,” in Proc. SSW , 2021, pp. 160–165
2021
Cited alongside, same era.
J.-h. Lin, Y. Y. Lin, C.-M. Chien, and H.-y. Lee, “S2VC: A framework for any-to-any voice conversion with self-supervised pretrained representations,” in Proc. Interspeech , 2021, pp. 836–840
2021
Cited alongside, same era.
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Unsupervised cross-lingual representation learning for speech recognition,” in Proc. Interspeech . ISCA, Aug. 2021, pp. 2426–2430
2021
Cited alongside, same era.
T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-SaruLab system for VoiceMOS Challenge 2022,” in Proc. Interspeech , 2022, pp. 4521–4525
2022
Later among the works it cites.
T. Saeki, S. Maiti, X. Li, S. Watanabe, S. Takamichi, and H. Saruwatari, “Learning to speak from text: Zero-shot multilingual text-to-speech with unsupervised text pretraining,” in Proc. IJCAI , 2023
2023
Closest in time.
Y. A. Li, C. Han, X. Jiang, and N. Mesgarani, “Phoneme-level BERT for enhanced prosody of text-to-speech with grapheme predictions,” in Proc. ICASSP . IEEE, 2023, pp. 1–5
2023
Closest in time.
L. T. Nguyen, T. Pham, and D. Q. Nguyen, “XPhoneBERT: A pre-trained multilingual model for phoneme representations for text-to-speech,” in Proc. Interspeech , 2023, pp. 5506–5510
2023
Closest in time.
2023
Closest in time.
S. Liu, Y. Guo, C. Du, X. Chen, and K. Yu, “DSE-TTS: Dual speaker embedding for cross-lingual text-to-speech,” in Proc. Interspeech , 2023, pp. 616–620
2023
Closest in time.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi et al. , “Audiolm: a language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Closest in time.
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour, “Speak, read and prompt: High-fidelity text-to-speech with minimal supervision,” Transactions of the Association for Computational Linguistics , vol. 11, pp. 1703–1718, 2023
2023
Closest in time.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Transactions on Machine Learning Research , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
R. Badlani, R. Valle, K. J. Shih, J. F. Santos, S. Gururani, and B. Catanzaro, “RAD-MMM: Multilingual multiaccented multispeaker text to speech,” in Proc. Interspeech , 2023, pp. 626–630
2023
Closest in time.
Y.-J. Zhang, C. Zhang, W. Song, Z. Zhang, Y. Wu, and X. He, “Prosody modelling with pre-trained cross-utterance representations for improved speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 2812–2823, 2023
2023
Closest in time.
L.-W. Chen, S. Watanabe, and A. Rudnicky, “A vector quantized approach for text to speech synthesis on real-world spontaneous speech,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 11, 2023, pp. 12 644–12 652
2023
Closest in time.
D. Wells, K. Richmond, and W. Lamb, “A Low-Resource Pipeline for Text-to-Speech from Found Data With Application to Scottish Gaelic,” in Proc. INTERSPEECH 2023 , 2023, pp. 4324–4328
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
H. Guo, F. Xie, X. Wu, F. K. Soong, and H. Meng, “MSMC-TTS: Multi-stage multi-codebook VQ-VAE based neural TTS,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 1811–1824, 2023
2023
Closest in time.
M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar et al. , “Voicebox: Text-guided multilingual universal speech generation at scale,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
T. Saeki, S. Maiti, X. Li, S. Watanabe, S. Takamichi, and H. Saruwatari, “Text-inductive graphone-based language adaptation for low-resource speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 1829–1844, 2024
2024
Closest in time.
2024
Closest in time.
Y. A. Li, C. Han, V. Raghavan, G. Mischler, and N. Mesgarani, “Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Y. Zhang, R. J. Weiss, H. Zen, Y. Wu, Z. Chen, R. J. Skerry-Ryan, Y. Jia, A. Rosenberg, and B. Ramabhadran, “Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning,” in Proc. Interspeech . ISCA, 2019, pp. 2080–2084
2084
Closest in time.