Fetching the paper…
Reading the bibliography…
Recent Text-to-Speech (TTS) systems trained on reading or acted corpora have achieved near human-level naturalness.
Language Models are Few-Shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 1901
Earlier work this paper cites.
DiscreTalk: Text-to-Speech as a Machine Translation Problem
Hayashi, T.; and Watanabe, S. 2020 · 2005
Earlier work this paper cites.
Multi-head Monotonic Chunkwise Attention For Online Speech Recognition
Liu, B.; Cao, S.; Sun, S.; Zhang, W.; and Ma, L. 2020 · 2005
Earlier work this paper cites.
Robust signal-to-noise ratio estimation based on waveform amplitude distribution analysis
Kim, C.; and Stern, R. M. 2008 · 2008
Earlier work this paper cites.
Computing and Visualizing Dynamic Time Warping Alignments in R: The dtw Package
Giorgino, T. 2009 · 2009
Earlier work this paper cites.
Shen, J.; Jia, Y.; Chrzanowski, M.; Zhang, Y.; Elias, I.; Zen, H.; and Wu, Y. 2020 · 2010
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Earlier work this paper cites.
The LJ Speech Dataset
Ito, K.; and Johnson, L. 2017 · 2017
Earlier work this paper cites.
Monotonic Chunkwise Attention
Chiu, C.; and Raffel, C. 2018 · 2018
Earlier work this paper cites.
Hierarchical Neural Story Generation
Fan, A.; Lewis, M.; and Dauphin, Y. 2018 · 2018
Earlier work this paper cites.
Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions
Shen, J.; Pang, R.; Weiss, R. J.; Schuster, M.; Jaitly, N.; Yang, Z.; Chen, Z.; Zhang, Y.; Wang, Y.; Skerrv-Ryan, R.; Saurous, R. A.; Agiomvrgiannakis, Y.; and Wu, Y. 2018 · 2018
Earlier work this paper cites.
Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
Wang, Y.; Stanton, D.; Zhang, Y.; Skerry-Ryan, R. J.; Battenberg, E.; Shor, J.; Xiao, Y.; Ren, F.; Jia, Y.; and Saurous, R. A. 2018 · 2018
Earlier work this paper cites.
ESPnet: End-to-End Speech Processing Toolkit
Watanabe, S.; Hori, T.; Karita, S.; Hayashi, T.; Nishitoba, J.; Unno, Y.; Enrique Yalta Soplin, N.; Heymann, J.; Wiesner, M.; Chen, N.; Renduchintala, A.; and Ochiai, T. 2018 · 2018
Cited alongside, same era.
Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS
He, M.; Deng, Y.; and He, L. 2019 · 2019
Cited alongside, same era.
Neural Speech Synthesis with Transformer Network
Li, N.; Liu, S.; Liu, Y.; Zhao, S.; and Liu, M. 2019 · 2019
Cited alongside, same era.
Building Naturalistic Emotionally Balanced Speech Corpus by Retrieving Emotional Speech From Existing Podcast Recordings
Lotfian, R.; and Busso, C. 2019 · 2019
Cited alongside, same era.
Voxceleb: Large-scale speaker verification in the wild
Nagrani, A.; Chung, J. S.; Xie, W.; and Zisserman, A. 2019 · 2019
Cited alongside, same era.
FastSpeech: Fast, Robust and Controllable Text to Speech
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Chen, G.; Chai, S.; Wang, G.; Du, J.; Zhang, W.-Q.; Weng, C.; Su, D.; Povey, D.; Trmal, J.; Zhang, J.; Jin, M.; Khudanpur, S.; Watanabe, S.; Zhao, S.; Zou, W.; Li, X.; Yao, X.; Wang, Y.; Wang, Y.; You, Z.; and Yan, Z. 2021 · 2021
Later among the works it cites.
Taming Transformers for High-Resolution Image Synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Later among the works it cites.
ESPnet2-TTS: Extending the Edge of TTS Research
Hayashi, T.; Yamamoto, R.; Yoshimura, T.; Wu, P.; Shi, J.; Saeki, T.; Ju, Y.; Yasuda, Y.; Takamichi, S.; and Watanabe, S. 2021 · 2021
Later among the works it cites.
Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
Kim, J.; Kong, J.; and Son, J. 2021 · 2021
Later among the works it cites.
On Generative Spoken Language Modeling from Raw Audio
Lakhotia, K.; Kharitonov, E.; Hsu, W.-N.; Adi, Y.; Polyak, A.; Bolte, B.; Nguyen, T.-A.; Copet, J.; Baevski, A.; Mohamed, A.; and Dupoux, E. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ren, Y.; Ruan, Y.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; and Liu, T.-Y. 2019 · 2019
Cited alongside, same era.
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92)
Yamagishi, J.; Veaux, C.; and MacDonald, K. 2019 · 2019
Cited alongside, same era.
vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
Baevski, A.; Schneider, S.; and Auli, M. 2020 · 2020
Cited alongside, same era.
Espnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
Hayashi, T.; Yamamoto, R.; Inoue, K.; Yoshimura, T.; Watanabe, S.; Toda, T.; Takeda, K.; Zhang, Y.; and Tan, X. 2020 · 2020
Cited alongside, same era.
The Curious Case of Neural Text Degeneration
Holtzman, A.; Buys, J.; Du, L.; Forbes, M.; and Choi, Y. 2020 · 2020
Cited alongside, same era.
Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
Kim, J.; Kim, S.; Kong, J.; and Yoon, S. 2020 · 2020
Cited alongside, same era.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Kong, J.; Kim, J.; and Bae, J. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Zero-Shot Text-to-Image Generation
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021 · 2021
Later among the works it cites.
AdaSpeech 3: Adaptive Text to Speech for Spontaneous Style
Yan, Y.; Tan, X.; Li, B.; Zhang, G.; Qin, T.; Zhao, S.; Shen, Y.; Zhang, W.; and Liu, T. 2021 · 2021
Later among the works it cites.
A study of latent monotonic attention variants
Zeyer, A.; Schlüter, R.; and Ney, H. 2021 · 2021
Later among the works it cites.
Denoispeech: Denoising Text to Speech with Frame-Level Noise Modeling
Zhang, C.; Ren, Y.; Tan, X.; Liu, J.; Zhang, K.; Qin, T.; Zhao, S.; and Liu, T. 2021 · 2021
Later among the works it cites.
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
Du, C.; Guo, Y.; Chen, X.; and Yu, K. 2022 · 2022
Later among the works it cites.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Press, O.; Smith, N.; and Lewis, M. 2022 · 2022
Later among the works it cites.
Dawn of the transformer era in speech emotion recognition: closing the valence gap
Wagner, J.; Triantafyllopoulos, A.; Wierstorf, H.; Schmitt, M.; Burkhardt, F.; Eyben, F.; and Schuller, B. W. 2022 · 2022
Later among the works it cites.