Fetching the paper…
Reading the bibliography…
We present FastPitch, a fully-parallel text-to-speech model based on FastSpeech, conditioned on fundamental frequency contours.
“Accurate short-term analysis of the fundamental frequency and the harmonics-to-noise ratio of a sampled sound,”
P. Boersma, · 1993
Earlier work this paper cites.
Example of the Glicko-2 system
M. E. Glickman, · 2013
Earlier work this paper cites.
“Using deep bidirectional recurrent neural networks for prosodic-target prediction in a unit-selection text-to-speech system.,”
R. Fernandez, A. Rendel, B. Ramabhadran, and R. Hoory, · 2015
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Earlier work this paper cites.
“Fast and easy crowdsourced perceptual audio evaluation,”
M. Cartwright, B. Pardo, G. J. Mysore, and M. Hoffman, · 2016
Earlier work this paper cites.
“The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset , 2017
K. Ito, · 2017
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using kaldi,”
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Large-Scale Speaker Ranking from Crowdsourced Pairwise Listener Ratings,”
T. Baumann, · 2017
Earlier work this paper cites.
“Neural tts voice conversion,”
Z. Kons, S. Shechtman, A. Sorin, R. Hoory, C. Rabinovitz, and E. Da Silva Morais, · 2018
Cited alongside, same era.
“Deep voice 3: Scaling text-to-speech with convolutional sequence learning,”
W. Ping, K. Peng, A. Gibiansky, S. Ö. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, · 2018
Cited alongside, same era.
“Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, and et al., · 2018
Cited alongside, same era.
“Mixed precision training,”
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, · 2018
Cited alongside, same era.
“Skill rating for generative models,”
C. Olsson, S. Bhupatiraju, T. Brown, A. Odena, and I. Goodfellow, · 2018
Cited alongside, same era.
“High fidelity speech synthesis with adversarial networks,”
M. Bińkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan, · 2019
Later among the works it cites.
“Durian: Duration informed attention network for multimodal synthesis,”
C. Yu, H. Lu, N. Hu, M. Yu, C. Weng, K. Xu, P. Liu, D. Tuo, S. Kang, G. Lei, D. Su, and D. Yu, · 2019
Later among the works it cites.
“Representation Mixing for TTS Synthesis,”
K. Kastner, J. F. Santos, Y. Bengio, and A. Courville, · 2019
Later among the works it cites.
“Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search,”
J. Kim, S. Kim, J. Kong, and S. Yoon, · 2020
Closest in time.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Fastspeech: Fast, robust and controllable text to speech,”
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2019
Cited alongside, same era.
R. Valle, J. Li, R. Prenger, and B. Catanzaro, · 2019
Cited alongside, same era.
“Waveglow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2019
Cited alongside, same era.
“Neural speech synthesis with transformer network,”
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, · 2019
Cited alongside, same era.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2020
Closest in time.
ForwardTacotron
C. Schäfer, · 2020
Closest in time.
“AlignTTS: Efficient Feed-Forward Text-to-Speech System without Explicit Alignment,”
Z. Zeng, J. Wang, N. Cheng, T. Xia, and J. Xiao, · 2020
Closest in time.
“Large batch optimization for deep learning: Training bert in 76 minutes,”
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh, · 2020
Closest in time.
“Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis,”
R. Valle, K. J. Shih, R. Prenger, and B. Catanzaro, · 2021
Closest in time.