Fetching the paper…
Reading the bibliography…
We propose TalkNet, a non-autoregressive convolutional neural model for speech synthesis with explicit pitch and duration prediction.
T. Weijters and J. Thole, “Speech synthesis with artificial neural networks,” in IEEE International Conference on Neural Networks , 1993
1993
Earlier work this paper cites.
C. Tuerk and T. Robinson, “Speech synthesis using artificial neural networks trained on cepstral coefficients,” in Eurospeech , 1993
1993
Earlier work this paper cites.
O. Karaali, G. Corrigan, and I. Gerson, “Speech synthesis with neural networks,” in World Congress on Neural Networks , 1996
1996
Earlier work this paper cites.
O. Karaali, G. Corrigan, N. Massey, C. Miller, O. Schnurr, and A. Mackie, “A high quality text-to-speech system composed of multiple neural networks,” in ICASSP , 1998
1998
Earlier work this paper cites.
P. Taylor, Text-to-Speech Synthesis . Cambridge University Press, 2009
2009
Earlier work this paper cites.
2014
Earlier work this paper cites.
H. Zen and H. Sak, “Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis,” in ICASSP , 2015
2015
Earlier work this paper cites.
R. Yamamoto, “pysptk,” https://github.com/r9y9/pysptk , 2015
2015
Earlier work this paper cites.
H. Zen, Y. Agiomyrgiannakis, N. Egberts, F. Henderson, and P. Szczepaniak, “Fast, compact, and high quality LSTM-RNN based statistical parametric speech synthesizers for mobile devices,” in INTERSPEECH , 2016
2016
Earlier work this paper cites.
Z. Wu, O. Watts, and S. King, “Merlin: An open source neural network speech synthesis system,” in ISCA Speech Synthesis Workshop , 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Gibiansky, S. Arik, G. Diamos, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, “Deep Voice 2: Multi-speaker neural text-to-speech,” in NIPS , 2017
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, “Tacotron: Towards end-to-end speech synthesis,” in INTERSPEECH , 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017
2017
Cited alongside, same era.
K. Ito and L. Johnson, “The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Cited alongside, same era.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal Forced Aligner: trainable text-speech alignment using Kaldi,” in Conference of the International Speech Communication Association , 2017
K. Peng, W. Ping, Z. Song, and K. Zhao, “Parallel neural text-to-speech,” arXiv:1905.08459 , 2019
2019
Later among the works it cites.
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, “Neural speech synthesis with Transformer network,” in AAAI , 2019
2019
Later among the works it cites.
K. Park and J. Kim, “g2pe,” https://github.com/Kyubyong/g2p , 2019
2019
Later among the works it cites.
H. Zen, R. Clark, R. J. Weiss, V. Dang, Y. Jia, Y. Wu, Y. Zhang, and Z. Chen, “LibriTTS: A corpus derived from LibriSpeech for text-to-speech,” in INTERSPEECH , 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning Wavenet on mel spectrogram predictions,” in ICASSP , 2018
2018
Cited alongside, same era.
W. Ping, K. Peng, A. Gibiansky, S. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, “Deep Voice 3: 2000-speaker neural text-to-speech,” in ICLR , 2018
2018
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “FastSpeech: Fast, robust and controllable text to speech,” in NeurIPS , 2019
2019
Cited alongside, same era.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A flow-based generative network for speech synthesis,” in ICASSP , 2019
2019
Cited alongside, same era.
2019
Later among the works it cites.
2020
Later among the works it cites.
S. Kriman, S. Beliaev, B. Ginsburg, J. Huang, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, and Y. Zhang, “QuartzNet: deep automatic speech recognition with 1D time-channel separable convolutions,” in ICASSP , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu, “Non-attentive tacotron: Robust and controllable neural tts synthesis including unsupervised duration modeling,” 2020
2020
Later among the works it cites.
A. Łańcucki, “Fastpitch: Parallel text-to-speech with pitch prediction,” in ICASSP , 2021
2021
Closest in time.