Fetching the paper…
Reading the bibliography…
This paper describes a novel text-to-speech (TTS) technique based on deep convolutional neural networks (CNN), without use of any recurrent units.
“TTS synthesis with bidirectional LSTM based recurrent neural networks,”
Y. Fan et al., · 1968
Earlier work this paper cites.
“Signal estimation from modified short-time fourier transform,”
D. Griffin and J. Lim, · 1984
Earlier work this paper cites.
“The handbook of brain theory and neural networks,”
Y. LeCun and Y. Bengio, · 1998
Earlier work this paper cites.
“Real-time signal estimation from modified short-time fourier transform magnitude spectra,”
X. Zhu et al., · 2007
Earlier work this paper cites.
“CrowdMOS: An approach for crowdsourcing mean opinion score studies,”
F. Ribeiro et al., · 2011
Earlier work this paper cites.
“Statistical parametric speech synthesis using deep neural networks,”
H. Zen et al., · 2013
Earlier work this paper cites.
“Learning phrase representations using RNN encoder-decoder for statistical machine translation,”
K. Cho et al., · 2014
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
I. Sutskever et al., · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
D Bahdanau et al., · 2014
Earlier work this paper cites.
“Convolutional neural networks for sentence classification,”
Y. Kim, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis,”
H. Zen and H. Sak, · 2015
Earlier work this paper cites.
“An investigation of recurrent neural network architectures for statistical parametric speech synthesis.,”
S. Achanta et al., · 2015
Earlier work this paper cites.
“A neural conversational model,”
O. Vinyals and Q. Le, · 2015
Cited alongside, same era.
“Character-level convolutional networks for text classification,”
X. Zhang et al., · 2015
Cited alongside, same era.
“Training very deep networks,”
R. K. Srivastava et al., · 2015
Cited alongside, same era.
“Chainer: A next-generation open source framework for deep learning,”
S. Tokui et al., · 2015
Cited alongside, same era.
“Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification,”
K. He et al., · 2015
Cited alongside, same era.
Deep Learning
I. Goodfellow et al., · 2016
Cited alongside, same era.
“Multi-scale context aggregation by dilated convolutions,”
F. Yu and V. Koltun, · 2016
Later among the works it cites.
“Tacotron: Towards end-to-end speech synthesis,”
Y. Wang et al., · 2017
Closest in time.
“Implementation of Google’s Tacotron in TensorFlow,” 2017,
A. Barron, · 2017
Closest in time.
“A TensorFlow implementation of Tacotron: A fully end-to-end text-to-speech synthesis model,” 2017,
K. Park, · 2017
Closest in time.
“Tacotron speech synthesis implemented in TensorFlow, with samples and a pre-trained model,” 2017,
K. Ito, · 2017
Closest in time.
“PyTorch implementation of Tacotron speech synthesis model,” 2017,
R. Yamamoto, · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Investigating gated recurrent networks for speech synthesis,”
Z. Wu and S. King, · 2016
Cited alongside, same era.
“WaveNet: A generative model for raw audio,”
A. van den Oord et al., · 2016
Cited alongside, same era.
“Building end-to-end dialogue systems using generative hierarchical neural network models.,”
I. V. Serban et al., · 2016
Cited alongside, same era.
“Neural machine translation in linear time,”
N. Kalchbrenner et al., · 2016
Cited alongside, same era.
“Language modeling with gated convolutional networks,”
Y. N. Dauphin et al., · 2016
Cited alongside, same era.
“Quasi-recurrent neural networks,”
J. Bradbury et al., · 2016
Cited alongside, same era.
Closest in time.
“Convolutional sequence to sequence learning,”
J. Gehring et al., · 2017
Closest in time.
“Char2wav: End-to-end speech synthesis,”
J. Sotelo et al., · 2017
Closest in time.
“Deep voice: Real-time neural text-to-speech,”
S. Arik et al., · 2017
Closest in time.
“Deep voice 2: Multi-speaker neural text-to-speech,”
S. Arik et al., · 2017
Closest in time.
“StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks,”
H. Zhang et al., · 2017
Closest in time.
“The LJ speech dataset,” 2017,
K. Ito, · 2017
Closest in time.