Fetching the paper…
Reading the bibliography…
We present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks.
Praat, a system for doing phonetics by computer
Boersma, Paulus Petrus Gerardus et al · 2002
Earlier work this paper cites.
Production Rendering, Design and Implementation
Stephenson, Ian · 2005
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Graves, Alex, Fernández, Santiago, Gomez, Faustino, and Schmidhuber, Jürgen · 2006
Earlier work this paper cites.
The CMU pronunciation dictionary 0.7
Weide, R · 2008
Earlier work this paper cites.
Text-to-Speech Synthesis
Taylor, Paul · 2009
Earlier work this paper cites.
Crowdmos: An approach for crowdsourcing mean opinion score studies
Ribeiro, Flávio, Florêncio, Dinei, Zhang, Cha, and Seltzer, Michael · 2011
Earlier work this paper cites.
The blizzard challenge 2013–indian language task
Prahallad, Kishore, Vadapalli, Anandaswarup, Elluru, Naresh, et al · 2013
Earlier work this paper cites.
Statistical parametric speech synthesis using deep neural networks
Zen, Heiga, Senior, Andrew, and Schuster, Mike · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, Junyoung, Gulcehre, Caglar, Cho, KyungHyun, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, Martín, Agarwal, Ashish, Barham, Paul, Brevdo, Eugene, Chen, Zhifeng, Citro, Craig, Corrado, Greg S., Davis, Andy, Dean, Jeffrey, Devin, Matthieu, Ghemawat, Sanjay, Goodfellow, Ian, Harp, Andrew, Irving, Geoffrey, Isard, Michael, Jia, Yangqing, Jozefowicz, Rafal, Kaiser, Lukasz, Kudlur, Manjunath, Levenberg, Josh, Mané, Dan, Monga, Rajat, Moore, Sherry, Murray, Derek, Olah, Chris, Schuster, Mike, Shlens, Jonathon, Steiner, Benoit, Sutskever, Ilya, Talwar, Kunal, Tucker, Paul, Vanhoucke, Vincent, Vasudevan, Vijay, Viégas, Fernanda, Vinyals, Oriol, Warden, Pete, Wattenberg, Martin, Wicke, Martin, Yu, Yuan, and Zheng, Xiaoqiang · 2015
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Amodei, Dario, Anubhai, Rishita, Battenberg, Eric, Case, Carl, Casper, Jared, Catanzaro, Bryan, Chen, Jingdong, Chrzanowski, Mike, Coates, Adam, Diamos, Greg, et al · 2015
Cited alongside, same era.
Peachpy meets opcodes: direct machine code generation from python
Dukhan, Marat · 2015
Cited alongside, same era.
Persistent rnns: Stashing recurrent weights on-chip
Diamos, Greg, Sengupta, Shubho, Catanzaro, Bryan, Chrzanowski, Mike, Coates, Adam, Elsen, Erich, Engel, Jesse, Hannun, Awni, and Satheesh, Sanjeev · 2016
Later among the works it cites.
Samplernn: An unconditional end-to-end neural audio generation model
Mehri, Soroush, Kumar, Kundan, Gulrajani, Ishaan, Kumar, Rithesh, Jain, Shubham, Sotelo, Jose, Courville, Aaron, and Bengio, Yoshua · 2016
Later among the works it cites.
World: a vocoder-based high-quality speech synthesis system for real-time applications
Morise, Masanori, Yokomori, Fumiya, and Ozawa, Kenji · 2016
Later among the works it cites.
Pixel recurrent neural networks
Oord, Aaron van den, Kalchbrenner, Nal, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grapheme-to-phoneme conversion using long short-term memory recurrent neural networks
Rao, Kanishka, Peng, Fuchun, Sak, Haşim, and Beaufays, Françoise · 2015
Cited alongside, same era.
A note on the evaluation of generative models
Theis, Lucas, Oord, Aäron van den, and Bethge, Matthias · 2015
Cited alongside, same era.
Sequence-to-sequence neural net models for grapheme-to-phoneme conversion
Yao, Kaisheng and Zweig, Geoffrey · 2015
Cited alongside, same era.
Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis
Zen, Heiga and Sak, Haşim · 2015
Cited alongside, same era.
Quasi-recurrent neural networks
Bradbury, James, Merity, Stephen, Xiong, Caiming, and Socher, Richard · 2016
Cited alongside, same era.
Paine, Tom Le, Khorrami, Pooya, Chang, Shiyu, Zhang, Yang, Ramachandran, Prajit, Hasegawa-Johnson, Mark A, and Huang, Thomas S · 2016
Later among the works it cites.
Multi-output rnn-lstm for multiple speaker speech synthesis with α \alpha -interpolation model
Pascual, Santiago and Bonafonte, Antonio · 2016
Later among the works it cites.
A template-based approach for speech synthesis intonation generation using lstms
Ronanki, Srikanth, Henter, Gustav Eje, Wu, Zhizheng, and King, Simon · 2016
Later among the works it cites.
Wavenet: A generative model for raw audio
van den Oord, Aäron, Dieleman, Sander, Zen, Heiga, Simonyan, Karen, Vinyals, Oriol, Graves, Alex, Kalchbrenner, Nal, Senior, Andrew, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Char2wav: End-to-end speech synthesis
Sotelo, Jose, Mehri, Soroush, Kumar, Kundan, Santos, Joao Felipe, Kastner, Kyle, Courville, Aaron, and Bengio, Yoshua · 2017
Closest in time.