Fetching the paper…
Reading the bibliography…
This paper introduces WaveNet, a deep neural network for generating raw audio waveforms.
Remaking speech
Dudley, Homer · 1939
Earlier work this paper cites.
The Vowel: Its Nature and Structure
Chiba, Tsutomu and Kajiyama, Masato · 1942
Earlier work this paper cites.
Acoustic Theory of Speech Production
Fant, Gunnar · 1970
Earlier work this paper cites.
A statistical method for estimation of speech spectral density and formant frequencies
Itakura, Fumitada and Saito, Shuzo · 1970
Earlier work this paper cites.
Line spectrum representation of linear predictor coefficients of speech signals
Itakura, Fumitada · 1975
Earlier work this paper cites.
Linear predictive hidden Markov models and the speech signal
Poritz, Alan B · 1982
Earlier work this paper cites.
Mixture autoregressive hidden Markov models for speech signals
Juang, Biing-Hwang and Rabiner, Lawrence · 1985
Earlier work this paper cites.
Unbiased estimation of log spectrum
Imai, Satoshi and Furuichi, Chieko · 1988
Earlier work this paper cites.
Recommendation G. 711
ITU-T · 1988
Earlier work this paper cites.
An implementation of the “algorithme à trous” to compute the wavelet transform
Dutilleux, Pierre · 1989
Earlier work this paper cites.
A real-time algorithm for signal analysis with the help of the wavelet transform
Holschneider, Matthias, Kronland-Martinet, Richard, Morlet, Jean, and Tchamitchian, Philippe · 1989
Earlier work this paper cites.
Pitch synchronous waveform processing techniques for text-to-speech synthesis using diphones
Moulines, Eric and Charpentier, Francis · 1990
Earlier work this paper cites.
ATR ν \nu -talk speech synthesis system
Sagisaka, Yoshinori, Kaiki, Nobuyoshi, Iwahashi, Naoto, and Mimura, Katsuhiko · 1992
Earlier work this paper cites.
DARPA TIMIT acoustic-phonetic continuous speech corpus CD-ROM. NIST speech disc 1-1.1
Garofolo, John S., Lamel, Lori F., Fisher, William M., Fiscus, Jonathon G., and Pallett, David S · 1993
Earlier work this paper cites.
Fundamentals of Speech Recognition
Rabiner, Lawrence and Juang, Biing-Hwang · 1993
Earlier work this paper cites.
Speech synthesis using artificial neural networks trained on cepstral coefficients
Tuerk, Christine and Robinson, Tony · 1993
Earlier work this paper cites.
Mixture density networks
Bishop, Christopher M · 1994
Earlier work this paper cites.
Unit selection in a concatenative speech synthesis system using a large speech database
Hunt, Andrew J. and Black, Alan W · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Text-to-speech conversion with neural networks: A recurrent TDNN approach
Karaali, Orhan, Corrigan, Gerald, Gerson, Ira, and Massey, Noel · 1997
Earlier work this paper cites.
Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based f 0 f_{0} extraction: possible role of a repetitive structure in sounds
Kawahara, Hideki, Masuda-Katsuse, Ikuyo, and de Cheveigné, Alain · 1999
Cited alongside, same era.
Aperiodicity extraction and control using mixed mode excitation and group delay manipulation for a high quality speech analysis, modification and synthesis system STRAIGHT
Kawahara, Hideki, Estill, Jo, and Fujimura, Osamu · 2001
Cited alongside, same era.
Nonlinear filter design: methodologies and challenges
Peltonen, Sari, Gabbouj, Moncef, and Astola, Jaakko · 2001
Cited alongside, same era.
Simultaneous modeling of phonetic and prosodic parameters, and characteristic conversion for HMM-based text-to-speech systems
Yoshimura, Takayoshi · 2002
Cited alongside, same era.
An example of context-dependent label format for HMM-based speech synthesis in English, 2006
Muthukumar, P. and Black, Alan W · 2014
Later among the works it cites.
Integration of spectral feature extraction and modeling for HMM-based speech synthesis
Nakamura, Kazuhiro, Hashimoto, Kei, Nankaku, Yoshihiko, and Tokuda, Keiichi · 2014
Later among the works it cites.
Acoustic modeling with deep neural networks using raw time signal for LVCSR
Tüske, Zoltán, Golik, Pavel, Schlüter, Ralf, and Ney, Hermann · 2014
Later among the works it cites.
Vocaine the vocoder and applications is speech synthesis
Agiomyrgiannakis, Yannis · 2015
Later among the works it cites.
Semantic image segmentation with deep convolutional nets and fully connected CRFs
Chen, Liang-Chieh, Papandreou, George, Kokkinos, Iasonas, Murphy, Kevin, and Yuille, Alan L · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zen, Heiga · 2006
Cited alongside, same era.
A speech parameter generation algorithm considering global variance for HMM-based speech synthesis
Toda, Tomoki and Tokuda, Keiichi · 2007
Cited alongside, same era.
Reformulating the HMM as a trajectory model by imposing explicit relationships between static and dynamic features
Zen, Heiga, Tokuda, Keiichi, and Kitamura, Tadashi · 2007
Cited alongside, same era.
Statistical approach to vocal tract transfer function estimation based on factor analyzed trajectory hmm
Toda, Tomoki and Tokuda, Keiichi · 2008
Cited alongside, same era.
Minimum generation error training with direct log spectral distortion on LSPs for HMM-based speech synthesis
Wu, Yi-Jian and Tokuda, Keiichi · 2008
Cited alongside, same era.
Input-agreement: a new mechanism for collecting data using human computation games
Law, Edith and Von Ahn, Luis · 2009
Cited alongside, same era.
Statistical parametric speech synthesis
Zen, Heiga, Tokuda, Keiichi, and Black, Alan W · 2009
Cited alongside, same era.
Speech analysis with multi-kernel linear prediction
Kameoka, Hirokazu, Ohishi, Yasunori, Mochihashi, Daichi, and Le Roux, Jonathan · 2010
Cited alongside, same era.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2015
Later among the works it cites.
Speech acoustic modeling from raw multichannel waveforms
Hoshen, Yedid, Weiss, Ron J., and Wilson, Kevin W · 2015
Later among the works it cites.
Learning the speech front-end with raw waveform CLDNNs
Sainath, Tara N., Weiss, Ron J., Senior, Andrew, Wilson, Kevin W., and Vinyals, Oriol · 2015
Later among the works it cites.
Generative image modeling using spatial LSTMs
Theis, Lucas and Bethge, Matthias · 2015
Later among the works it cites.
Directly modeling speech waveforms by neural networks for statistical parametric speech synthesis
Tokuda, Keiichi and Zen, Heiga · 2015
Later among the works it cites.
Modelling acoustic feature dependencies with artificial neural networks: Trajectory-RNADE
Uria, Benigno, Murray, Iain, Renals, Steve, Valentini-Botinhao, Cassia, and Bridle, John · 2015
Later among the works it cites.
Recent advances in Google real-time HMM-driven unit selection synthesizer
Gonzalvo, Xavi, Tazari, Siamak, Chan, Chun-an, Becker, Markus, Gutkin, Alexander, and Silen, Hanna · 2016
Closest in time.
Exploring the limits of language modeling
Józefowicz, Rafal, Vinyals, Oriol, Schuster, Mike, Shazeer, Noam, and Wu, Yonghui · 2016
Closest in time.
WORLD: A vocoder-based high-quality speech synthesis system for real-time applications
Morise, Masanori, Yokomori, Fumiya, and Ozawa, Kenji · 2016
Closest in time.
A deep auto-encoder based low-dimensional feature extraction from FFT spectral envelopes for statistical parametric speech synthesis
Takaki, Shinji and Yamagishi, Junichi · 2016
Closest in time.
Postfilters to modify the modulation spectrum for statistical parametric speech synthesis
Takamichi, Shinnosuke, Toda, Tomoki, Black, Alan W., Neubig, Graham, Sakriani, Sakti, and Nakamura, Satoshi · 2016
Closest in time.
Directly modeling voiced and unvoiced components in speech waveforms by neural networks
Tokuda, Keiichi and Zen, Heiga · 2016
Closest in time.
Multi-scale context aggregation by dilated convolutions
Yu, Fisher and Koltun, Vladlen · 2016
Closest in time.
Zen, Heiga, Agiomyrgiannakis, Yannis, Egberts, Niels, Henderson, Fergus, and Szczepaniak, Przemysław · 2016
Closest in time.