Fetching the paper…
Reading the bibliography…
Neural waveform models such as WaveNet have demonstrated better performance than conventional vocoders for statistical parametric speech synthesis.
J. R. Carson and T. C. Fry, “Variable frequency electric circuit theory with application to the theory of frequency-modulation,” Bell System Technical Journal , vol. 16, no. 4, pp. 513–540, 1937
1937
Earlier work this paper cites.
D. B. Fry, A. S. Abramson, P. D. Eimas, and A. M. Liberman, “The identification and discrimination of synthetic vowels,” Language and speech , vol. 5, no. 4, pp. 171–189, 1962
1962
Earlier work this paper cites.
A. M. Liberman, F. S. Cooper, D. P. Shankweiler, and M. Studdert-Kennedy, “Perception of the speech code,” Psychological review , vol. 74, no. 6, p. 431, 1967
1967
Earlier work this paper cites.
Y. Fan, Y. Qian, F. Xie, and F. K. Soong, “TTS synthesis with bidirectional LSTM based recurrent neural networks,” in Proc. Interspeech , 2014, pp. 1964–1968
1968
Earlier work this paper cites.
T. Parks and J. McClellan, “Chebyshev approximation for nonrecursive digital filters with linear phase,” IEEE Transactions on Circuit Theory , vol. 19, no. 2, pp. 189–194, 1972
1972
Earlier work this paper cites.
J. Makhoul, R. Viswanathan, R. Schwartz, and A. Huggins, “A mixed-source model for speech compression and synthesis,” The Journal of the Acoustical Society of America , vol. 64, no. 6, pp. 1577–1581, 1978
1978
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal estimation from modified short-time Fourier transform,” IEEE Trans. ASSP , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
D. Griffin and J. Lim, “A new model-based speech analysis/synthesis system,” in Proc. ICASSP , vol. 10. IEEE, 1985, pp. 513–516
1985
Earlier work this paper cites.
D. W. Griffin and J. S. Lim, “Multiband excitation vocoder,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 36, no. 8, pp. 1223–1235, Aug 1988
1988
Earlier work this paper cites.
R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” Neural computation , vol. 1, no. 2, pp. 270–280, 1989
1989
Earlier work this paper cites.
W. Strange, “Evolving theories of vowel perception,” The Journal of the Acoustical Society of America , vol. 85, no. 5, pp. 2081–2087, 1989
1989
Earlier work this paper cites.
A. J. Abrantes, J. S. Marques, and I. M. Trancoso, “Hybrid sinusoidal modeling of speech without voicing decision,” in Proc. Eurospeech , 1991, pp. 231–234
1991
Earlier work this paper cites.
J. Laroche, Y. Stylianou, and E. Moulines, “HNS: Speech modification based on a harmonic+ noise model,” in Proc. ICASSP , vol. 2. IEEE, 1993, pp. 550–553
1993
Earlier work this paper cites.
K. Tokuda, T. Kobayashi, T. Masuko, and S. Imai, “Mel-generalized cepstral analysis a unified approach,” in Proc. ICSLP , 1994, pp. 1043–1046
1994
Earlier work this paper cites.
A. V. McCree and T. P. Barnwell, “A mixed excitation LPC vocoder model for low bit rate speech coding,” IEEE Transactions on Speech and audio Processing , vol. 3, no. 4, pp. 242–250, 1995
1995
Earlier work this paper cites.
Y. Stylianou, “Harmonic plus noise models for speech, combined with statistical methods, for speech and speaker modification,” Ph.D. dissertation, Ecole Nationale Superieure des Telecommunications, 1996
1996
Earlier work this paper cites.
J. M. Hillenbrand and T. M. Nearey, “Identification of resynthesized/hvd/utterances: Effects of formant contour,” The Journal of the Acoustical Society of America , vol. 105, no. 6, pp. 3509–3523, 1999
1999
Earlier work this paper cites.
D. D. Lee and H. S. Seung, “Algorithms for non-negative matrix factorization,” in Proc. NIPS , 2001, pp. 556–562
2001
Earlier work this paper cites.
H. Kawai, T. Toda, J. Ni, M. Tsuzaki, and K. Tokuda, “XIMERA: A new TTS from ATR based on corpus-based technologies,” in Proc. SSW5 , 2004, pp. 179–184
2004
Earlier work this paper cites.
A. Graves, “Supervised Sequence Labelling with Recurrent Neural Networks,” Ph.D. dissertation, Technische Universität München, 2008
2008
Earlier work this paper cites.
P. Taylor, Text-to-Speech Synthesis . Cambridge University Press, 2009
2009
Cited alongside, same era.
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,” Speech Communication , vol. 51, pp. 1039–1064, 2009
2009
Cited alongside, same era.
N. Bell and J. Hoberock, “Thrust: A productivity-oriented library for CUDA,” in GPU computing gems Jade edition . Elsevier, 2011, pp. 359–371
2011
Cited alongside, same era.
I. Saratxaga, I. Hernaez, M. Pucher, E. Navas, and I. Sainz, “Perceptual importance of the phase related information in speech,” in Proc. Interspeech , 2012
2012
Cited alongside, same era.
K. Tokuda, Y. Nankaku, T. Toda, H. Zen, J. Yamagishi, and K. Oura, “Speech synthesis based on hidden Markov models,” Proceedings of the IEEE , vol. 101, no. 5, pp. 1234–1252, 2013
2013
Cited alongside, same era.
B. Chen, T. Bian, and K. Yu, “Discrete duration model for speech synthesis,” in Proc. Interspeech , 2017, pp. 789–793
2017
Later among the works it cites.
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio, “SampleRNN: An unconditional end-to-end neural audio generation model,” in Proc. ICLR , 2017
2017
Later among the works it cites.
S. Takaki, H. Kameoka, and J. Yamagishi, “Direct modeling of frequency spectra and waveform generation based on phase recovery for DNN-based speech synthesis,” in Proc. Interspeech , 2017, pp. 1128–1132
2017
Later among the works it cites.
X. Wang, S. Takaki, and J. Yamagishi, “Autoregressive neural F0 model for statistical parametric speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 8, pp. 1406–1419, 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in Proc. ICASSP , 2013, pp. 7962–7966
2013
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. ICLR , 2014, p. unknown
2014
Cited alongside, same era.
K. Yao and G. Zweig, “Sequence-to-sequence neural net models for grapheme-to-phoneme conversion,” in Proc. Interspeech , 2015, pp. 3330–3334
2015
Cited alongside, same era.
D. Rezende and S. Mohamed, “Variational inference with normalizing flows,” in Proc. ICML , 2015, pp. 1530–1538
2015
Cited alongside, same era.
F. Weninger, J. Bergmann, and B. Schuller, “Introducing CURRENT: The Munich open-source CUDA recurrent neural network toolkit,” The Journal of Machine Learning Research , vol. 16, no. 1, pp. 547–551, 2015
2015
Cited alongside, same era.
M. S. Ribeiro, O. Watts, and J. Yamagishi12, “Parallel and cascaded deep neural networks for text-to-speech synthesis,” in Proc. SSW , 2016, pp. 100–105
2016
Cited alongside, same era.
G. E. Henter, S. Ronanki, O. Watts, M. Wester, Z. Wu, and S. King, “Robust TTS duration modelling using DNNs,” in Proc. ICASSP , 2016, pp. 5130–5134
2016
Cited alongside, same era.
X. Wang, J. Lorenzo-Trueba, S. Takaki, L. Juvela, and J. Yamagishi, “A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,” in Proc. ICASSP , 2018, pp. 4804–4808
2018
Later among the works it cites.
Y. Ai, H.-C. Wu, and Z.-H. Ling, “SampleRNN-based neural vocoder for statistical parametric speech synthesis,” in Proc. ICASSP . IEEE, 2018, pp. 5659–5663
2018
Later among the works it cites.
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, “Parallel WaveNet: Fast high-fidelity speech synthesis,” in Proc. ICML , 2018, pp. 3918–3926
2018
Later among the works it cites.
2018
Later among the works it cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in Proc. ICML , vol. 80, 10–15 Jul 2018, pp. 2410–2419
2018
Later among the works it cites.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “FFTNet: A real-time speaker-dependent neural vocoder,” in Proc. ICASSP . IEEE, 2018, pp. 2251–2255
2018
Later among the works it cites.
T. Okamoto, K. Tachibana, T. Toda, Y. Shiga, and H. Kawai, “An investigation of subband WaveNet vocoder covering entire audible frequency range with limited acoustic features,” in Proc. ICASSP . IEEE, 2018, pp. 5654–5658
2018
Later among the works it cites.
X. Wang, S. Takaki, and J. Yamagishi, “Investigation of WaveNet for text-to-speech synthesis,” SIG Technical Reports, Tech. Rep. 6, feb 2018
2018
Later among the works it cites.
H.-T. Luong, X. Wang, J. Yamagishi, and N. Nishizawa, “Investigating accuracy of pitch-accent annotations in neural-network-based speech synthesis and denoising effects,” in Proc. Interspeech , 2018, pp. 37–41
2018
Later among the works it cites.
L. Juvela, B. Bollepalli, X. Wang, H. Kameoka, M. Airaksinen, J. Yamagishi, and P. Alku, “Speech waveform synthesis from MFCC sequences with generative adversarial networks,” in Proc. ICASSP . IEEE, 2018, pp. 5679–5683
2018
Later among the works it cites.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A flow-based generative network for speech synthesis,” in Proc. ICASSP , 2019, pp. 3617–3621
2019
Closest in time.
W. Ping, K. Peng, and J. Chen, “ClariNet: Parallel wave generation in end-to-end text-to-speech,” in Proc. ICLP , 2019
2019
Closest in time.
X. Wang, S. Takaki, and J. Yamagishi, “Neural source-filter-based waveform model for statistical parametric speech synthesis,” in Proc. ICASSP , 2019, pp. 5916–5920
2019
Closest in time.
J.-M. Valin and J. Skoglund, “LPCNet: Improving neural speech synthesis through linear prediction,” in Proc. ICASSP , 2019, pp. 5891–5895
2019
Closest in time.
T. Okamoto, T. Toda, Y. Shiga, and H. Kawai, “Investigations of real-time neural vocoders with fundamental frequency and Mel-cepstra,” in Proc. ASJ spring meeting , 2019, p. (in Japanese)
2019
Closest in time.
Y. Yasuda, X. Wang, S. Takaki, and J. Yamagishi, “Investigation of enhanced Tacotron text-to-speech synthesis systems with self-attention for pitch accent language,” in Proc. ICASSP . IEEE, 2019, pp. 6905–6909
2019
Closest in time.