Fetching the paper…
Reading the bibliography…
Neural speech synthesis models have recently demonstrated the ability to synthesize high quality speech for text-to-speech and compression applications.
“Speech analysis and synthesis by linear prediction of the speech wave,”
B. S. Atal and S. L. Hanauer, · 1971
Earlier work this paper cites.
“A linear prediction vocoder simulation based upon the autocorrelation method,”
J. Markel and A. Gray, · 1974
Earlier work this paper cites.
“Linear prediction: A tutorial review,”
J. Makhoul, · 1975
Earlier work this paper cites.
“A new model of LPC excitation for producing natural-sounding speech at low bit rates,”
B. S. Atal and J. Remde, · 1982
Earlier work this paper cites.
“A new model-based speech analysis/synthesis system,”
D. Griffin and J. Lim, · 1985
Earlier work this paper cites.
“Code-excited linear prediction (CELP): High-quality speech at very low bit rates,”
M. Schroeder and B.S. Atal, · 1985
Earlier work this paper cites.
Recommendation G.711: Pulse Code Modulation (PCM) of voice frequencies
ITU-T, · 1988
Earlier work this paper cites.
“A 2.4 kbit/s MELP coder candidate for the new us federal standard,”
A. McCree, K. Truong, E. B. George, T. P. Barnwell, and V. Viswanathan, · 1996
Earlier work this paper cites.
ITU-R, · 2001
Cited alongside, same era.
An introduction to the psychology of hearing
B.C.J. Moore, · 2012
Cited alongside, same era.
“On the properties of neural machine translation: Encoder-decoder approaches,”
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio, · 2014
Cited alongside, same era.
“WaveNet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Cited alongside, same era.
“Speaker-independent raw waveform model for glottal excitation,”
L. Juvela, V. Tsiaras, B. Bollepalli, M. Airaksinen, J. Yamagishi, and P. Alku, · 2016
Cited alongside, same era.
“Char2wav: End-to-end speech synthesis,”
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. Courville, and Y. Bengio, · 2017
Later among the works it cites.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, et al., · 2018
Closest in time.
“WaveNet based low rate speech coding,”
W. B. Kleijn, F. SC Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, · 2018
Closest in time.
“FFTNet: A real-time speaker-dependent neural vocoder,”
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, · 2018
Closest in time.
“Efficient neural audio synthesis,”
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio, · 2016
Cited alongside, same era.
“Deep voice 2: Multi-speaker neural text-to-speech,”
S. Arik, G. Diamos, A. Gibiansky, J. Miller, K. Peng, W. Ping, J. Raiman, and Y. Zhou, · 2017
Cited alongside, same era.
J.-M. Valin, · 2018
Closest in time.
“On the convergence of adam and beyond,”
S. J. Reddi, S. Kale, and S. Kumar, · 2018
Closest in time.