Fetching the paper…
Reading the bibliography…
This paper explores the potential universality of neural vocoders.
“Signal estimation from modified short-time fourier transform,”
D. Griffin and J. Lim, · 1984
Earlier work this paper cites.
“A four-parameter model of glottal flow,”
G. Fant and Q. Liljencrants, J.and Lin, · 1985
Earlier work this paper cites.
“Speech concatenation and synthesis using an overlap-add sinusoidal model,”
M. W. Macon and M. A. Clements, · 1996
Earlier work this paper cites.
“Aperiodicity extraction and control using mixed mode excitation and group delay manipulation for a high quality speech analysis, modification and synthesis system straight,”
H. Kawahara, J. Estill, and O. Fujimura, · 2001
Earlier work this paper cites.
“Bs. 1534-1. method for the subjective assessment of intermediate sound quality (mushra),”
I. Recommendation, · 2001
Earlier work this paper cites.
“The hmm-based speech synthesis system (hts) version 2.0.,”
H. Zen, T. Nose, J. Yamagishi, et al., · 2007
Earlier work this paper cites.
“Comparison of multiple voice source parameters in different phonation types,”
M. Airas and P. Alku, · 2007
Earlier work this paper cites.
“Voice cloning toolkit for festival and hts,” 2010
J. Yamagishi and K. Edwards, · 2010
Earlier work this paper cites.
“The deterministic plus stochastic model of the residual signal and its applications,”
T. Drugman and T. Dutoit, · 2012
Earlier work this paper cites.
“When voices get emotional: a corpus of nonverbal vocalizations for research on emotion processing,”
C. F. Lima, S. L. Castro, and S. K. Scott, · 2013
Earlier work this paper cites.
“Investigating source and filter contributions, and their interaction, to statistical parametric speech synthesis,”
T. Merritt, T. Raitio, and S. King, · 2014
Earlier work this paper cites.
“Attributing modelling errors in HMM synthesis by stepping gradually from natural to modelled speech,”
T. Merritt, J. Latorre, and S. King, · 2015
Earlier work this paper cites.
“librosa: Audio and music signal analysis in python,”
B. McFee, C. Raffel, D. Liang, et al., · 2015
Cited alongside, same era.
“Wavenet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, et al., · 2016
Cited alongside, same era.
“WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
M. Morise, F. Yokomori, and K. Ozawa, · 2016
Cited alongside, same era.
“Blizzard Challenge 2016,” 2016,
S. King and V. Karaiskos, · 2016
Cited alongside, same era.
“Reverberant speech database for training speech dereverberation algorithms and tts models,” 2016
C. Valentini-Botinhao et al., · 2016
Cited alongside, same era.
“Collecting resources in sub-saharan african languages for automatic speech recognition: a case study of wolof,”
“Efficient neural audio synthesis,”
N. Kalchbrenner, E. Elsen, K. Simonyan, et al., · 2018
Closest in time.
“Clarinet: Parallel wave generation in end-to-end text-to-speech,”
W. Ping, K. Peng, and J. Chen, · 2018
Closest in time.
“Comprehensive evaluation of statistical speech waveform synthesis,”
T. Merritt, B. Putrycz, A. Nadolski, et al., · 2018
Closest in time.
“Waveglow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2018
Closest in time.
“Fftnet: A real-time speaker-dependent neural vocoder,”
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Gauthier, L. Besacier, S. Voisin, et al., · 2016
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, et al., · 2017
Cited alongside, same era.
“An investigation of multi-speaker training for wavenet vocoder,”
T. Hayashi, A. Tamamori, K. Kobayashi, et al., · 2017
Cited alongside, same era.
“Parallel wavenet: Fast high-fidelity speech synthesis,”
A. van den Oord, Y. Li, I. Babuschkin, et al., · 2017
Cited alongside, same era.
“Noisy speech database for training speech enhancement algorithms and tts models,” 2017
C. Valentini-Botinhao et al., · 2017
Cited alongside, same era.
“Noisy reverberant speech database for training speech enhancement algorithms and tts models,” 2017
C. Valentini-Botinhao et al., · 2017
Cited alongside, same era.
“The 2016 signal separation evaluation campaign,”
A. Liutkus, F.-R. Stöter, Z. Rafii, et al., · 2017
Cited alongside, same era.
“A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,”
X. Wang, J. Lorenzo-Trueba, S. Takaki, et al., · 2018
Closest in time.
“Rapid style adaptation using residual error embedding for expressive speech synthesis,”
X. Wu, Y. Cao, M. Wang, et al., · 2018
Closest in time.
“A voice conversion framework with tandem feature sparse representation and speaker-adapted wavenet vocoder,”
B. Sisman, M. Zhang, and H. Li, · 2018
Closest in time.
“Wavenet vocoder with limited training data for voice conversion,”
L. Liu, Z. Ling, Y. Jiang, et al., · 2018
Closest in time.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Y. Jia, Y. Zhang, R. Weiss, et al., · 2018
Closest in time.
“Collapsed speech segment detection and suppression for wavenet vocoder,”
Y. Wu, K. Kobayashi, T. Hayashi, et al., · 2018
Closest in time.
“Fast spectrogram inversion using multi-head convolutional neural networks,”
S. Ö. Arık, H. Jun, and G. Diamos, · 2019
Closest in time.