Fetching the paper…
Reading the bibliography…
WaveCycleGAN has recently been proposed to bridge the gap between natural and synthesized speech waveforms in statistical parametric speech synthesis and provides fast inference with a moving average model rather than an autoregressive model and high-quality speech synthesis with the adversarial training.
1904
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal estimation from modified short-time Fourier transform,”
1984
Earlier work this paper cites.
C. E. Shannon, “Communication in the presence of noise,”
1998
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. De Cheveigne, “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based F0 extraction: Possible role of a repetitive structure in sounds,”
1999
Earlier work this paper cites.
H. Kawahara, J. Estill, and O. Fujimura, “Aperiodicity extraction and control using mixed mode excitation and group delay manipulation for a high quality speech analysis, modification and synthesis system STRAIGHT,” in
2001
Earlier work this paper cites.
L. Atlas and S. A. Shamma, “Joint acoustic and modulation frequency,”
2003
Earlier work this paper cites.
J. Kominek and A. W. Black, “The CMU Arctic speech databases,” in
2004
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
2007
Earlier work this paper cites.
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,”
2009
Earlier work this paper cites.
H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in
2013
Earlier work this paper cites.
S. Takamichi, T. Toda, G. Neubig, S. Sakti, and S. Nakamura, “A postfilter to modify the modulation spectrum in HMM-based speech synthesis,” in
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in
2014
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in
2014
Earlier work this paper cites.
J. Mairal, P. Koniusz, Z. Harchaoui, and C. Schmid, “Convolutional kernel networks,” in
2014
Earlier work this paper cites.
2014
Cited alongside, same era.
S. Takamichi, K. Kobayashi, K. Tanaka, T. Toda, and S. Nakamura, “The NAIST text-to-speech system for the blizzard challenge 2015,” in
2015
Cited alongside, same era.
2016
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: a vocoder-based high-quality speech synthesis system for real-time applications,”
2016
Cited alongside, same era.
D. He, Y. Xia, T. Qin, L. Wang, N. Yu, T. Liu, and W.-Y. Ma, “Dual learning for machine translation,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
T. Zhou, P. Krahenbuhl, M. Aubry, Q. Huang, and A. A. Efros, “Learning dense correspondence via 3D-guided cycle consistency,” in
2016
Cited alongside, same era.
Y. Taigman, A. Polyak, and L. Wolf, “Unsupervised cross-domain image generation,”
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang
2016
Cited alongside, same era.
K. Ito, “The LJ speech dataset,”
2017
Cited alongside, same era.
2017
Cited alongside, same era.
T. Kaneko, H. Kameoka, N. Hojo, Y. Ijima, K. Hiramatsu, and K. Kashino, “Generative adversarial network-based postfilter for statistical parametric speech synthesis,” in
2017
Cited alongside, same era.
2018
Later among the works it cites.
Y. Saito, S. Takamichi, H. Saruwatari, Y. Saito, S. Takamichi, and H. Saruwatari, “Statistical parametric speech synthesis incorporating generative adversarial networks,”
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Sun, J. Pang, J. Shi, S. Yi, and W. Ouyang, “FishNet: A versatile backbone for image, region, and pixel level prediction,” in
2018
Later among the works it cites.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with SincNet,”
2018
Later among the works it cites.
Y. Gong and C. Poellabauer, “Impact of aliasing on deep CNN-based end-to-end acoustic models,”
2018
Later among the works it cites.
2018
Later among the works it cites.
P. L. Tobing, Y.-C. Wu, T. Hayashi, K. Kobayashi, and T. Toda, “Voice conversion with cyclic recurrent neural network and fine-tuned WaveNet vocoder,” in
2019
Closest in time.
S. Ma, D. Mcduff, and Y. Song, “A generative adversarial network for style modeling in a text-to-speech system,” in
2019
Closest in time.