Fetching the paper…
Reading the bibliography…
We present a modification to the spectrum differential based direct waveform modification for voice conversion (DIFFVC) so that it can be directly applied as a waveform generation module to voice conversion models.
B. S. Atal and S. L. Hanauer, “Speech analysis and synthesis by linear prediction of the speech wave,”
1971
Earlier work this paper cites.
J. Makhoul and M. Berouti, “High-frequency regeneration in speech coding systems,” in
1979
Earlier work this paper cites.
Y.-C. Wu, K. Kobayashi, T. Hayashi, P. L. Tobing, and T. Toda, “Collapsed speech segment detection and suppression for wavenet vocoder,” in
1992
Earlier work this paper cites.
W. Verhelst and M. Roelands, “An overlap-add technique based on waveform similarity (wsola) for high quality time-scale modification of speech,” in
1993
Earlier work this paper cites.
K. Tokuda, T. Kobayashi, T. Masuko, and S. Imai, “Mel-generalized cepstral analysis - a unified approach to speech spectral estimation,” in
1994
Earlier work this paper cites.
K. Chen, B. Chen, J. Lai, and K. Yu, “High-quality voice conversion using spectrogram-based wavenet vocoder,” in
1997
Earlier work this paper cites.
Y. Stylianou, O. Cappe, and E. Moulines, “Continuous probabilistic transform for voice conversion,”
1998
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. de Cheveigné, “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds,”
1999
Earlier work this paper cites.
S. Chennoukh, A. Gerrits, G. Miet, and R. Sluijter, “Speech enhancement via frequency bandwidth extension using line spectral frequencies,” in
2001
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
2007
Earlier work this paper cites.
S. Desai, A. W. Black, B. Yegnanarayana, and K. Prahallad, “Spectral mapping using artificial neural networks for voice conversion,”
2010
Earlier work this paper cites.
R. Takashima, T. Takiguchi, and Y. Ariki, “Exemplar-based voice conversion in noisy environment,” in
2012
Earlier work this paper cites.
H. Silén, E. Hel, J. Nurminen, and M. Gabbouj, “Ways to implement global variance in statistical speech synthesis,” in
2012
Earlier work this paper cites.
L. H. Chen, Z. H. Ling, L. J. Liu, and L. R. Dai, “Voice conversion using deep neural networks with layer-wise generative training,”
2014
Cited alongside, same era.
Z. Wu, T. Virtanen, E. S. Chng, and H. Li, “Exemplar-based sparse representation with residual compensation for voice conversion,”
2014
Cited alongside, same era.
Y.-C. Wu, H.-T. Hwang, C.-C. Hsu, Y. Tsao, and H.-M. Wang, “Locally linear embedding for exemplar-based spectral conversion,” in
2016
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications,”
2016
Cited alongside, same era.
2016
K. Kobayashi, T. Hayashi, A. Tamamori, and T. Toda, “Statistical voice conversion with wavenet-based waveform generation,” in
2017
Later among the works it cites.
2018
Later among the works it cites.
Z. Jin, A. Finkelstein, G. J. Mysore, and J. Lu, “Fftnet: A real-time speaker-dependent neural vocoder,” in
2018
Later among the works it cites.
B. Sisman, M. Zhang, S. Sakti, H. Li, and S. Nakamura, “Adaptive wavenet vocoder for residual compensation in gan-based voice conversion,” in
2018
Later among the works it cites.
K. Kobayashi and T. Toda, “sprocket: Open-source voice conversion software,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
K. Kobayashi, S. Takamichi, S. Nakamura, and T. Toda, “The nu-naist voice conversion system for the voice conversion challenge 2016,” in
2016
Cited alongside, same era.
T. Toda, L.-H. Chen, D. Saito, F. Villavicencio, M. Wester, Z. Wu, and J. Yamagishi, “The voice conversion challenge 2016,” in
2016
Cited alongside, same era.
K. Kobayashi, T. Toda, and S. Nakamura, “Implementation of f0 transformation for statistical singing voice conversion based on direct waveform modification,”
2016
Cited alongside, same era.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in
2016
Cited alongside, same era.
K. Kobayashi, T. Toda, and S. Nakamura, “F0 transformation techniques for statistical voice conversion with direct waveform modification with spectral differential,”
2016
Cited alongside, same era.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent wavenet vocoder,” in
2017
Cited alongside, same era.
T. Hayashi, A. Tamamori, K. Kobayashi, K. Takeda, and T. Toda, “An investigation of multi-speaker training for wavenet vocoder,” in
2017
Cited alongside, same era.
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, “The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,” in
2018
Later among the works it cites.
J.-M. Valin and J. Skoglund, “Lpcnet: Improving neural speech synthesis through linear prediction,” in
2019
Closest in time.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in
2019
Closest in time.
X. Wang, S. Takaki, and J. Yamagishi, “Neural source-filter-based waveform model for statistical parametric speech synthesis,” in
2019
Closest in time.
2019
Closest in time.
J.-X. Zhang, Z.-H. Ling, L.-J. Liu, Y. Jiang, and L.-R. Dai, “Sequence-to-sequence acoustic modeling for voice conversion,”
2019
Closest in time.
W.-C. Huang, Y.-C. Wu, C.-C. Lo, P. Lumban Tobing, T. Hayashi, K. Kobayashi, T. Toda, Y. Tsao, and H.-M. Wang, “Investigation of F0 conditioning and Fully Convolutional Networks in Variational Autoencoder based Voice Conversion,”
2019
Closest in time.