Fetching the paper…
Reading the bibliography…
We present a wav-to-wav generative model for the task of singing voice conversion from any identity.
B. C. Moore, B. R. Glasberg, and T. Baer, “A model for the prediction of thresholds, loudness, and partial loudness,”
1997
Earlier work this paper cites.
K. Saino, H. Zen, Y. Nankaku, A. Lee, and K. Tokuda, “An hmm-based singing voice synthesis system,” in
2006
Earlier work this paper cites.
T. Nakatani
2008
Earlier work this paper cites.
W. Chu and A. Alwan, “Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,” in
2009
Earlier work this paper cites.
F. Villavicencio and J. Bonada, “Applying voice conversion to concatenative singing-voice synthesis.” in
2010
Earlier work this paper cites.
K. Nakamura
2014
Earlier work this paper cites.
——, “Statistical singing voice conversion with direct waveform modification based on the spectrum differential,” in
2014
Earlier work this paper cites.
K. Kobayashi
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Bonada, M. Umbert, and M. Blaauw, “Expressive singing synthesis based on unit selection for the singing synthesis challenge 2016,” in
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in
2016
Earlier work this paper cites.
R. Sennrich
2016
Earlier work this paper cites.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in
2016
Earlier work this paper cites.
S. Mehri
2017
Earlier work this paper cites.
M. Blaauw and J. Bonada, “A Neural Parametric Singing Synthesizer,” in
2017
Earlier work this paper cites.
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, “Least squares generative adversarial networks,” in
2017
Cited alongside, same era.
K. Ito, “The lj speech dataset,”
2017
Cited alongside, same era.
C. Veaux
2017
Cited alongside, same era.
M. McAuliffe
2017
Cited alongside, same era.
N. Kalchbrenner
2018
Cited alongside, same era.
A. van den Oord
2018
Cited alongside, same era.
S. Ö. Arık, H. Jun, and G. Diamos, “Fast spectrogram inversion using multi-head convolutional neural networks,” in
2018
Cited alongside, same era.
S. Kim, S.-G. Lee, J. Song, J. Kim, and S. Yoon, “FloWaveNet : A generative flow for raw audio,” in
2019
Later among the works it cites.
R. Prenger, R. Valle, and B. Catanzaro, “Waveglow: A flow-based generative network for speech synthesis,” in
2019
Later among the works it cites.
K. Kumar
2019
Later among the works it cites.
X. Wang, S. Takaki, and J. Yamagishi, “Neural source-filter-based waveform model for statistical parametric speech synthesis,” in
2019
Later among the works it cites.
M. Blaauw, J. Bonada, and R. Daido, “Data efficient voice cloning for neural singing synthesis,” in
2019
Later among the works it cites.
P. Chandna, M. Blaauw, J. Bonada, and E. Gómez, “WGANSing: A multi-voice singing voice synthesizer based on the wasserstein-gan,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. W. Kim, J. Salamon, P. Li, and J. P. Bello, “CREPE: A convolutional representation for pitch estimation,” in
2018
Cited alongside, same era.
A. v. d. Oord
2018
Cited alongside, same era.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in
2018
Cited alongside, same era.
D. Rethage, J. Pons, and X. Serra, “A wavenet for speech denoising,” in
2018
Cited alongside, same era.
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani, “VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop,” in
2018
Cited alongside, same era.
2019
Later among the works it cites.
E. Nachmani and L. Wolf, “Unsupervised Singing Voice Conversion,”
2019
Later among the works it cites.
X. Chen, W. Chu, J. Guo, and N. Xu, “Singing voice conversion with non-parallel data,”
2019
Later among the works it cites.
C. Hawthorne
2019
Later among the works it cites.
L. Crew,
2019
Later among the works it cites.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in
2020
Closest in time.
J. Engel, L. Hantrakul, C. Gu, and A. Roberts, “DDSP: Differentiable Digital Signal Processing,”
2020
Closest in time.
R. Valle, J. Li, R. Prenger, and B. Catanzaro, “Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,”
2020
Closest in time.
M. Blaauw and J. Bonada, “Sequence-to-Sequence Singing Synthesis Using the Feed-Forward Transformer,” in
2020
Closest in time.
Y.-J. Luo
2020
Closest in time.
G. Xiaoxue
2020
Closest in time.