Fetching the paper…
Reading the bibliography…
A vocoder is a conditional audio generation model that converts acoustic features such as mel-spectrograms into waveforms.
Y. Ren, X. Tan, T. Qin, J. Luan, Z. Zhao, and T.-Y. Liu, “DeepSinger: Singing voice synthesis with data mined from the web,” in ACM Int. Conf. Knowledge Discovery & Data Mining , 2020, pp. 1979–1989
1989
Earlier work this paper cites.
J. Sundberg, The Science of the Singing Voice . Northern Illinois University Press, 1989
1989
Earlier work this paper cites.
X. Serra and J. Smith, “Spectral modeling synthesis: A sound analysis/synthesis system based on a deterministic plus stochastic decomposition,” Computer Music Journal , vol. 14, no. 4, pp. 12–24, 1990
1990
Earlier work this paper cites.
P. R. Cook, “Singing voice synthesis: History, current work, and future directions,” Computer Music Journal , vol. 20, no. 3, pp. 38–46, 1996
1996
Earlier work this paper cites.
J. Lane, D. Hoory, E. Martinez, and P. Wang, “Modeling analog synthesis with DSPs,” Computer Music Journal , vol. 21, no. 4, pp. 32–41, 1997
1997
Earlier work this paper cites.
A. Huovilainen and V. Välimäki, “New approaches to digital subtractive synthesis,” in Inr. Computer Music Conf. , 2005
2005
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Morise, F. Yokomori, and K. Ozawa, “WORLD: A vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Transactions on Information and Systems , vol. E99.D, no. 7, pp. 1877–1884, 2016
2016
Earlier work this paper cites.
K. Ito, “The LJ speech dataset,” 2017
2017
Earlier work this paper cites.
N. Kalchbrenner et al. , “Efficient neural audio synthesis,” in Int. Conf. Machine Learning , 2018
2018
Earlier work this paper cites.
J. W. Kim, J. Salamon, P. Li, and J. P. Bello, “CREPE: A convolutional representation for pitch estimation,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2018, pp. 161–165
2018
Earlier work this paper cites.
Y. Jadoul, B. Thompson, and B. de Boer, “Introducing Parselmouth: A Python interface to Praat,” Journal of Phonetics , vol. 71, pp. 1–15, 2018
2018
Earlier work this paper cites.
J. Lee, H.-S. Choi, C.-B. Jeon, J. Koo, and K. Lee, “Adversarially trained end-to-end Korean singing voice synthesis system,” in INTERSPEECH , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
K. Kilgour et al. , “Fréchet Audio Distance: A metric for evaluating music enhancement algorithms,” arXiv preprint arXiv: 1812.08466 , 2019
2019
Earlier work this paper cites.
M. Blaauw and J. Bonada, “Sequence-to-sequence singing synthesis using the feed-forward Transformer,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2020, pp. 7229–7233
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
X. Wang, S. Takaki, and J. Yamagishi, “Neural source-filter waveform models for statistical parametric speech synthesis,” IEEE/ACM Trans. Audio Speech Lang. Process. , vol. 28, pp. 402–415, 2020
2020
Earlier work this paper cites.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2020, pp. 6199–6203
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Advances in Neural Information Processing Systems , 2020
2020
Cited alongside, same era.
C.-C. Chu, F.-R. Yang, Y.-J. Lee, Y.-W. Liu, and S.-H. Wu, “MPop600: A Mandarin popular song database with aligned audio, lyrics, and musical scores for singing voice synthesis,” in Asia-Pacific Signal and Information Processing Association Annual Summit and Conf. , 2020, pp. 1647–1652
2020
Cited alongside, same era.
Z. Liu, K.-T. Chen, and K. Yu, “Neural homomorphic vocoder,” in INTERSPEECH , 2020
2020
Cited alongside, same era.
S. Nercessian, “End-to-end zero-shot voice conversion using a DDSP vocoder,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics , 2021, pp. 1–5
2021
Later among the works it cites.
G. Greshler, T. Shaham, and T. Michaeli, “Catch-a-waveform: Learning to generate audio from a single short example,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
M. A. M. Ramírez, O. Wang, P. Smaragdis, and N. J. Bryan, “Differentiable signal processing with black-box audio effects,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2021, pp. 66–70
2021
Later among the works it cites.
C. J. Steinmetz, J. Pons, S. Pascual, and J. Serrà, “Automatic multitrack mixing with a differentiable mixing console of neural audio effects,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2021, pp. 71–75
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Nercessian, “Zero-shot singing voice conversion,” in Int. Soc. Music Information Retrieval Conf. , 2020, pp. 70–76
2020
Cited alongside, same era.
B. Kuznetsov, J. D. Parker, and F. Esqueda, “Differentiable IIR filters for machine learning applications,” in Int. Conf. Digital Audio Effects , 2020
2020
Cited alongside, same era.
A. Gulati et al. , “Conformer: Convolution-augmented Transformer for speech recognition,” in INTERSPEECH , 2020
2020
Cited alongside, same era.
Y. Hono et al. , “Sinsy: A deep neural network-based singing voice synthesis system,” IEEE Trans. Audio, Speech and Lang. Proc. , vol. 29, pp. 2803–2815, 2021
2021
Cited alongside, same era.
Y.-P. Cho, F.-R. Yang, Y.-C. Chang, C.-T. Cheng, X.-H. Wang, and Y.-W. Liu, “A survey on recent deep learning-driven singing voice synthesis systems,” in IEEE Int. Conf. Artificial Intelligence and Virtual Reality , 2021
2021
Cited alongside, same era.
Y. Hono et al. , “PeriodNet: A non-autoregressive waveform generation model with a structure separating periodic and aperiodic components,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2021
2021
Cited alongside, same era.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” in Int. Conf. Learning Representations , 2021
2021
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” in Int. Conf. Learning Representations , 2021
2021
Cited alongside, same era.
S. Nercessian, A. Sarroff, and K. J. Werner, “Lightweight and interpretable neural modeling of an audio distortion effect using hyperconditioned differentiable biquads,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2021, pp. 890–894
2021
Later among the works it cites.
2021
Later among the works it cites.
C. J. Steinmetz and J. D. Reiss, “pyloudnorm: A simple yet flexible loudness meter in Python,” in Proc. AES Convention , 2021
2021
Later among the works it cites.
C.-F. Liao, J.-Y. Liu, and Y.-H. Yang, “KaraSinger: Score-free singing voice synthesis with VQ-VAE using Mel-spectrograms,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2022
2022
Closest in time.
J. Liu, C. Li, Y. Ren, F. Chen, P. Liu, and Z. Zhao, “DiffSinger: Singing voice synthesis via shallow diffusion mechanism,” in AAAI Conference on Artificial Intelligence , 2022
2022
Closest in time.
H. Guo, Z. Zhou, F. Meng, and K. Liu, “Improving adversarial waveform generation based singing voice conversion with harmonic signals,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2022
2022
Closest in time.
F. Chen, R. Huang, C. Cui, Y. Ren, J. Liu, and Z. Zhao, “SingGAN: Generative adversarial network for high-fidelity singing voice generation,” in ACM Multimedia , 2022
2022
Closest in time.
R. Huang, M. W. Y. Lam, J. Wang, D. Su, D. Yu, Y. Ren, and Z. Zhao, “FastDiff: A fast conditional diffusion model for high-quality speech synthesis,” in Int. Joint Conf. Artificial Intelligence , 2022
2022
Closest in time.
M. W. Y. Lam, J. Wang, D. Su, and D. Yu, “BDDM: Bilateral denoising diffusion models for fast and high-quality speech synthesis,” in Int. Conf. Learning Representations , 2022
2022
Closest in time.
S. Shan, L. Hantrakul, J. Chen, M. Avent, and D. Trevelyan, “Differentiable wavetable synthesis,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2022, pp. 4598–4602
2022
Closest in time.
B.-Y. Chen, W.-H. Hsu, W.-H. Liao, R. M. A. Martínez, Y. Mitsufuji, and Y.-H. Yang, “Automatic DJ transitions with differentiable audio effects and generative adversarial networks,” in IEEE Int. Conf. Acoustics, Speech and Signal Proc. , 2022, pp. 466–470
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
E. Deruty, M. Grachten, S. Lattner, J. Nistal, and C. Aouameur, “On the development and practice of AI technology for contemporary popular music production,” Transactions of the International Society for Music Information Retrieval , vol. 5, no. 1, pp. 35–49, 2022
2022
Closest in time.