Fetching the paper…
Reading the bibliography…
Recently, denoising diffusion models have demonstrated remarkable performance among generative models in various domains.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual Evaluation of Speech Quality (PESQ)-A New Method for Speech Quality Assessment of Telephone Networks and Codecs,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , vol. 2. IEEE, 2001, pp. 749–752
2001
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,” in International Conference on Machine Learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
M. Müller, “Dynamic Time Warping,” Information retrieval for music and motion , pp. 69–84, 2007
2007
Earlier work this paper cites.
2013
Earlier work this paper cites.
D. Rezende and S. Mohamed, “Variational Inference with Normalizing Flows,” in International Conference on Machine Learning . PMLR, 2015, pp. 1530–1538
2015
Earlier work this paper cites.
G. Degottex, L. Ardaillon, and A. Roebel, “Multi-Frame Amplitude Envelope Estimation for Modification of Singing Voice,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 7, pp. 1242–1254, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther, “Autoencoding beyond Pixels Using a Learned Similarity Metric,” in International Conference on Machine Learning . PMLR, 2016, pp. 1558–1566
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural Discrete Representation Learning,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “Melgan: Generative Adversarial Networks for Conditional Waveform Synthesis,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
R. Prenger, R. Valle, and B. Catanzaro, “WaveGlow: A Flow-based Generative Network for Speech Synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 3617–3621
2019
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative Adversarial Networks,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel WaveGAN: A Fast Waveform Generation Model based on Generative Adversarial Networks with Multi-resolution Spectrogram,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2020, pp. 6199–6203
2020
Earlier work this paper cites.
Y. Ai and Z.-H. Ling, “A Neural Vocoder with Hierarchical Generation of Amplitude and Phase Spectra for Statistical Parametric Speech Synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 839–851, 2020
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” Advances in Neural Information Processing Systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A Simple Framework for Contrastive Learning of Visual Representations,” in International Conference on Machine Learning . PMLR, 2020, pp. 1597–1607
2020
Cited alongside, same era.
S. Choi, W. Kim, S. Park, S. Yong, and J. Nam, “Children’s Song Dataset for Singing Voice Research,” in International Society for Music Information Retrieval Conference (ISMIR) , 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-supervised Learning of Speech Representations,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 449–12 460, 2020
2020
Cited alongside, same era.
S. Choi and J. Nam, “A Melody-Unsupervision Model for Singing Voice Synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7242–7246
2022
Later among the works it cites.
J. Liu, C. Li, Y. Ren, F. Chen, and Z. Zhao, “DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 10, 2022, pp. 11 020–11 028
2022
Later among the works it cites.
2022
Later among the works it cites.
S.-H. Lee, H.-R. Noh, W.-J. Nam, and S.-W. Lee, “Duration Controllable Voice Conversion via Phoneme-based Information Bottleneck,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1173–1183, 2022
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-tts: A Generative Flow for Text-to-Speech via Monotonic Alignment Search,” Advances in Neural Information Processing Systems , vol. 33, pp. 8067–8077, 2020
2020
Cited alongside, same era.
Y. Hono, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Sinsy: A Deep Neural Network-based Singing Voice Synthesis System,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2803–2815, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An End-to-End Neural Audio Codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
P. Dhariwal and A. Nichol, “Diffusion Models Beat GANs on Image Synthesis,” Advances in Neural Information Processing Systems , vol. 34, pp. 8780–8794, 2021
2021
Cited alongside, same era.
Later among the works it cites.
Y. Zhang, J. Cong, H. Xue, L. Xie, P. Zhu, and M. Bi, “ViSinger: Variational Inference with Adversarial Learning for End-to-End Singing Voice Synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7237–7241
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution Image Synthesis with Latent Diffusion Models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 684–10 695
2022
Later among the works it cites.
2022
Later among the works it cites.
H. Kim, S. Kim, and S. Yoon, “Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier Guidance,” in International Conference on Machine Learning . PMLR, 2022, pp. 11 119–11 133
2022
Later among the works it cites.
2022
Later among the works it cites.
S.-H. Lee, S.-B. Kim, J.-H. Lee, E. Song, M.-J. Hwang, and S.-W. Lee, “HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis,” in Advances in Neural Information Processing Systems , 2022
2022
Later among the works it cites.
K. Qian, Y. Zhang, H. Gao, J. Ni, C.-I. Lai, D. Cox, M. Hasegawa-Johnson, and S. Chang, “Contentvec: An Improved Self-supervised Speech Representation by Disentangling Speakers,” in International Conference on Machine Learning . PMLR, 2022, pp. 18 003–18 017
2022
Later among the works it cites.
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, “Diffsound: Discrete Diffusion model for Text-to-Sound Generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.