Fetching the paper…
Reading the bibliography…
Latent diffusion models have shown promising results in audio generation, making notable advancements over traditional methods.
P. Taylor, Text-to-speech synthesis . Cambridge university press, 2009
2009
Earlier work this paper cites.
F. R. Sharma and S. G. Wasson, “Speech recognition and synthesis tool: assistive technology for physically disabled persons,” International Journal of Computer Science and Telecommunications , vol. 3, no. 4, pp. 86–91, 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18 . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
H.-W. Dong, W.-Y. Hsiao, L.-C. Yang, and Y.-H. Yang, “Musegan: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2019, pp. 119–132
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684–10 695
2022
Cited alongside, same era.
R. Huang, Z. Zhao, H. Liu, J. Liu, C. Cui, and Y. Ren, “Prodiff: Progressive fast diffusion model for high-quality text-to-speech,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 2595–2605
2022
Cited alongside, same era.
I. Vovk, T. Sadekova, V. Gogoryan, V. Popov, M. A. Kudinov, and J. Wei, “Fast grad-tts: Towards efficient diffusion-based speech generation on cpu.” in Interspeech , 2022, pp. 838–842
2022
Cited alongside, same era.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans et al. , “Photorealistic text-to-image diffusion models with deep language understanding,” Advances in neural information processing systems , vol. 35, pp. 36 479–36 494, 2022
2023
Later among the works it cites.
J. Chen, X. Song, Z. Peng, B. Zhang, F. Pan, and Z. Wu, “Lightgrad: Lightweight diffusion probabilistic model for text-to-speech,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman, “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 22 500–22 510
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
Z. Chen, X. Tan, K. Wang, S. Pan, D. Mandic, L. He, and S. Zhao, “Infergrad: Improving diffusion models for vocoder by considering inference in training,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 8432–8436
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, “Diffsound: Discrete diffusion model for text-to-sound generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Z. Ye, W. Xue, X. Tan, J. Chen, Q. Liu, and Y. Guo, “Comospeech: One-step speech and singing voice synthesis via consistency model,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 1831–1839
2023
Cited alongside, same era.
2023
Later among the works it cites.
D. Bolya and J. Hoffman, “Token merging for fast stable diffusion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 4598–4602
2023
Later among the works it cites.
X. Yang, D. Zhou, J. Feng, and X. Wang, “Diffusion probabilistic model made slim,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 22 552–22 562
2023
Later among the works it cites.
2023
Later among the works it cites.
B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “Clap learning audio concepts from natural language supervision,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
G. Fang, X. Ma, and X. Wang, “Structural pruning for diffusion models,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.