Fetching the paper…
Reading the bibliography…
Despite recent progress in text-to-audio (TTA) generation, we show that the state-of-the-art models, such as AudioLDM, trained on datasets with an imbalanced class distribution, such as AudioCaps, are biased in their generation performance.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “AudioSet: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2017, pp. 776–780
2017
Earlier work this paper cites.
X. Chen, N. Mishra, M. Rohaninejad, and P. Abbeel, “PixelSNAIL: An improved autoregressive generative model,” in International Conference on Machine Learning , 2018, pp. 864–872
2018
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Annual Conference of the North American Chapter of the Association for Computational Linguistics , 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
T. Heittola, A. Mesaros, and T. Virtanen, “TAU Urban Acoustic Scenes 2019, Development dataset,” Mar. 2019. [Online]. Available: https://doi.org/10.5281/zenodo.2589280
2019
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Neural Information Processing Systems , 2020
2020
Earlier work this paper cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Earlier work this paper cites.
X. Liu, T. Iqbal, J. Zhao, Q. Huang, M. Plumbley, and W. Wang, “Conditional sound generation using neural discrete time-frequency representation learning,” IEEE International Workshop on Machine Learning for Signal Processing , 2021
2021
Cited alongside, same era.
V. Iashin and E. Rahtu, “Taming visually guided sound generation,” in British Machine Vision Conference , 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning , 2021, pp. 8748–8763
2021
Cited alongside, same era.
B. Li, P. H. Torr, and T. Lukasiewicz, “Memory-driven text-to-image generation,” in British Machine Vision Conference , 2022
2022
Cited alongside, same era.
2023
Closest in time.
Y. Yuan, H. Liu, J. Liang, X. Liu, M. D. Plumbley, and W. Wang, “Leveraging pre-trained audioldm for sound generation: A benchmark study,” in European Association for Signal Processing , 2023
2023
Closest in time.
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, “Diffsound: Discrete diffusion model for text-to-sound generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Closest in time.
S. Sheynin, O. Ashual, A. Polyak, U. Singer, O. Gafni, E. Nachmani, and Y. Taigman, “KNN-diffusion: Image generation via large-scale retrieval,” in International Conference on Learning Representations , 2023
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Blattmann, R. Rombach, K. Oktay, J. Müller, and B. Ommer, “Retrieval-augmented diffusion models,” Advances in Neural Information Processing Systems , vol. 35, pp. 15 309–15 324, 2022
2022
Cited alongside, same era.
H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, C. Feichtenhofer et al. , “Masked autoencoders that listen,” arXiv preprint:2207.06405 , 2022
2022
Cited alongside, same era.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-Audio generation with latent diffusion models,” in International Conference on Machine Learning , 2023
2023
Cited alongside, same era.
R. Huang, J. Huang, D. Yang, Y. Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,” in International Conference on Machine Learning , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Chen, H. Hu, C. Saharia, and W. W. Cohen, “Re-Imagen: Retrieval-augmented text-to-image generator,” in International Conference on Learning Representations , 2023
2023
Closest in time.
Y. Yuan, H. Liu, X. Kang, P. Wu, M. D. Plumbley, and W. Wang, “Text-driven foley sound generation with latent diffusion model,” in Proceedings of the Detection and Classification of Acoustic Scenes and Events Workshop , 2023, pp. 231–235
2023
Closest in time.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP , 2023
2023
Closest in time.
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi, “AudioGen: textually guided audio generation,” in International Conference on Learning Representations , 2023
2023
Closest in time.