Fetching the paper…
Reading the bibliography…
This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in
2017
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in
2019
Earlier work this paper cites.
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.” in
2019
Earlier work this paper cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
2020
Earlier work this paper cites.
K. Choi, S. Oh, M. Kang, and B. McFee, “A proposal for foley sound synthesis challenge,”
2022
Earlier work this paper cites.
K. Choi, J. Im, L. Heller, B. Mcfee, K. Imoto, Y. Okamoto, M. Lagrange, and S. Takamichi, “Foley sound synthesis at the dcase 2023 challenge,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Tailleur, J. Lee, M. Lagrange, K. Choi, L. M. Heller, K. Imoto, and Y. Okamoto, “Correlation of fréchet audio distance with human perception of environmental audio is embedding dependent,” in
2024
Earlier work this paper cites.
X. ZhiDong, L. XinYu, L. HaiCheng, Z. XiaoYan, and S. Yu, “Sound scene synthesis with audioldm and tango2 for dcase 2024 task7,” Samsung Research China-Nanjing, Nanjing, China, Tech. Rep., July 2024
2024
Earlier work this paper cites.
H. C. Chung and J. H. Jung, “Sound scene synthesis based on gan using contrastive learning and effective time-frequency swap cross attention mechanism,” KT Corporation, Seoul, Republic of Korea, Tech. Rep., July 2024
2024
Earlier work this paper cites.
Y. Yuan, H. Liu, X. Liu, M. D. Plumbley, and W. Wang, “Diffusion based sound scene synthesis for dcase challenge 2024 task 7,” University of Surrey, Guildford, United Kingdom, Tech. Rep., July 2024
2024
Earlier work this paper cites.
S. Ghosh, G. Verma, S. N. Shakya, S. Sharma, and S. Singh, “Sound scene synthesis based on fine-tuned latent diffusion model for dcase challenge 2024 task 7,” Indian Institute of Technology Mandi, Kamand, Mandi, India, Tech. Rep., July 2024
2024
Earlier work this paper cites.
J. Lee, M. Tailleur, L. M. Heller, K. Choi, M. Lagrange, B. McFee, K. Imoto, and Y. Okamoto, “Challenge on sound scene synthesis: Evaluating text-to-audio generation,” in
2024
Cited alongside, same era.
Y. Chung, J. Lee, and J. Nam, “T-foley: A controllable waveform-domain diffusion model for temporal-event-guided foley sound synthesis,” in
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Z. Guo, J. Mao, R. Tao, L. Yan, K. Ouchi, H. Liu, and X. Wang, “Audio generation with multiple conditional diffusion model,” in
2024
Cited alongside, same era.
Z. Xie, S. Yu, Q. He, and M. Li, “Sonicvisionlm: Playing sound with vision language models,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley, “Audioldm 2: Learning holistic audio generation with self-supervised pretraining,”
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Anonymous, “Fugatto 1: Foundational generative audio transformer opus 1,” in
2024
Cited alongside, same era.
Z. Evans, J. D. Parker, C. Carr, Z. Zukowski, J. Taylor, and J. Pons, “Stable audio open,”
2024
Cited alongside, same era.
2024
Cited alongside, same era.
M. Comunità, R. F. Gramaccioni, E. Postolache, E. Rodolà, D. Comminiello, and J. D. Reiss, “Syncfusion: Multimodal onset-synchronized video-to-audio foley synthesis,” in
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
I. Viertola, V. Iashin, and E. Rahtu, “Temporally aligned audio for video with autoregression,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Pascual, C. Yeh, I. Tsiamas, and J. Serrà, “Masked generative video-to-audio transformers with enhanced synchronicity,” in
2025
Closest in time.