Fetching the paper…
Reading the bibliography…
Visual and auditory perception are two crucial ways humans experience the world.
2013
Earlier work this paper cites.
S. Böck and G. Widmer, “Maximum filter vibrato suppression for onset detection,” in Proc. of the 16th Int. Conf. on Digital Audio Effects (DAFx). Maynooth, Ireland (Sept 2013) , vol. 7. Citeseer, 2013, p. 4
2013
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in neural information processing systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “Vggsound: A large-scale audio-visual dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 721–725
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
V. Iashin and E. Rahtu, “Taming visually guided sound generation,” in BMVA , 2021
2021
Earlier work this paper cites.
A. Ali, I. Schwartz, T. Hazan, and L. Wolf, “Video and text matching with conditioned embeddings,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision , 2022, pp. 1565–1574
2022
Earlier work this paper cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598 , 2022
2022
Earlier work this paper cites.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “Audioldm: text-to-audio generation with latent diffusion models,” in Proceedings of the 40th International Conference on Machine Learning , 2023, pp. 21 450–21 474
2023
Earlier work this paper cites.
D. Ghosal, N. Majumder, A. Mehrish, and S. Poria, “Text-to-audio generation using instruction guided latent diffusion model,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 3590–3598
2023
Earlier work this paper cites.
L. Ruan, Y. Ma, H. Yang, H. He, B. Liu, J. Fu, N. J. Yuan, Q. Jin, and B. Guo, “Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 10 219–10 228
2023
Earlier work this paper cites.
R. Sheffer and Y. Adi, “I hear your true colors: Image guided audio generation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Cited alongside, same era.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3836–3847
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Cited alongside, same era.
H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley, “Audioldm 2: Learning holistic audio generation with self-supervised pretraining,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
Cited alongside, same era.
Y. Xing, Y. He, Z. Tian, X. Wang, and Q. Chen, “Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 7151–7161
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
S. Luo, C. Yan, C. Hu, and H. Zhao, “Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
M. Comunità, R. F. Gramaccioni, E. Postolache, E. Rodolà, D. Comminiello, and J. D. Reiss, “Syncfusion: Multimodal onset-synchronized video-to-audio foley synthesis,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 936–940
2024
Closest in time.
G. Yariv, I. Gat, S. Benaim, L. Wolf, I. Schwartz, and Y. Adi, “Diverse and aligned audio-to-video generation via text-to-video model adaptation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, 2024, pp. 6639–6647
2024
Closest in time.
S. Zhao, D. Chen, Y.-C. Chen, J. Bao, S. Hao, L. Yuan, and K.-Y. K. Wong, “Uni-controlnet: All-in-one control to text-to-image diffusion models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
I. Schwartz, S. Yu, T. Hazan, and A. G. Schwing, “Factor graph attention,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 2039–2048
2048
Closest in time.