Fetching the paper…
Reading the bibliography…
Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience.
Loudness, its definition, measurement and calculation
Fletcher, H.; and Munson, W. A. 1933 · 1933
Earlier work this paper cites.
Exponentially weighted moving average control schemes: properties and enhancements
Lucas, J. M.; and Saccucci, M. S. 1990 · 1990
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Visually indicated sounds
Owens, A.; Isola, P.; McDermott, J.; Torralba, A.; Adelson, E. H.; and Freeman, W. T. 2016 · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F.; Ellis, D. P.; Freedman, D.; Jansen, A.; Lawrence, W.; Moore, R. C.; Plakal, M.; and Ritter, M. 2017 · 2017
Earlier work this paper cites.
CNN architectures for large-scale audio classification
Hershey, S.; Chaudhuri, S.; Ellis, D. P.; Gemmeke, J. F.; Jansen, A.; Moore, R. C.; Plakal, M.; Platt, D.; Saurous, R. A.; Seybold, B.; et al. 2017 · 2017
Earlier work this paper cites.
Visual to sound: Generating natural sound for videos in the wild
Zhou, Y.; Wang, Z.; Fang, C.; Bui, T.; and Berg, T. L. 2018 · 2018
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Kim, C. D.; Kim, B.; Lee, H.; and Kim, G. 2019 · 2019
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Chen, H.; Xie, W.; Vedaldi, A.; and Zisserman, A. 2020a · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R. L.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
Taming visually guided sound generation
Iashin, V.; and Rahtu, E. 2021 · 2021
Earlier work this paper cites.
Efficient training of audio transformers with patchout
Koutini, K.; Schlüter, J.; Eghbal-Zadeh, H.; and Widmer, G. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Imagen video: High definition video generation with diffusion models
Ho, J.; Chan, W.; Saharia, C.; Whang, J.; Gao, R.; Gritsenko, A.; Kingma, D. P.; Poole, B.; Norouzi, M.; Fleet, D. J.; et al. 2022 · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J.; and Salimans, T. 2022 · 2022
Cited alongside, same era.
Masked autoencoders that listen
Huang, P.-Y.; Xu, H.; Li, J.; Baevski, A.; Auli, M.; Galuba, W.; Metze, F.; and Feichtenhofer, C. 2022 · 2022
Cited alongside, same era.
Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models
Lu, C.; Zhou, Y.; Bao, F.; Chen, J.; Li, C.; and Zhu, J. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents. arXiv 2022
I hear your true colors: Image guided audio generation
Sheffer, R.; and Adi, Y. 2023 · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; and Dubnov, S. 2023 · 2023
Later among the works it cites.
DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
Xing, J.; Xia, M.; Zhang, Y.; Chen, H.; Yu, W.; Liu, H.; Wang, X.; Wong, T.-T.; and Shan, Y. 2023 · 2023
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Yang, D.; Yu, J.; Wang, H.; Wang, W.; Weng, C.; Zou, Y.; and Yu, D. 2023 · 2023
Later among the works it cites.
Video Generation Models as World Simulators
Brooks, T.; Peebles, B.; Holmes, C.; DePue, W.; Guo, Y.; Jing, L.; Schnurr, D.; Taylor, J.; Luhman, T.; Luhman, E.; Ng, C.; Wang, R.; and Ramesh, A. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Wav2clip: Learning robust audio representations from clip
Wu, H.-H.; Seetharaman, P.; Kumar, K.; and Bello, J. P. 2022 · 2022
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y.; English, Z.; Voleti, V.; Letts, A.; et al. 2023 · 2023
Cited alongside, same era.
PixArt- α \alpha : Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Chen, J.; Yu, J.; Ge, C.; Yao, L.; Xie, E.; Wu, Y.; Wang, Z.; Kwok, J.; Luo, P.; Lu, H.; and Li, Z. 2023 · 2023
Cited alongside, same era.
AnimateAnything: Fine-Grained Open Domain Image Animation with Motion Guidance
Dai, Z.; Zhang, Z.; Yao, Y.; Qiu, B.; Zhu, S.; Qin, L.; and Wang, W. 2023 · 2023
Cited alongside, same era.
Clap learning audio concepts from natural language supervision
Elizalde, B.; Deshmukh, S.; Al Ismail, M.; and Wang, H. 2023 · 2023
Cited alongside, same era.
ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model With Knowledge-Enhanced Mixture-of-Denoising-Experts
Feng, Z.; Zhang, Z.; Yu, X.; Fang, Y.; Li, L.; Chen, X.; Lu, Y.; Liu, J.; Yin, W.; Feng, S.; Sun, Y.; Chen, L.; Tian, H.; Wu, H.; and Wang, H. 2023 · 2023
Cited alongside, same era.
Yolo-world: Real-time open-vocabulary object detection
Cheng, T.; Song, L.; Ge, Y.; Liu, W.; Wang, X.; and Shan, Y. 2024 · 2024
Closest in time.
Scaling instruction-finetuned language models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2024 · 2024
Closest in time.
Fast Timing-Conditioned Latent Audio Diffusion
Evans, Z.; Carr, C.; Taylor, J.; Hawley, S. H.; and Pons, J. 2024 · 2024
Closest in time.
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Guo, Y.; Yang, C.; Rao, A.; Liang, Z.; Wang, Y.; Qiao, Y.; Agrawala, M.; Lin, D.; and Dai, B. 2024 · 2024
Closest in time.
Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models
Luo, S.; Yan, C.; Hu, C.; and Zhao, H. 2024 · 2024
Closest in time.
Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research
Mei, X.; Meng, C.; Liu, H.; Kong, Q.; Ko, T.; Zhao, C.; Plumbley, M. D.; Zou, Y.; and Wang, W. 2024 · 2024
Closest in time.
Fast High-Resolution Image Synthesis with Latent Adversarial Diffusion Distillation
Sauer, A.; Boesel, F.; Dockhorn, T.; Blattmann, A.; Esser, P.; and Rombach, R. 2024 · 2024
Closest in time.
V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models
Wang, H.; Ma, J.; Pascual, S.; Cartwright, R.; and Cai, W. 2024 · 2024
Closest in time.
SonicVisionLM: Playing Sound with Vision Language Models
Xie, Z.; Yu, S.; Li, M.; He, Q.; Chen, C.; and Jiang, Y.-G. 2024 · 2024
Closest in time.
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
Xing, Y.; He, Y.; Tian, Z.; Wang, X.; and Chen, Q. 2024 · 2024
Closest in time.