Fetching the paper…
Reading the bibliography…
Audio synthesis has broad applications in multimedia.
An Ecological Approach to Auditory Event Perception
Gaver, W. W. 1993 · 1993
Earlier work this paper cites.
Recognition of sound sources and events
McAdams, S. 1993 · 1993
Earlier work this paper cites.
Methods for subjective determination of transmission quality
Sector, I. T. U. T. S. 1996 · 1996
Earlier work this paper cites.
On clustering validation techniques
Halkidi, M.; Batistakis, Y.; and Vazirgiannis, M. 2001 · 2001
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Hadsell, R.; Chopra, S.; and LeCun, Y. 2006 · 2006
Earlier work this paper cites.
Finding a “kneedle” in a haystack: Detecting knee points in system behavior
Satopaa, V.; Albrecht, J.; Irwin, D.; and Raghavan, B. 2011 · 2011
Earlier work this paper cites.
AVA: A large-scale database for aesthetic visual analysis
Murray, N.; Marchesotti, L.; and Perronnin, F. 2012 · 2012
Earlier work this paper cites.
Sound synthesis and sampling
Russ, M. 2012 · 2012
Earlier work this paper cites.
Visually indicated sounds
Owens, A.; Isola, P.; McDermott, J.; Torralba, A.; Adelson, E. H.; and Freeman, W. T. 2016 · 2016
Earlier work this paper cites.
Deep cross-modal audio-visual generation
Chen, L.; Srivastava, S.; Duan, Z.; and Xu, C. 2017 · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Tarvainen, A.; and Valpola, H. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Visually indicated sound generation by perceptually optimized classification
Chen, K.; Zhang, C.; Fang, C.; Wang, Z.; Bui, T.; and Nevatia, R. 2018 · 2018
Earlier work this paper cites.
CMCGAN: A uniform framework for cross-modal visual-audio mutual generation
Hao, W.; Zhang, Z.; and Guan, H. 2018 · 2018
Earlier work this paper cites.
The sound of pixels
Zhao, H.; Gan, C.; Rouditchenko, A.; Vondrick, C.; McDermott, J.; and Torralba, A. 2018 · 2018
Earlier work this paper cites.
Visual to sound: Generating natural sound for videos in the wild
Zhou, Y.; Wang, Z.; Fang, C.; Bui, T.; and Berg, T. L. 2018 · 2018
Earlier work this paper cites.
Fréchet Audio Distance: A Reference-free Metric for Evaluating Music Enhancement Algorithms
Roblek, D.; Kilgour, K.; Sharifi, M.; and Zuluaga, M. 2019 · 2019
Earlier work this paper cites.
Applications of deep learning to audio generation
Zhao, Y.; Xia, X.; and Togneri, R. 2019 · 2019
Earlier work this paper cites.
Deep learning in video multi-object tracking: A survey
Ciaparrone, G.; Sánchez, F. L.; Tabik, S.; Troiano, L.; Tagliaferri, R.; and Herrera, F. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Audio in VR: Effects of a soundscape and movement-triggered step sounds on presence
Kern, A. C.; and Ellermeier, W. 2020 · 2020
Earlier work this paper cites.
Localizing Visual Sounds the Hard Way
Chen, H.; Xie, W.; Afouras, T.; Nagrani, A.; Vedaldi, A.; and Zisserman, A. 2021 · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Cited alongside, same era.
FSD50K: An Open Dataset of Human-Labeled Sound Events
Fonseca, E.; Favory, X.; Pons, J.; Font, F.; and Serra, X. 2021 · 2021
Cited alongside, same era.
The benefit of temporally-strong labels in audio event classification
Hershey, S.; Ellis, D. P.; Fonseca, E.; Jansen, A.; Liu, C.; Moore, R. C.; and Plakal, M. 2021 · 2021
Cited alongside, same era.
Classifier-Free Diffusion Guidance
Ho, J.; and Salimans, T. 2021 · 2021
Cited alongside, same era.
Taming Visually Guided Sound Generation
Iashin, V.; and Rahtu, E. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; and Plumbley, M. D. 2023 · 2023
Later among the works it cites.
Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Liu, X.; Gong, C.; and Liu, Q. 2023 · 2023
Later among the works it cites.
Diff-Foley: Synchronized video-to-audio synthesis with latent diffusion models
Luo, S.; Yan, C.; Hu, C.; and Zhao, H. 2023 · 2023
Later among the works it cites.
Audio Deepfake Approaches
Shaaban, O. A.; Yildirim, R.; and Alguttar, A. A. 2023 · 2023
Later among the works it cites.
I hear your true colors: Image guided audio generation
Sheffer, R.; and Adi, Y. 2023 · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; and Dubnov, S. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient attention: Attention with linear complexities
Shen, Z.; Zhang, M.; Zhao, H.; Yi, S.; and Li, H. 2021 · 2021
Cited alongside, same era.
MaskGIT: Masked generative image transformer
Chang, H.; Zhang, H.; Jiang, L.; Liu, C.; and Freeman, W. T. 2022 · 2022
Cited alongside, same era.
Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors
Iashin, V. E.; Xie, W.; Rahtu, E.; and Zisserman, A. 2022 · 2022
Cited alongside, same era.
Mind the gap: understanding the modality gap in multi-modal contrastive representation learning
Liang, W.; Zhang, Y.; Kwon, Y.; Yeung, S.; and Zou, J. 2022 · 2022
Cited alongside, same era.
https://github.com/openai/CLIP
OpenAI. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with CLIP latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
A survey of deep learning audio generation methods
Božić, M.; and Horvat, M. 2024 · 2024
Closest in time.
The digital Foley: what Foley artists say about using audio synthesis
Di Donato, B.; and McGregor, I. 2024 · 2024
Closest in time.
Synchformer: Efficient synchronization from sparse cues
Iashin, V.; Xie, W.; Rahtu, E.; and Zisserman, A. 2024 · 2024
Closest in time.
https://huggingface.co/nousr/conditioned-prior/tree/main/vit-l-14/aesthetic
LAION. 2024 · 2024
Closest in time.
Object-Aware Audio-Visual Sound Generation
Li, T.; Huang, B.; Zhuang, X.; Jia, D.; Chen, J.; Wang, Y.; Anumanchipalli, G.; Chen, Z.; and Wang, Y. 2024 · 2024
Closest in time.
Cyclic Learning for Binaural Audio Generation and Localization
Li, Z.; Zhao, B.; and Yuan, Y. 2024 · 2024
Closest in time.
AudioLDM 2: Learning holistic audio generation with self-supervised pretraining
Liu, H.; Yuan, Y.; Liu, X.; Mei, X.; Kong, Q.; Tian, Q.; Wang, Y.; Wang, W.; Wang, Y.; and Plumbley, M. D. 2024 · 2024
Closest in time.
https://storage.googleapis.com/openimages/web/index.html
OpenImagesV7. 2024 · 2024
Closest in time.
Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity
Pascual, S.; Yeh, C.; Tsiamas, I.; and Serrà, J. 2024 · 2024
Closest in time.
CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor
Sun, S.; Li, R.; Torr, P.; Gu, X.; and Li, S. 2024 · 2024
Closest in time.
https://github.com/gudgud96/frechet-audio-distance
Tan, H. 2024 · 2024
Closest in time.
Seeing and Hearing: Open-domain visual-audio generation with diffusion latent aligners
Xing, Y.; He, Y.; Tian, Z.; Wang, X.; and Chen, Q. 2024 · 2024
Closest in time.
Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
Yang, Q.; Mao, B.; Wang, Z.; Nie, X.; Gao, P.; Guo, Y.; Zhen, C.; Yan, P.; and Xiang, S. 2024 · 2024
Closest in time.
Efficient video to audio mapper with visual scene detection
Yi, M.; and Li, M. 2024 · 2024
Closest in time.
Video-guided foley sound generation with multimodal controls
Chen, Z.; Seetharaman, P.; Russell, B.; Nieto, O.; Bourgin, D.; Owens, A.; and Salamon, J. 2025 · 2025
Closest in time.
Taming multimodal joint training for high-quality video-to-audio synthesis
Cheng, H. K.; Ishii, M.; Hayakawa, A.; Shibuya, T.; Schwing, A.; and Mitsufuji, Y. 2025 · 2025
Closest in time.