Fetching the paper…
Reading the bibliography…
Spectrograms are 2D representations of sound that look very different from the images found in our visual world.
Lucy in the sky with diamonds, 1967
The Beatles · 1967
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
D. Griffin and J. Lim · 1984
Earlier work this paper cites.
Formula, 1994
Aphex Twin · 1994
Earlier work this paper cites.
Songs about my cats, 2001
Venetian Snares · 2001
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G. E. Hinton · 2002
Earlier work this paper cites.
Pixels that sound
E. Kidron, Y. Y. Schechner, and M. Elad · 2005
Earlier work this paper cites.
10,000 days, 2006
Tool · 2006
Earlier work this paper cites.
Year zero, 2007
Nine Inch Nails · 2007
Earlier work this paper cites.
Phase-controlled sound transfer based on maximally-inconsistent spectrograms
J. Le Roux · 2011
Earlier work this paper cites.
A fast griffin-lim algorithm
N. Perraudin, P. Balazs, and P. L. Søndergaard · 2013
Earlier work this paper cites.
Deep image features in music information retrieval
G. Gwardys and D. Grzywczak · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Audio texture synthesis and style transfer
D. Ulyanov · 2016
Earlier work this paper cites.
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Earlier work this paper cites.
Fun with spectrograms! how to make an image using sound and music
Classical Music Reimagined · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter · 2017
Earlier work this paper cites.
Objects that sound
R. Arandjelovic and A. Zisserman · 2018
Earlier work this paper cites.
Self-supervised generation of spatial audio for 360 video
P. Morgado, N. Nvasconcelos, T. Langlois, and O. Wang · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
Earlier work this paper cites.
2.5d visual sound
R. Gao and K. Grauman · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Y. Song and S. Ermon · 2019
Earlier work this paper cites.
Hiding video in audio via reversible generative models
H. Yang, H. Ouyang, V. Koltun, and Q. Chen · 2019
Earlier work this paper cites.
The sound of motions
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba · 2019
Earlier work this paper cites.
Self-supervised learning of audio-visual objects from video
T. Afouras, A. Owens, J. S. Chung, and A. Zisserman · 2020
Earlier work this paper cites.
Labelling unlabelled videos from scratch with multi-modal self-supervision
Y. Asano, M. Patrick, C. Rupprecht, and A. Vedaldi · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman · 2020
Earlier work this paper cites.
Compositional visual generation with energy based models
Y. Du, S. Li, and I. Mordatch · 2020
Earlier work this paper cites.
Visualechoes: Spatial visual representation learning through echolocation
R. Gao, C. Chen, Z. Al-Halah, C. Schissler, and K. Grauman · 2020
Earlier work this paper cites.
Listen to look: Action recognition by previewing audio
R. Gao, T.-H. Oh, K. Grauman, and L. Torresani · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
J. Kong, J. Kim, and J. Bae · 2020
Earlier work this paper cites.
Rethinking cnn models for audio classification
K. Palanisamy, D. Singhania, and A. Yao · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2020
Earlier work this paper cites.
Fourier features let networks learn high frequency functions in low dimensional domains
M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ramamoorthi, J. Barron, and R. Ng · 2020
Earlier work this paper cites.
Telling left from right: Learning spatial correspondence of sight and sound
K. Yang, B. Russell, and J. Salamon · 2020
Earlier work this paper cites.
Audio-visual synchronisation in the wild
H. Chen, W. Xie, T. Afouras, A. Nagrani, A. Vedaldi, and A. Zisserman · 2021
Earlier work this paper cites.
Structure from silence: Learning scene structure from ambient sound
Z. Chen, X. Hu, and A. Owens · 2021
Earlier work this paper cites.
Ilvr: Conditioning method for denoising diffusion probabilistic models
J. Choi, S. Kim, Y. Jeong, Y. Gwon, and S. Yoon · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
Visualvoice: Audio-visual speech separation with cross-modal consistency
R. Gao and K. Grauman · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi · 2021
Earlier work this paper cites.
Taming visually guided sound generation
V. Iashin and E. Rahtu · 2021
Cited alongside, same era.
Diffusion probabilistic models for 3d point cloud generation
S. Luo and W. Hu · 2021
Cited alongside, same era.
Sdedit: Guided image synthesis and editing with stochastic differential equations
C. Meng, Y. He, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon · 2021
Cited alongside, same era.
Audio-visual instance discrimination with cross-modal agreement
P. Morgado, N. Vasconcelos, and I. Misra · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2021
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Cited alongside, same era.
Text2room: Extracting textured 3d meshes from 2d text-to-image models
L. Höllein, A. Cao, A. Owens, J. Johnson, and M. Nießner · 2023
Later among the works it cites.
Egocentric audio-visual object localization
C. Huang, Y. Tian, A. Kumar, and C. Xu · 2023
Later among the works it cites.
Epic-sounds: A large-scale dataset of actions that sound
J. Huh, J. Chalk, E. Kazakos, D. Damen, and A. Zisserman · 2023
Later among the works it cites.
Vision transformers are parameter-efficient audio-visual learners
Y.-B. Lin, Y.-L. Sung, J. Lei, M. Bansal, and G. Bertasius · 2023
Later among the works it cites.
Audioldm: Text-to-audio generation with latent diffusion models
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley · 2023
Later among the works it cites.
Audioldm 2: Learning holistic audio generation with self-supervised pretraining
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Spectrogram art: A short history of musicians hiding visuals inside their tracks
B. Buckle · 2022
Cited alongside, same era.
Visual acoustic matching
C. Chen, R. Gao, P. Calamia, and K. Grauman · 2022
Cited alongside, same era.
Sound localization by self-supervised time delay estimation
Z. Chen, D. F. Fouhey, and A. Owens · 2022
Cited alongside, same era.
Riffusion - Stable diffusion for real-time music generation, 2022
S. Forsgren and H. Martiros · 2022
Cited alongside, same era.
Contrastive audio-visual masked autoencoder
Y. Gong, A. Rouditchenko, A. H. Liu, D. Harwath, L. Karlinsky, H. Kuehne, and J. Glass · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y. Pritch, and D. Cohen-Or · 2022
Cited alongside, same era.
H. Liu, Q. Tian, Y. Yuan, X. Liu, X. Mei, Q. Kong, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley · 2023
Later among the works it cites.
Zero-1-to-3: Zero-shot one image to 3d object
R. Liu, R. Wu, B. Van Hoorick, P. Tokmakov, S. Zakharov, and C. Vondrick · 2023
Later among the works it cites.
Learning spatial features from audio-visual correspondence in egocentric videos
S. Majumder, Z. Al-Halah, and K. Grauman · 2023
Later among the works it cites.
Foleygen: Visually-guided audio generation
X. Mei, V. Nagaraja, G. L. Lan, Z. Ni, E. Chang, Y. Shi, and V. Chandra · 2023
Later among the works it cites.
Audio-visual glance network for efficient video recognition
M. A. Nugroho, S. Woo, S. Lee, and C. Kim · 2023
Later among the works it cites.
Sympawnies: animal portraits made of musical notations
N. Oxman · 2023
Later among the works it cites.
Dreamfusion: Text-to-3d using 2d diffusion
B. Poole, A. Jain, J. T. Barron, and B. Mildenhall · 2023
Later among the works it cites.
Fatezero: Fusing attentions for zero-shot text-based video editing
C. Qi, X. Cun, Y. Zhang, C. Lei, X. Wang, Y. Shan, and Q. Chen · 2023
Later among the works it cites.
Sound source localization is all about cross-modal alignment
A. Senocak, H. Ryu, J. Kim, T.-H. Oh, H. Pfister, and J. S. Chung · 2023
Later among the works it cites.
Eventfulness for interactive video alignment
J. Sun, L. Deng, T. Afouras, A. Owens, and A. Davis · 2023
Later among the works it cites.
Sound to visual scene generation by audio-to-visual latent alignment
K. Sung-Bin, A. Senocak, H. Ha, A. Owens, and T.-H. Oh · 2023
Later among the works it cites.
Zero-shot image restoration using denoising diffusion null-space model
Y. Wang, J. Yu, and J. Zhang · 2023
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, Y. Shi, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov · 2023
Later among the works it cites.
Lumiere: A space-time diffusion model for video generation
O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, Y. Li, T. Michaeli, et al · 2024
Closest in time.
R. Bensadoun, T. Monnier, Y. Kleiman, F. Kokkinos, Y. Siddiqui, M. Kariya, O. Harosh, R. Shapovalov, B. Graham, E. Garreau, et al · 2024
Closest in time.
Lightplane: Highly-scalable components for neural 3d fields
A. Cao, J. Johnson, A. Vedaldi, and D. Novotny · 2024
Closest in time.
Real acoustic fields: An audio-visual room acoustics dataset and benchmark
Z. Chen, I. D. Gebru, C. Richardt, A. Kumar, W. Laney, A. Owens, and A. Richard · 2024
Closest in time.
Fast timing-conditioned latent audio diffusion
Z. Evans, C. Carr, J. Taylor, S. H. Hawley, and J. Pons · 2024
Closest in time.
Cat3d: Create anything in 3d with multi-view diffusion models
R. Gao, A. Holynski, P. Henzler, A. Brussee, R. Martin-Brualla, P. Srinivasan, J. T. Barron, and B. Poole · 2024
Closest in time.
Motion guidance: Diffusion-based image editing with differentiable motion estimators
D. Geng and A. Owens · 2024
Closest in time.
Factorized diffusion: Perceptual illusions by noise decomposition
D. Geng, I. Park, and A. Owens · 2024
Closest in time.
Visual anagrams: Generating multi-view optical illusions with diffusion models
D. Geng, I. Park, and A. Owens · 2024
Closest in time.
Synchformer: Efficient synchronization from sparse cues
V. Iashin, W. Xie, E. Rahtu, and A. Zisserman · 2024
Closest in time.
Av-nerf: Learning neural fields for real-world audio-visual scene synthesis
S. Liang, C. Huang, Y. Tian, A. Kumar, and C. Xu · 2024
Closest in time.
Siamese vision transformers are scalable audio-visual learners
Y.-B. Lin and G. Bertasius · 2024
Closest in time.
T-vsl: Text-guided visual sound source localization in mixtures
T. Mahmud, Y. Tian, and D. Marculescu · 2024
Closest in time.
Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization
N. Majumder, C.-Y. Hung, D. Ghosal, W.-N. Hsu, R. Mihalcea, and S. Poria · 2024
Closest in time.
Maskmark: Robust neuralwatermarking for real and synthetic speech
P. O’Reilly, Z. Jin, J. Su, and B. Pardo · 2024
Closest in time.
Can clip help sound source localization?
S. Park, A. Senocak, and J. S. Chung · 2024
Closest in time.
Proactive detection of voice cloning with localized watermarking
R. San Roman, P. Fernandez, H. Elsahar, A. Défossez, T. Furon, and T. Tran · 2024
Closest in time.
Self-supervised visual acoustic matching
A. Somayazulu, C. Chen, and K. Grauman · 2024
Closest in time.
Sonicvisionlm: Playing sound with vision language models
Z. Xie, S. Yu, M. Li, Q. He, C. Chen, and Y.-G. Jiang · 2024
Closest in time.
Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation
J. Xue, Y. Deng, Y. Gao, and Y. Li · 2024
Closest in time.
Cameras as rays: Pose estimation via ray diffusion
J. Y. Zhang, A. Lin, M. Kumar, T.-H. Yang, D. Ramanan, and S. Tulsiani · 2024
Closest in time.
Thinimg: Cross-modal steganography for presenting talking heads in images
L. Zhao, H. Li, X. Ning, and X. Jiang · 2024
Closest in time.