Fetching the paper…
Reading the bibliography…
The sound effects that designers add to videos are designed to convey a particular artistic effect and, thus, may be quite different from a scene's true sound.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Audio-visual sound separation via hidden markov models
John Hershey and Michael Casey · 2001
Earlier work this paper cites.
Image analogies
Aaron Hertzmann, Charles E Jacobs, Nuria Oliver, Brian Curless, and David H Salesin · 2001
Earlier work this paper cites.
Foleyautomatic: physically-based sound effects for interactive simulation and animation
Kees Van Den Doel, Paul G Kry, and Dinesh K Pai · 2001
Earlier work this paper cites.
Matching local self-similarities across images and videos
Eli Shechtman and Michal Irani · 2007
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
The Foley grail: The art of performing sound for film, games, and animation
Vanessa Theme Ament · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Inverse-foley animation: synchronizing rigid-body motions to sound
Timothy R Langlois and Doug L James · 2014
Earlier work this paper cites.
Mir_eval: A transparent implementation of common mir metrics
Colin Raffel, Brian McFee, Eric J Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, Daniel PW Ellis, and C Colin Raffel · 2014
Earlier work this paper cites.
A neural algorithm of artistic style
Leon A Gatys, Alexander S Ecker, and Matthias Bethge · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Image style transfer using convolutional neural networks
Leon A Gatys, Alexander S Ecker, and Matthias Bethge · 2016
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution, 2016
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Earlier work this paper cites.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman · 2016
Earlier work this paper cites.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman · 2016
Earlier work this paper cites.
Audio texture synthesis and style transfer
Dmitry Ulyanov · 2016
Earlier work this paper cites.
Vid2speech: speech reconstruction from silent video
Ariel Ephrat and Shmuel Peleg · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron Weiss, and Kevin Wilson · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Real-time user-guided image colorization with learned deep priors
Richard Zhang, Jun-Yan Zhu, Phillip Isola, Xinyang Geng, Angela S Lin, Tianhe Yu, and Alexei A Efros · 2017
Cited alongside, same era.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros · 2017
Cited alongside, same era.
Objects that sound
Relja Arandjelovic and Andrew Zisserman · 2018
Cited alongside, same era.
Javier Nistal, Stefan Lattner, and Gael Richard · 2020
Later among the works it cites.
Learning individual speaking styles for accurate lip to speech synthesis
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Later among the works it cites.
Finding, visualizing, and quantifying latent structure across diverse animal vocal repertoires
Tim Sainburg, Marvin Thielk, and Timothy Q Gentner · 2020
Later among the works it cites.
Multi-instrumentalist net: Unsupervised generation of music from body movements
Kun Su, Xiulong Liu, and Eli Shlizerman · 2020
Later among the works it cites.
Into the wild with audioscope: Unsupervised audio-visual separation of on-screen sounds
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abe Davis and Maneesh Agrawala · 2018
Cited alongside, same era.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Cited alongside, same era.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Cited alongside, same era.
Timbretron: A wavenet (cyclegan (cqt (audio))) pipeline for musical timbre transfer
Sicong Huang, Qiyang Li, Cem Anil, Xuchan Bao, Sageev Oore, and Roger B Grosse · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Neural style transfer for audio spectograms
Prateek Verma and Julius O Smith · 2018
Cited alongside, same era.
Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey, Tal Remez, Daniel PW Ellis, and John R Hershey · 2020
Later among the works it cites.
Few-shot sound event detection
Yu Wang, Justin Salamon, Nicholas J Bryan, and Juan Pablo Bello · 2020
Later among the works it cites.
Telling left from right: Learning spatial correspondence of sight and sound
Karren Yang, Bryan Russell, and Justin Salamon · 2020
Later among the works it cites.
Structure from silence: Learning scene structure from ambient sound
Ziyang Chen, Xixi Hu, and Andrew Owens · 2021
Later among the works it cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Later among the works it cites.
Geometry-aware multi-task learning for binaural audio generation from video
Rishabh Garg, Ruohan Gao, and Kristen Grauman · 2021
Later among the works it cites.
Sanchita Ghose and John J Prevost · 2021
Later among the works it cites.
Taming visually guided sound generation
Vladimir Iashin and Esa Rahtu · 2021
Later among the works it cites.
Sound-guided semantic image manipulation
Seung Hyun Lee, Wonseok Roh, Wonmin Byeon, Sang Ho Yoon, Chan Young Kim, Jinkyu Kim, and Sangpil Kim · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
How does it sound?
Kun Su, Xiulong Liu, and Eli Shlizerman · 2021
Later among the works it cites.
Who calls the shots? rethinking few-shot learning for audio
Yu Wang, Nicholas J Bryan, Justin Salamon, Mark Cartwright, and Juan Pablo Bello · 2021
Later among the works it cites.
Visually informed binaural audio generation without binaural audios
Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin · 2021
Later among the works it cites.
Repetitive activity counting by sight and sound
Yunhua Zhang, Ling Shao, and Cees GM Snoek · 2021
Later among the works it cites.
Changan Chen, Ruohan Gao, Paul Calamia, and Kristen Grauman · 2022
Later among the works it cites.
Chenye Cui, Yi Ren, Jinglin Liu, Rongjie Huang, and Zhou Zhao · 2022
Later among the works it cites.
Sparse in space and time: Audio-visual synchronisation with trainable selectors
Vladimir Iashin, Weidi Xie, Esa Rahtu, and Andrew Zisserman · 2022
Later among the works it cites.
Learning visual styles from audio-visual associations
Tingle Li, Yichen Liu, Andrew Owens, and Hang Zhao · 2022
Later among the works it cites.
It’s time for artistic correspondence in music and video
Dídac Surís, Carl Vondrick, Bryan Russell, and Justin Salamon · 2022
Later among the works it cites.