Fetching the paper…
Reading the bibliography…
Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel.
Visual to sound: Generating natural sound for videos in the wild
Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg · 1904
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Prediction-driven computational auditory scene analysis
Daniel Patrick Whittlesey Ellis · 1996
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
John R Hershey and Javier R Movellan · 2000
Earlier work this paper cites.
Independent component analysis: algorithms and applications
Aapo Hyvärinen and Erkki Oja · 2000
Earlier work this paper cites.
Learning joint statistical models for audio-visual fusion and segregation
John W Fisher III, Trevor Darrell, William T Freeman, and Paul A Viola · 2001
Earlier work this paper cites.
Audio/visual independent components
Paris Smaragdis and Michael Casey · 2003
Earlier work this paper cites.
Sound source separation using sparse coding with temporal continuity objective
Tuomas Virtanen · 2003
Earlier work this paper cites.
Pixels that sound
Einat Kidron, Yoav Y Schechner, and Michael Elad · 2005
Earlier work this paper cites.
Cosegmentation of image pairs by histogram matching-incorporating a global constraint into mrfs
Carsten Rother, Tom Minka, Andrew Blake, and Vladimir Kolmogorov · 2006
Earlier work this paper cites.
Harmony in motion
Zohar Barzelay and Yoav Y Schechner · 2007
Earlier work this paper cites.
Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria
Tuomas Virtanen · 2007
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Nonnegative matrix factorization with the itakura-saito divergence: With application to music analysis
Cédric Févotte, Nancy Bertin, and Jean-Louis Durrieu · 2009
Earlier work this paper cites.
Source-filter based clustering for monaural blind source separation
Martin Spiertz and Volker Gnann · 2009
Earlier work this paper cites.
Clustering nmf basis functions using shifted nmf for monaural sound source separation
Rajesh Jaiswal, Derry FitzGerald, Dan Barry, Eugene Coyle, and Scott Rickard · 2011
Earlier work this paper cites.
Nmf-based environmental sound source separation using time-variant gain features
Satoshi Innami and Hiroyuki Kasai · 2012
Earlier work this paper cites.
Joint and individual variation explained (jive) for integrated analysis of multiple data types
Eric F Lock, Katherine A Hoadley, James Stephen Marron, and Andrew B Nobel · 2013
Earlier work this paper cites.
Selective search for object recognition
Jasper RR Uijlings, Koen EA van de Sande, Theo Gevers, and Arnold WM Smeulders · 2013
Cited alongside, same era.
Deep learning for monaural speech separation
Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, and Paris Smaragdis · 2014
Cited alongside, same era.
mir_eval: A transparent implementation of common mir metrics
Colin Raffel, Brian McFee, Eric J Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, Daniel PW Ellis, and C Colin Raffel · 2014
Cited alongside, same era.
Nmf-based blind source separation using a linear predictive coding error clustering criterion
Xin Guo, Stefan Uhlich, and Yuki Mitsufuji · 2015
Cited alongside, same era.
Joint optimization of masks and deep recurrent neural networks for monaural source separation
Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, and Paris Smaragdis · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Audio-visual object localization and separation using low-rank and sparsity
Jie Pu, Yannis Panagakis, Stavros Petridis, and Maja Pantic · 2017
Later among the works it cites.
Permutation invariant training of deep models for speaker-independent multi-talker speech separation
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen · 2017
Later among the works it cites.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Later among the works it cites.
The conversation: Deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Later among the works it cites.
Objects that sound
Relja Arandjelović and Andrew Zisserman · 2018
Later among the works it cites.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Cited alongside, same era.
Deep karaoke: Extracting vocals from musical mixtures using a convolutional deep neural network
Andrew JR Simpson, Gerard Roma, and Mark D Plumbley · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Deep clustering: Discriminative embeddings for segmentation and separation
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe · 2016
Cited alongside, same era.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman · 2016
Cited alongside, same era.
Two multimodal approaches for single microphone source separation
Farnaz Sedighin, Massoud Babaie-Zadeh, Bertrand Rivet, and Christian Jutten · 2016
Cited alongside, same era.
Later among the works it cites.
Visual speech enhancement
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg · 2018
Later among the works it cites.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Later among the works it cites.
Self-supervised generation of spatial audio for 360
Pedro Morgado, Nono Vasconcelos, Timothy Langlois, and Oliver Wang · 2018
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Later among the works it cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Later among the works it cites.
Adversarial semi-supervised audio source separation applied to singing voice extraction
Daniel Stoller, Sebastian Ewert, and Simon Dixon · 2018
Later among the works it cites.
Audio-visual event localization in unconstrained videos
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu · 2018
Later among the works it cites.
Supervised speech separation based on deep learning: An overview
DeLiang Wang and Jitong Chen · 2018
Later among the works it cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Later among the works it cites.
2.5d visual sound
Ruohan Gao and Kristen Grauman · 2019
Closest in time.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Closest in time.