Fetching the paper…
Reading the bibliography…
Binaural audio provides a listener with 3D sound sensation, allowing a rich perceptual experience of the scene.
Subjective effects in binaural hearing
W Koenig · 1950
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
John R Hershey and Javier R Movellan · 2000
Earlier work this paper cites.
Learning joint statistical models for audio-visual fusion and segregation
John W Fisher III, Trevor Darrell, William T Freeman, and Paul A Viola · 2001
Earlier work this paper cites.
Real-time speaker localization and speech separation by audio-visual integration
Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi G Okuno, and Hiroaki Kitano · 2002
Earlier work this paper cites.
A 3d ambisonic based binaural sound reproduction system
Markus Noisternig, Alois Sontacchi, Thomas Musil, and Robert Holdrich · 2003
Earlier work this paper cites.
Audio/visual independent components
Paris Smaragdis and Michael Casey · 2003
Earlier work this paper cites.
Blind separation of speech mixtures via time-frequency masking
Ozgur Yilmaz and Scott Rickard · 2004
Earlier work this paper cites.
Pixels that sound
Einat Kidron, Yoav Y Schechner, and Michael Elad · 2005
Earlier work this paper cites.
Harmony in motion
Zohar Barzelay and Yoav Y Schechner · 2007
Earlier work this paper cites.
Supervised and semi-supervised separation of sounds from single-channel mixtures
Paris Smaragdis, Bhiksha Raj, and Madhusudana Shashanka · 2007
Earlier work this paper cites.
Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria
Tuomas Virtanen · 2007
Earlier work this paper cites.
Source separation based on binaural cues and source model constraints
Ron J Weiss, Michael I Mandel, and Daniel PW Ellis · 2008
Earlier work this paper cites.
Source-filter based clustering for monaural blind source separation
Martin Spiertz and Volker Gnann · 2009
Earlier work this paper cites.
Under-determined reverberant audio source separation using a full-rank spatial covariance model
Ngoc QK Duong, Emmanuel Vincent, and Rémi Gribonval · 2010
Earlier work this paper cites.
Single image depth estimation from predicted semantic labels
Beyang Liu, Stephen Gould, and Daphne Koller · 2010
Earlier work this paper cites.
The cocktail party robot: Sound source separation and localisation with an active binaural head
Antoine Deleforge and Radu Horaud · 2012
Earlier work this paper cites.
Deep learning for monaural speech separation
Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, and Paris Smaragdis · 2014
Earlier work this paper cites.
Personalization of head-related transfer functions in the median plane based on the anthropometry of the listener’s pinnae
Kazuhiro Iida, Yohji Ishii, and Shinsuke Nishioka · 2014
Cited alongside, same era.
Spatial transformations for the alteration of ambisonic recordings
Matthias Kronlachner · 2014
Cited alongside, same era.
mir_eval: A transparent implementation of common mir metrics
Colin Raffel, Brian McFee, Eric J Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, Daniel PW Ellis, and C Colin Raffel · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Cited alongside, same era.
Personalization of head-related transfer functions (hrtf) based on automatic photo-anthropometry and inference from a database
Deep learning based binaural speech separation in reverberant environments
Xueliang Zhang and DeLiang Wang · 2017
Later among the works it cites.
Generative modeling of audible shapes for object perception
Zhoutong Zhang, Jiajun Wu, Qiujia Li, Zhengjia Huang, James Traer, Josh H. McDermott, Joshua B. Tenenbaum, and William T. Freeman · 2017
Later among the works it cites.
The conversation: Deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Closest in time.
Objects that sound
Relja Arandjelović and Andrew Zisserman · 2018
Closest in time.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Closest in time.
Visual speech enhancement
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Edgar A Torres-Gallegos, Felipe Orduna-Bustamante, and Fernando Arámbula-Cosío · 2015
Cited alongside, same era.
Seeing the sound: a new multimodal imaging device for computer vision
A. Zunino, M. Crocco, S. Martelli, A. Trucco, A. Bue, and V. Murino · 2015
Cited alongside, same era.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman · 2016
Cited alongside, same era.
Ambient sound provides supervision for visual learning
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba · 2016
Cited alongside, same era.
Two multimodal approaches for single microphone source separation
Farnaz Sedighin, Massoud Babaie-Zadeh, Bertrand Rivet, and Christian Jutten · 2016
Cited alongside, same era.
Closest in time.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Closest in time.
Im2flow: Motion hallucination from static images for action recognition
Ruohan Gao, Bo Xiong, and Kristen Grauman · 2018
Closest in time.
Jointly discovering visual objects and spoken words from raw sensory input
David Harwath, Adrià Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass · 2018
Closest in time.
Co-training of audio and video representations from self-supervised temporal synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Closest in time.
Scene-aware audio for 360° videos
Dingzeyu Li, Timothy R. Langlois, and Changxi Zheng · 2018
Closest in time.
Self-supervised generation of spatial audio for 360
Pedro Morgado, Nono Vasconcelos, Timothy Langlois, and Oliver Wang · 2018
Closest in time.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Closest in time.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Closest in time.
Audio-visual event localization in unconstrained videos
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu · 2018
Closest in time.
Supervised speech separation based on deep learning: An overview
DeLiang Wang and Jitong Chen · 2018
Closest in time.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Closest in time.
Visual to sound: Generating natural sound for videos in the wild
Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg · 2018
Closest in time.