Fetching the paper…
Reading the bibliography…
Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings.
New method of measuring reverberation time
Manfred R Schroeder · 1965
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Openair: An interactive auralization web resource and database
Damian T Murphy and Simon Shelley · 2010
Earlier work this paper cites.
Freesound technical demo
Frederic Font, Gerard Roma, and Xavier Serra · 2013
Earlier work this paper cites.
A fast griffin-lim algorithm
Nathanaël Perraudin, Peter Balazs, and Peter L Søndergaard · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Temporal multimodal learning in audiovisual speech recognition
Di Hu, Xuelong Li, et al · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Matterport3d: Learning from RGB-D data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang · 2017
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Earlier work this paper cites.
Acoustic classification and optimization for multi-modal rendering of real-world scenes
Carl Schissler, Christian Loftin, and Dinesh Manocha · 2017
Earlier work this paper cites.
The conversation: Deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
Objects that sound
Relja Arandjelović and Andrew Zisserman · 2018
Earlier work this paper cites.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Earlier work this paper cites.
Visual speech enhancement
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg · 2018
Cited alongside, same era.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Cited alongside, same era.
Co-training of audio and video representations from self-supervised temporal synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Self-supervised generation of spatial audio for 360 ∘ video
Pedro Morgado, Nono Vasconcelos, Timothy Langlois, and Oliver Wang · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Cited alongside, same era.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Cited alongside, same era.
Talking face generation by adversarially disentangled audio-visual representation
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang · 2019
Later among the works it cites.
Batvision: Learning to see 3d spatial layout with two ears
Jesper Haahr Christensen, Sascha Hornauer, and X Yu Stella · 2020
Later among the works it cites.
Facefilter: Audio-visual speech separation using still images
Soo-Whan Chung, Soyeon Choe, Joon Son Chung, and Hong-Goo Kang · 2020
Later among the works it cites.
See, hear, explore: Curiosity via audio-visual association
Victoria Dean, Shubham Tulsiani, and Abhinav Gupta · 2020
Later among the works it cites.
Visualechoes: Spatial image representation learning through echolocation
Ruohan Gao, Changan Chen, Ziad Al-Halab, Carl Schissler, and Kristen Grauman · 2020
Later among the works it cites.
Discriminative sounding objects localization via self-supervised audiovisual matching
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Audio-visual event localization in unconstrained videos
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu · 2018
Cited alongside, same era.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Cited alongside, same era.
Visual to sound: Generating natural sound for videos in the wild
Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg · 2018
Cited alongside, same era.
My lips are concealed: Audio-visual speech enhancement through obstructions
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2019
Cited alongside, same era.
Gansynth: Adversarial neural audio synthesis
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts · 2019
Cited alongside, same era.
Self-supervised audio spatialization with correspondence classifier
Yu-Ding Lu, Hsin-Ying Lee, Hung-Yu Tseng, and Ming-Hsuan Yang · 2019
Cited alongside, same era.
Di Hu, Rui Qian, Minyue Jiang, Xiao Tan, Shilei Wen, Errui Ding, Weiyao Lin, and Dejing Dou · 2020
Later among the works it cites.
Learning representations from audio-visual spatial alignment
Pedro Morgado, Yi Li, and Nuno Nvasconcelos · 2020
Later among the works it cites.
Scene-aware audio rendering via deep acoustic analysis
Zhenyu Tang, Nicholas J Bryan, Dingzeyu Li, Timothy R Langlois, and Dinesh Manocha · 2020
Later among the works it cites.
Unified multisensory perception: Weakly-supervised audio-visual video parsing
Yapeng Tian, Dingzeyu Li, and Chenliang Xu · 2020
Later among the works it cites.
Telling left from right: Learning spatial correspondence of sight and sound
Karren Yang, Bryan Russell, and Justin Salamon · 2020
Later among the works it cites.
Audio-visual recognition of overlapped speech for the lrs2 dataset
Jianwei Yu, Shi-Xiong Zhang, Jian Wu, Shahram Ghorbani, Bo Wu, Shiyin Kang, Shansong Liu, Xunying Liu, Helen Meng, and Dong Yu · 2020
Later among the works it cites.
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu · 2020
Later among the works it cites.
Semantic audio-visual navigation
Changan Chen, Ziad Al-Halah, and Kristen Grauman · 2021
Closest in time.
Move2Hear: Active audio-visual source separation
Sagnik Majumder, Ziad Al-Halah, and Kristen Grauman · 2021
Closest in time.
Neural synthesis of binaural speech from mono audio
Alexander Richard, Dejan Markovic, Israel D Gebru, Steven Krenn, Gladstone Butler, Fernando de la Torre, and Yaser Sheikh · 2021
Closest in time.
Into the wild with audioscope: Unsupervised audio-visual separation of on-screen sounds
Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey, Tal Remez, Daniel PW Ellis, and John R Hershey · 2021
Closest in time.
Visually informed binaural audio generation without binaural audios
Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin · 2021
Closest in time.