Fetching the paper…
Reading the bibliography…
The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean.
Self-supervised learning by cross-modal audio-video clustering
Alwassel, H.; Mahajan, D.; Torresani, L.; Ghanem, B.; and Tran, D. 2019 · 1911
Earlier work this paper cites.
The merging of the senses
Stein, B. E.; and Meredith, M. A. 1993 · 1993
Earlier work this paper cites.
Multi-modal self-supervision from generalized data transformations
Patrick, M.; Asano, Y. M.; Fong, R.; Henriques, J. F.; Zweig, G.; and Vedaldi, A. 2020 · 2003
Earlier work this paper cites.
Multiple Sound Sources Localization from Coarse to Fine
Qian, R.; Hu, D.; Dinkel, H.; Wu, M.; Xu, N.; and Lin, W. 2020 · 2007
Earlier work this paper cites.
Revisiting Mid-Level Patterns for Distant-Domain Few-Shot Recognition
Zou, Y.; Zhang, S.; Moura, J. M.; Yu, J.; and Tian, Y. 2020 · 2008
Earlier work this paper cites.
Discriminative Sounding Objects Localization via Self-supervised Audiovisual Matching
Hu, D.; Qian, R.; Jiang, M.; Tan, X.; Wen, S.; Ding, E.; Lin, W.; and Dou, D. 2020 · 2010
Earlier work this paper cites.
Multisensory perceptual learning and sensory substitution
Proulx, M. J.; Brown, D. J.; Pasqualotto, A.; and Meijer, P. 2014 · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J.; Clune, J.; Bengio, Y.; and Lipson, H. 2014 · 2014
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Aytar, Y.; Vondrick, C.; and Torralba, A. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Look, listen and learn
Arandjelovic, R.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017 · 2017
Cited alongside, same era.
What makes audio event detection harder than classification?
Phan, H.; Koch, P.; Katzberg, F.; Maass, M.; Mazur, R.; McLoughlin, I.; and Mertins, A. 2017 · 2017
Cited alongside, same era.
Audio-visual object localization and separation using low-rank and sparsity
Pu, J.; Panagakis, Y.; Petridis, S.; and Pantic, M. 2017 · 2017
Cited alongside, same era.
Sound event localization and detection of overlapping sources using convolutional recurrent neural networks
Adavanne, S.; Politis, A.; Nikunen, J.; and Virtanen, T. 2018 · 2018
Cited alongside, same era.
Objects that sound
Arandjelovic, R.; and Zisserman, A. 2018 · 2018
Cited alongside, same era.
Deep clustering for unsupervised learning of visual features
Caron, M.; Bojanowski, P.; Joulin, A.; and Douze, M. 2018 · 2018
The sound of motions
Zhao, H.; Gan, C.; Ma, W.-C.; and Torralba, A. 2019 · 2019
Later among the works it cites.
Self-supervised learning of audio-visual objects from video
Afouras, T.; Owens, A.; Chung, J. S.; and Zisserman, A. 2020 · 2020
Later among the works it cites.
Music Gesture for Visual Sound Separation
Gan, C.; Huang, D.; Zhao, H.; Tenenbaum, J. B.; and Torralba, A. 2020 · 2020
Later among the works it cites.
Vggsound: a large-scale audio-visual dataset
Vedaldi, A.; Zisserman, A.; Chen, H.; and Xie, W. 2020 · 2020
Later among the works it cites.
Sep-Stereo: Visually Guided Stereophonic Audio Generation by Associating Source Separation
Zhou, H.; Xu, X.; Lin, D.; Wang, X.; and Liu, Z. 2020 · 2020
Later among the works it cites.
Localizing Visual Sounds the Hard Way
Chen, H.; Xie, W.; Afouras, T.; Nagrani, A.; Vedaldi, A.; and Zisserman, A. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
Korbar, B.; Tran, D.; and Torresani, L. 2018 · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Owens, A.; and Efros, A. A. 2018 · 2018
Cited alongside, same era.
Learning to localize sound source in visual scenes
Senocak, A.; Oh, T.-H.; Kim, J.; Yang, M.-H.; and So Kweon, I. 2018 · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos
Tian, Y.; Shi, J.; Li, B.; Duan, Z.; and Xu, C. 2018 · 2018
Cited alongside, same era.
The sound of pixels
Zhao, H.; Gan, C.; Rouditchenko, A.; Vondrick, C.; McDermott, J.; and Torralba, A. 2018 · 2018
Cited alongside, same era.
Co-separating sounds of visual objects
Gao, R.; and Grauman, K. 2019 · 2019
Cited alongside, same era.
Unsupervised Sound Localization via Iterative Contrastive Learning
Lin, Y.-B.; Tseng, H.-Y.; Lee, H.-Y.; Lin, Y.-Y.; and Yang, M.-H. 2021 · 2021
Later among the works it cites.
Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation
Tian, Y.; Hu, D.; and Xu, C. 2021 · 2021
Later among the works it cites.
Can audio-visual integration strengthen robustness under multimodal attacks?
Tian, Y.; and Xu, C. 2021 · 2021
Later among the works it cites.
Improving On-Screen Sound Separation for Open Domain Videos with Audio-Visual Self-attention
Tzinis, E.; Wisdom, S.; Remez, T.; and Hershey, J. R. 2021 · 2021
Later among the works it cites.
Visually Informed Binaural Audio Generation without Binaural Audios
Xu, X.; Zhou, H.; Liu, Z.; Dai, B.; Wang, X.; and Lin, D. 2021 · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Zhou, H.; Sun, Y.; Wu, W.; Loy, C. C.; Wang, X.; and Liu, Z. 2021 · 2021
Later among the works it cites.