Fetching the paper…
Reading the bibliography…
While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from the feet of a walker.
Smote: synthetic minority over-sampling technique
N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer · 2002
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Audio-visual speech recognition using deep learning
K. Noda, Y. Yamaguchi, K. Nakadai, H. G. Okuno, and T. Ogata · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Comparing modeled and measurement-based spherical harmonic encoding filters for spherical microphone arrays
A. Politis and H. Gamper · 2017
Earlier work this paper cites.
Sound event localization and detection of overlapping sources using convolutional recurrent neural networks
S. Adavanne, A. Politis, J. Nikunen, and T. Virtanen · 2018
Earlier work this paper cites.
Deep audio-visual speech recognition
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman · 2018
Earlier work this paper cites.
Learning to separate object sounds by watching unlabeled video
R. Gao, R. Feris, and K. Grauman · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
Earlier work this paper cites.
End-to-end audiovisual speech recognition
S. Petridis, T. Stafylakis, P. Ma, F. Cai, G. Tzimiropoulos, and M. Pantic · 2018
Earlier work this paper cites.
Learning to localize sound source in visual scenes
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon · 2018
Earlier work this paper cites.
The sound of pixels
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba · 2018
Earlier work this paper cites.
A multi-room reverberant dataset for sound event localization and detection
S. Adavanne, A. Politis, and T. Virtanen · 2019
Earlier work this paper cites.
Learning imbalanced datasets with label-distribution-aware margin loss
K. Cao, C. Wei, A. Gaidon, N. Arechiga, and T. Ma · 2019
Earlier work this paper cites.
Joint measurement of localization and detection of sound events
A. Mesaros, S. Adavanne, A. Politis, T. Heittola, and T. Virtanen · 2019
Earlier work this paper cites.
The sound of motions
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba · 2019
Earlier work this paper cites.
Avecl-umons database for audio-visual event classification and localization
M. Brousmiche, S. Dupont, and J. Rouat · 2020
Earlier work this paper cites.
Secl-umons database for sound event classification and localization
M. Brousmiche, J. Rouat, and S. Dupont · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Learning representations from audio-visual spatial alignment
P. Morgado, Y. Li, and N. Nvasconcelos · 2020
Cited alongside, same era.
A sequence matching network for polyphonic sound event localization and detection
T. N. T. Nguyen, D. L. Jones, and W.-S. Gan · 2020
Cited alongside, same era.
A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection
A. Politis, S. Adavanne, and T. Virtanen · 2020
Cited alongside, same era.
ACCDOA: Activity-coupled cartesian direction of arrival representation for sound event localization and detection
K. Shimada, Y. Koyama, N. Takahashi, S. Takahashi, and Y. Mitsufuji · 2021
Later among the works it cites.
Assessment of self-attention on learned features for sound event localization and detection
P. Sudarsanam, A. Politis, and K. Drossos · 2021
Later among the works it cites.
Q. Wang, J. Du, H.-X. Wu, J. Pan, F. Ma, and C.-H. Lee · 2021
Later among the works it cites.
Pano-AVQA: Grounded audio-visual question answering on 360deg videos
H. Yun, Y. Yu, W. Yang, K. Lee, and G. Kim · 2021
Later among the works it cites.
Soundspaces 2.0: A simulation platform for visual-acoustic learning
C. Chen, C. Schissler, S. Garg, P. Kobernik, A. Clegg, P. Calamia, D. Batra, P. Robinson, and K. Grauman · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Overview and evaluation of sound event localization and detection in DCASE 2019
A. Politis, A. Mesaros, S. Adavanne, T. Heittola, and T. Virtanen · 2020
Cited alongside, same era.
Sound event localization and detection based on crnn using rectangular filters and channel rotation data augmentation
F. Ronchini, D. Arteaga, and A. Pérez-López · 2020
Cited alongside, same era.
Telling left from right: Learning spatial correspondence of sight and sound
K. Yang, B. Russell, and J. Salamon · 2020
Cited alongside, same era.
Sound event localization based on sound intensity vector refined by DNN-based denoising and source separation
M. Yasuda, Y. Koizumi, S. Saito, H. Uematsu, and K. Imoto · 2020
Cited alongside, same era.
Deep triplet neural networks with cluster-cca for audio-visual cross-modal retrieval
D. Zeng, Y. Yu, and K. Oyama · 2020
Cited alongside, same era.
An improved event-independent network for polyphonic sound event localization and detection
Y. Cao, T. Iqbal, Q. Kong, F. An, W. Wang, and M. D. Plumbley · 2021
Cited alongside, same era.
Localizing visual sounds the hard way
H. Chen, W. Xie, T. Afouras, A. Nagrani, A. Vedaldi, and A. Zisserman · 2021
Cited alongside, same era.
Later among the works it cites.
Wearable seld dataset: Dataset for sound event localization and detection using wearable devices around head
K. Nagatomo, M. Yasuda, K. Yatabe, S. Saito, and Y. Oikawa · 2022
Later among the works it cites.
Salsa-lite: A fast and effective feature for polyphonic sound event localization and detection with microphone arrays
T. N. T. Nguyen, D. L. Jones, K. N. Watcharasupat, H. Phan, and W.-S. Gan · 2022
Later among the works it cites.
A. Politis, K. Shimada, P. Sudarsanam, S. Adavanne, D. Krause, Y. Koyama, N. Takahashi, S. Takahashi, Y. Mitsufuji, and T. Virtanen · 2022
Later among the works it cites.
Audio-visual cross-attention network for robotic speaker tracking
X. Qian, Z. Wang, J. Wang, G. Guan, and H. Li · 2022
Later among the works it cites.
On sorting and padding multiple targets for sound event localization and detection with permutation invariant and location-based training
R. Scheibler, T. Komatsu, Y. Fujita, and M. Hentschel · 2022
Later among the works it cites.
Multi-accdoa: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating permutation invariant training
K. Shimada, Y. Koyama, S. Takahashi, N. Takahashi, E. Tsunoo, and Y. Mitsufuji · 2022
Later among the works it cites.
Deep learning based audio-visual multi-speaker doa estimation using permutation-free loss function
Q. Wang, H. Chen, Y. Jiang, Z. Wang, Y. Wang, J. Du, and C.-H. Lee · 2022
Later among the works it cites.
Self-supervised learning of audio representations from audio-visual data using spatial alignment
S. Wang, A. Politis, A. Mesaros, and T. Virtanen · 2022
Later among the works it cites.
Audio–visual segmentation
J. Zhou, J. Wang, J. Zhang, W. Sun, J. Zhang, S. Birchfield, D. Guo, L. Kong, M. Wang, and Y. Zhong · 2022
Later among the works it cites.
Novel-view acoustic synthesis
C. Chen, A. Richard, R. Shapovalov, V. K. Ithapu, N. Neverova, K. Grauman, and A. Vedaldi · 2023
Closest in time.
Imagebind: One embedding space to bind them all
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra · 2023
Closest in time.
An experimental study on sound event localization and detection under realistic testing conditions
S. Niu, J. Du, Q. Wang, L. Chai, H. Wu, Z. Nian, L. Sun, Y. Fang, J. Pan, and C.-H. Lee · 2023
Closest in time.
Any-to-any generation via composable diffusion
Z. Tang, Z. Yang, C. Zhu, M. Zeng, and M. Bansal · 2023
Closest in time.
The nerc-slip system for sound event localization and detection of dcase2023 challenge
Q. Wang, Y. Jiang, S. Cheng, M. Hu, Z. Nian, P. Hu, Z. Liu, Y. Dong, M. Cai, J. Du, and C.-H. Lee · 2023
Closest in time.
Multi event localization by audio-visual fusion with omnidirectional camera and microphone array
W. Zheng, R. Yoshihashi, R. Kawakami, I. Sato, and A. Kanezaki · 2023
Closest in time.