Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel four-stage data augmentation approach to ResNet-Conformer based acoustic modeling for sound event localization and detection (SELD).
C. Knapp and G. Carter, “The generalized correlation method for estimation of time delay,”
1976
Earlier work this paper cites.
R. Schmidt, “Multiple emitter location and signal parameter estimation,”
1986
Earlier work this paper cites.
R. Roy and T. Kailath, “ESPRIT-estimation of signal parameters via rotational invariance techniques,”
1989
Earlier work this paper cites.
H. Wang and P. Chu, “Voice source localization for automatic camera pointing system in videoconferencing,” in
1997
Earlier work this paper cites.
M. S. Brandstein and H. F. Silverman, “A robust method for speech signal time-delay estimation in reverberant rooms,” in
1997
Earlier work this paper cites.
J. Daniel, “Représentation de champs acoustiques, application à la transmission et à la reproduction de scènes sonores complexes dans un contexte multimédia,” Ph.D. dissertation, Univ. of Paris VI, France, 2000. [Online]. Available:
2000
Earlier work this paper cites.
P. Y. Simard, D. Steinkraus, J. C. Platt
2003
Earlier work this paper cites.
H. Teutsch and W. Kellermann, “Acoustic source detection and localization based on wavefield decomposition using circular microphone arrays,”
2006
Earlier work this paper cites.
G. Valenzise, L. Gerosa, M. Tagliasacchi, F. Antonacci, and A. Sarti, “Scream and gunshot detection and localization for audio-surveillance systems,” in
2007
Earlier work this paper cites.
H. Do, H. F. Silverman, and Y. Yu, “A real-time SRP-PHAT source location implementation using stochastic region contraction (SRC) on a large-aperture microphone array,” in
2007
Earlier work this paper cites.
E. Warsitz and R. Haeb-Umbach, “Blind acoustic beamforming based on generalized eigenvalue decomposition,”
2007
Earlier work this paper cites.
T. Heittola, A. Mesaros, A. Eronen, and T. Virtanen, “Audio context recognition using audio event histograms,” in
2010
Earlier work this paper cites.
A. Mesaros, T. Heittola, A. Eronen, and T. Virtanen, “Acoustic event detection in real life recordings,” in
2010
Earlier work this paper cites.
N. Q. Duong, E. Vincent, and R. Gribonval, “Under-determined reverberant audio source separation using a full-rank spatial covariance model,”
2010
Earlier work this paper cites.
T. Heittola, A. Mesaros, A. Eronen, and T. Virtanen, “Context-dependent sound event detection,”
2013
Earlier work this paper cites.
J. F. Gemmeke, L. Vuegen, P. Karsmakers, B. Vanrumste
2013
Earlier work this paper cites.
mh acoustics, “EM32 Eigenmike microphone array release notes (v17. 0),” Oct. 2013. [Online]. Available:
2013
Earlier work this paper cites.
P. Swietojanski, A. Ghoshal, and S. Renals, “Convolutional neural networks for distant speech recognition,”
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
P. Foggia, N. Petkov, A. Saggese, N. Strisciuglio, and M. Vento, “Audio surveillance of roads: A system for detecting anomalous sounds,”
2015
Earlier work this paper cites.
A. Mesaros, T. Heittola, O. Dikmen, and T. Virtanen, “Sound event detection in real life recordings using coupled matrix factorization of spectral representations and class activity annotations,” in
2015
Earlier work this paper cites.
I. McLoughlin, H. Zhang, Z. Xie, Y. Song, and W. Xiao, “Robust sound event classification using deep neural networks,”
2015
Earlier work this paper cites.
K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in
2015
Earlier work this paper cites.
H. Zhang, I. McLoughlin, and Y. Song, “Robust sound event recognition using convolutional neural networks,” in
2015
Earlier work this paper cites.
D. Pavlidi, S. Delikaris-Manias, V. Pulkki, and A. Mouchtaris, “3D localization of multiple sound sources with intensity vector estimates in single source zones,” in
2015
Earlier work this paper cites.
X. Cui, V. Goel, and B. Kingsbury, “Data augmentation for deep neural network acoustic modeling,”
2015
Earlier work this paper cites.
E. Kurz, F. Pfahler, and M. Frank, “Comparison of first-order Ambisonics microphone arrays,” in
2015
Earlier work this paper cites.
H. Phan, L. Hertel, M. Maass, and A. Mertins, “Robust audio event recognition with 1-max pooling convolutional neural networks,” in
2016
Cited alongside, same era.
Y. Wang, L. Neves, and F. Metze, “Audio-based multimedia event detection using deep recurrent neural networks,” in
2016
Cited alongside, same era.
G. Parascandolo, H. Huttunen, and T. Virtanen, “Recurrent neural networks for polyphonic sound event detection in real life recordings,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
T. Higuchi, N. Ito, T. Yoshioka, and T. Nakatani, “Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise,” in
2016
L. Mazzon, Y. Koizumi, M. Yasuda, and N. Harada, “First order ambisonics domain spatial augmentation for DNN-based direction of arrival estimation,” in
2019
Later among the works it cites.
2019
Later among the works it cites.
J. Zhang, W. Ding, and L. He, “Data augmentation and prior knowledge-based regularization for sound event localization and detection,” DCASE2019 Challenge, Tech. Rep., June 2019
2019
Later among the works it cites.
S. Adavanne, A. Politis, and T. Virtanen, “A multi-room reverberant dataset for sound event localization and detection,” in
2019
Later among the works it cites.
L. Mazzon, M. Yasuda, Y. Koizumi, and N. Harada, “Sound event localization and detection using foa domain spatial augmentation,” DCASE2019 Challenge, Tech. Rep., June 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
T. Hayashi, S. Watanabe, T. Toda, T. Hori, J. Le Roux, and K. Takeda, “Duration-controlled LSTM for polyphonic sound event detection,”
2017
Cited alongside, same era.
S. Sabour, N. Frosst, and G. E. Hinton, “Dynamic routing between capsules,” in
2017
Cited alongside, same era.
E. Cakır, G. Parascandolo, T. Heittola, H. Huttunen, and T. Virtanen, “Convolutional recurrent neural networks for polyphonic sound event detection,”
2017
Cited alongside, same era.
S. Hafezi, A. H. Moore, and P. A. Naylor, “Augmented intensity vectors for direction of arrival estimation in the spherical harmonic domain,”
2017
Cited alongside, same era.
J. Salamon and J. P. Bello, “Deep convolutional neural networks and data augmentation for environmental sound classification,”
2017
Cited alongside, same era.
L. Perez and J. Wang, “The effectiveness of data augmentation in image classification using deep learning,”
2017
Cited alongside, same era.
2019
Later among the works it cites.
Y. Cao, T. Iqbal, Q. Kong, M. Galindo, W. Wang, and M. Plumbley, “Two-stage sound event localization and detection using intensity vector and generalized cross-correlation,” DCASE2019 Challenge, Tech. Rep., June 2019
2019
Later among the works it cites.
A. Mesaros, S. Adavanne, A. Politis, T. Heittola, and T. Virtanen, “Joint measurement of localization and detection of sound events,” in
2019
Later among the works it cites.
M. Yasuda, Y. Koizumi, S. Saito, H. Uematsu, and K. Imoto, “Sound event localization based on sound intensity vector refined by DNN-based denoising and source separation,” in
2020
Later among the works it cites.
A. Politis, S. Adavanne, and T. Virtanen, “A dataset of reverberant spatial sound scenes with moving sources for sound event localization and detection,” in
2020
Later among the works it cites.
M. Brousmiche, J. Rouat, and S. Dupont, “SECL-UMons database for sound event classification and localization,” in
2020
Later among the works it cites.
C. Evers, H. W. Löllmann, H. Mellmann, A. Schmidt, H. Barfuss, P. A. Naylor, and W. Kellermann, “The LOCATA Challenge: Acoustic source localization and tracking,”
2020
Later among the works it cites.
2020
Later among the works it cites.
Q. Wang, H. Wu, Z. Jing, F. Ma, Y. Fang, Y. Wang, T. Chen, J. Pan, J. Du, and C.-H. Lee, “The USTC-iFlytek system for sound event localization and detection of DCASE2020 challenge,” DCASE2020 Challenge, Tech. Rep., July 2020
2020
Later among the works it cites.
DCASE2020, “Sound event localization and detection challenge results,” DCASE2020 Challenge, Tech. Rep., July 2020. [Online]. Available:
2020
Later among the works it cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu
2020
Later among the works it cites.
2020
Later among the works it cites.
K. Miyazaki, T. Komatsu, T. Hayashi, S. Watanabe, T. Toda, and K. Takeda, “Convolution-augmented transformer for semi-supervised sound event detection,” DCASE2020 Challenge, Tech. Rep., June 2020
2020
Later among the works it cites.
P. Pertilä, E. Cakir, A. Hakala, E. Fagerlund, T. Virtanen, A. Politis, and A. Eronen, “Mobile microphone array speech detection and localization in diverse everyday environments,” in
2021
Closest in time.
A. Politis, S. Adavanne, D. Krause, A. Deleforge, P. Srivastava, and T. Virtanen, “A dataset of dynamic reverberant sound scenes with directional interferers for sound event localization and detection,” in
2021
Closest in time.
W. He, P. Motlicek, and J.-M. Odobez, “Neural network adaptation and data augmentation for multi-speaker direction-of-arrival estimation,”
2021
Closest in time.
E. Guizzo, R. F. Gramaccioni, S. Jamili, C. Marinoni, E. Massaro, C. Medaglia, G. Nachira, L. Nucciarelli, L. Paglialunga, M. Pennese
2021
Closest in time.
Q. Wang, H. Wu, Z. Jing, F. Ma, Y. Fang, Y. Wang, T. Chen, J. Pan, J. Du, and C.-H. Lee, “A model ensemble approach for sound event localization and detection,” in
2021
Closest in time.
K. Nagatomo, M. Yasuda, K. Yatabe, S. Saito, and Y. Oikawa, “Wearable SELD dataset: Dataset for sound event localization and detection using wearable devices around head,” in
2022
Closest in time.
E. Guizzo, C. Marinoni, M. Pennese, X. Ren, X. Zheng, C. Zhang, B. Masiero, A. Uncini, and D. Comminiello, “L3DAS22 Challenge: Learning 3D audio sources in a real office environment,” in
2022
Closest in time.
Q. Wang, L. Chai, H. Wu, Z. Nian, S. Niu, S. Zheng, Y. Wang, L. Sun, Y. Fang, J. Pan, J. Du, and C.-H. Lee, “The NERC-SLIP system for sound event localization and detection of DCASE2022 challenge,” DCASE2022 Challenge, Tech. Rep., June 2022
2022
Closest in time.
DCASE2022, “Sound event localization and detection challenge results,” DCASE2022 Challenge, Tech. Rep., June 2022. [Online]. Available:
2022
Closest in time.