Fetching the paper…
Reading the bibliography…
In the Sound Event Localization and Detection (SELD) task, Transformer-based models have demonstrated impressive capabilities.
I. Misra, A. Shrivastava, A. Gupta, and M. Hebert, “Cross-stitch networks for multi-task learning,” in Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 3994–4003
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Neural Information Processing Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
S. Adavanne, A. Politis, J. Nikunen, and T. Virtanen, “Sound event localization and detection of overlapping sources using convolutional recurrent neural networks,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 1, pp. 34–48, 2018
2018
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,” in Neural Information Processing Systems (NeurIPS) , 2019
2019
Earlier work this paper cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Proc. Interspeech , 2020, pp. 5036–5040
2020
Earlier work this paper cites.
Y. Cao, T. Iqbal, Q. Kong, F. An, W. Wang, and M. D. Plumbley, “An improved event-independent network for polyphonic sound event localization and detection,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 885–889
2021
Earlier work this paper cites.
J. Hu, Y. Cao, M. Wu, Q. Kong, F. Yang, M. D. Plumbley, and J. Yang, “A track-wise ensemble event independent network for polyphonic sound event localization and detection,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 9196–9200
2022
Earlier work this paper cites.
S. Niu, J. Du, Q. Wang, L. Chai, H. Wu, Z. Nian, L. Sun, Y. Fang, J. Pan, and C.-H. Lee, “An experimental study on sound event localization and detection under realistic testing conditions,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Shul and J.-W. Choi, “Cst-former: Transformer with channel-spectro-temporal attention for sound event localization and detection,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 8686–8690
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations (ICLR)
Cited in the paper.
2024
Closest in time.
D. A. Krause and A. Politis, “[DCASE2024 Task 3] Synthetic SELD mixtures for baseline training,” 2024. [Online]. Available: https://doi.org/10.5281/zenodo.10932241
2024
Closest in time.