Fetching the paper…
Reading the bibliography…
In this paper, we present SpecAugment++, a novel data augmentation method for deep neural networks based acoustic scene classification (ASC).
H. Zhang, I. McLoughlin, and Y. Song, “Robust sound event recognition using convolutional neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 559–563
2015
Earlier work this paper cites.
K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in IEEE 25th International Workshop on Machine Learning for Signal Processing (MLSP) . IEEE, 2015, pp. 1–6
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Q. Kong, I. Sobieraj, W. Wang, and M. Plumbley, “Deep neural network baseline for DCASE challenge 2016,” Proceedings of Detection and Classification of Acoustic Scenes and Events (DCASE) , 2016
2016
Earlier work this paper cites.
M. Valenti, A. Diment, G. Parascandolo, S. Squartini, and T. Virtanen, “DCASE 2016 acoustic scene classification using convolutional neural networks,” in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2016, pp. 95–99
2016
Earlier work this paper cites.
G. Parascandolo, H. Huttunen, and T. Virtanen, “Recurrent neural networks for polyphonic sound event detection in real life recordings,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 6440–6444
2016
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “SoundNet: learning sound representations from unlabeled video,” in Proceedings of the 30th International Conference on Neural Information Processing Systems , 2016, pp. 892–900
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
E. Cakır, G. Parascandolo, T. Heittola, H. Huttunen, and T. Virtanen, “Convolutional recurrent neural networks for polyphonic sound event detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 6, pp. 1291–1303, 2017
2017
Earlier work this paper cites.
Y. Tokozume and T. Harada, “Learning environmental sounds with end-to-end convolutional neural network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 2721–2725
2017
Cited alongside, same era.
J. Salamon and J. P. Bello, “Deep convolutional neural networks and data augmentation for environmental sound classification,” IEEE Signal Processing Letters , vol. 24, no. 3, pp. 279–283, 2017
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Tokozume, Y. Ushiku, and T. Harada, “Learning from between-class examples for deep sound recognition,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
A. Mesaros, T. Heittola, and T. Virtanen, “Acoustic scene classification in DCASE 2019 challenge: Closed and open set classification and data mismatch setups,” in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2019, pp. 164–168
2019
Later among the works it cites.
K. Koutini, H. Eghbal-Zadeh, M. Dorfer, and G. Widmer, “The receptive field as a regularizer in deep convolutional neural networks for acoustic scene classification,” in 27th European Signal Processing Conference (EUSIPCO) . IEEE, 2019, pp. 1–5
2019
Later among the works it cites.
H. Wang, Y. Zou, and D. Chong, “Acoustic scene classification with spectrogram processing strategies,” in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2020, pp. 210–214
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Mesaros, T. Heittola, and T. Virtanen, “A multi-device dataset for urban acoustic scene classification,” in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2018, pp. 9–13
2018
Cited alongside, same era.
Y. Sakashita and M. Aono, “Acoustic scene classification by ensemble of spectrograms based on adaptive temporal divisions,” in Detection and Classification of Acoustic Scenes and Events (DCASE) Challenge , 2018
2018
Cited alongside, same era.
A. Mesaros, A. Diment, B. Elizalde, T. Heittola, E. Vincent, B. Raj, and T. Virtanen, “Sound event detection in the DCASE 2017 challenge,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 6, pp. 992–1006, 2019
2019
Cited alongside, same era.
D. Zou and Q. Gu, “An improved analysis of training over-parameterized deep neural networks,” Advances in Neural Information Processing Systems , 2019
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” Proc. Interspeech , pp. 2613–2617, 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
H. Wang, Y. Zou, D. Chong, and W. Wang, “Modeling label dependencies for audio tagging with graph convolutional network,” IEEE Signal Processing Letters , vol. 27, pp. 1560–1564, 2020
2020
Later among the works it cites.
A. Jindal, N. E. Ranganatha, A. Didolkar, A. G. Chowdhury, D. Jin, R. Sawhney, and R. R. Shah, “SpeechMix—augmenting deep sound recognition using hidden space interpolations,” Proc. Interspeech , pp. 861–865, 2020
2020
Later among the works it cites.
H. Wang, Y. Zou, D. Chong, and W. Wang, “Environmental sound classification with parallel temporal-spectral attention,” in Proc. Interspeech , 2020, pp. 821–825
2020
Later among the works it cites.
K. Koutini, F. Henkel, H. Eghbal-Zadeh, and G. Widmer, “Low-complexity models for acoustic scene classification based on receptive field regularization and frequency damping,” in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2020, pp. 86–90
2020
Later among the works it cites.
H. Wang, Y. Zou, and W. Wang, “A global-local attention framework for weakly labelled audio tagging,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 351–355
2021
Closest in time.