Fetching the paper…
Reading the bibliography…
Sound event detection (SED) is a task to detect sound events in an audio recording.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in Proceedings of International Conference on Machine Learning (ICML) , 2010, pp. 807–814
2010
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur, “Recurrent neural network based language model,” in Conference on International Speech Communication Association (ISCA) , 2010, pp. 1724–1734
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NIPS) , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
D. Giannoulis, E. Benetos, D. Stowell, M. Rossignol, M. Lagrange, and M. D. Plumbley, “Detection and classification of acoustic scenes and events: An IEEE AASP challenge,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2013
2013
Earlier work this paper cites.
G. E. Dahl, T. N. Sainath, and G. E. Hinton, “Improving deep neural networks for lvcsr using rectified linear units and dropout,” in ICASSP , 2013, pp. 8609–8613
2013
Earlier work this paper cites.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 22, no. 10, pp. 1533–1545, 2014
2014
Earlier work this paper cites.
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014
2014
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” International Conference for Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, and M. D. Plumbley, “Detection and classification of acoustic scenes and events,” IEEE Transactions on Multimedia , vol. 17, no. 10, pp. 1733–1746, 2015
2015
Earlier work this paper cites.
E. Cakir, T. Heittola, H. Huttunen, and T. Virtanen, “Polyphonic sound event detection using multi label deep neural networks,” in International Joint Conference on Neural Networks (IJCNN) , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Machine Learning (ICML) , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Mesaros, T. Heittola, and T. Virtanen, “TUT database for acoustic scene classification and sound event detection,” in IEEE European Signal Processing Conference (EUSIPCO) , 2016, pp. 1128–1132
2016
Earlier work this paper cites.
E. Cakır, T. Heittola, and T. Virtanen, “Domestic audio tagging with convolutional neural networks,” Tech. Rep., 2016. [Online]. Available: http://dcase.community/challenge2016
2016
Earlier work this paper cites.
A. Kumar and B. Raj, “Audio event detection using weakly labeled data,” in Proceedings of ACM on Multimedia Conference , 2016, pp. 1038–1047
2016
Earlier work this paper cites.
J. L. Dai Wei, P. Pham, S. Das, S. Qu, and F. Metze, “Sound event detection for real life audio DCASE challenge,” DCASE2016 Challenge, Tech. Rep., 2016. [Online]. Available: http://dcase.community/challenge2016
2016
Earlier work this paper cites.
H. Phan, L. Hertel, M. Maass, and A. Mertins, “Robust audio event recognition with 1-max pooling convolutional neural networks,” in INTERSPEECH , 2016, pp. 3653–3657
2016
Earlier work this paper cites.
K. Choi, G. Fazekas, and M. Sandler, “Automatic tagging using deep convolutional neural networks,” International Society of Music Information Retrieval (ISMIR) , 2016
2016
Cited alongside, same era.
A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for polyphonic sound event detection,” Applied Sciences , vol. 6, no. 6, p. 162, 2016
2016
Cited alongside, same era.
A. Mesaros, T. Heittola, A. Diment, B. Elizalde, A. Shah, E. Vincent, B. Raj, and T. Virtanen, “DCASE 2017 challenge setup: Tasks, datasets and baseline system,” in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2017, pp. 85–92
2017
Cited alongside, same era.
“DCASE 2017 Task 4,” http://www.cs.tut.fi/sgn/arg/dcase2017/challenge/task-large-scale-sound-event-detection
2017
Cited alongside, same era.
E. Cakir, G. Parascandolo, T. Heittola, H. Huttunen, and T. Virtanen, “Convolutional recurrent neural networks for polyphonic sound event detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 6, pp. 1291–1303, 2017
2017
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NIPS) , 2017, pp. 5998–6008
2017
Later among the works it cites.
J. Thickstun, Z. Harchaoui, and S. Kakade, “Learning features of music from scratch,” in International Conference on Learning Representations (ICLR) , 2017
2017
Later among the works it cites.
A. Mesaros, T. Heittola, and T. Virtanen, “A multi-device dataset for urban acoustic scene classification,” Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 776–780
2017
Cited alongside, same era.
Y. Xu, Q. Kong, Q. Huang, W. Wang, and M. D. Plumbley, “Convolutional gated recurrent neural network incorporating spatial features for audio tagging,” in International Joint Conference on Neural Networks (IJCNN) , 2017, pp. 3461–3466
2017
Cited alongside, same era.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “CNN architectures for large-scale audio classification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 131–135
2017
Cited alongside, same era.
T.-W. Su, J.-Y. Liu, and Y.-H. Yang, “Weakly-supervised audio event detection using event-specific gaussian filters and fully convolutional networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 791–795
2017
Cited alongside, same era.
J. Salamon, B. McFee, and P. Li, “DCASE 2017 submission: Multiple instance learning for sound event detection,” DCASE2017 Challenge, Tech. Rep., 2017. [Online]. Available: http://dcase.community/challenge2017
2017
Cited alongside, same era.
K. Lee, D. Lee, S. Lee, and Y. Han, “Ensemble of convolutional neural networks for weakly-supervised sound event detection using multiple scale input,” DCASE2017 Challenge, Tech. Rep., September 2017. [Online]. Available: http://dcase.community/challenge2017
2017
Cited alongside, same era.
S. Chou, J. Jang, and Y.-H. Yang, “FrameCNN: a weakly-supervised learning framework for frame-wise acoustic event detection and classification,” DCASE2017 Challenge, Tech. Rep., 2017. [Online]. Available: http://dcase.community/challenge2017
2017
Cited alongside, same era.
S. Adavanne, P. Pertilä, and T. Virtanen, “Sound event detection using spatial features and convolutional recurrent neural network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 771–775
2017
Cited alongside, same era.
Y. Xu, Q. Kong, W. Wang, and M. D. Plumbley, “Large-scale weakly supervised audio classification using gated convolutional neural network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 121–125
2018
Later among the works it cites.
S.-Y. Chou, J.-S. R. Jang, and Y.-H. Yang, “Learning to recognize transient sound events using attentional supervision.” in International Joint Conference on Artificial Intelligence (IJCAI) , 2018, pp. 3336–3342
2018
Later among the works it cites.
S. Gururani, C. Summers, and A. Lerch, “Instrument activity detection in polyphonic music using deep neural networks.” in International Society for Music Information Retrieval (ISMIR) , 2018, pp. 569–576
2018
Later among the works it cites.
Q. Kong, T. Iqbal, Y. Xu, W. Wang, and M. D. Plumbley, “DCASE 2018 challenge baseline with convolutional neural networks,” Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE) , pp. 217–221, 2018
2018
Later among the works it cites.
R. Serizel, N. Turpault, H. Eghbal-Zadeh, and A. P. Shah, “Large-scale weakly labeled semi-supervised sound event detection in domestic environments,” in Detection Classification Acoust. Scenes Events Workshop (DCASE) , 2018, pp. 19–23
2018
Later among the works it cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” International Conference on Learning Representations (ICLR) (ICLR) , 2018
2018
Later among the works it cites.
L. Ford, H. Tang, F. Grondin, and J. Glass, “A deep residual network for large-scale acoustic scene analysis,” in INTERSPEECH , 2019, pp. 2568–2572
2019
Closest in time.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics (NAACL) , 2019, pp. 4171–4186
2019
Closest in time.
2019
Closest in time.
L. Cances, P. Guyot, and T. Pellegrini, “Evaluation of post-processing algorithms for polyphonic sound event detection,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2019, pp. 318–322
2019
Closest in time.
Q. Kong, C. Yu, Y. Xu, T. Iqbal, W. Wang, and M. D. Plumbley, “Weakly labelled audioset tagging with attention neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2019
2019
Closest in time.
N. Turpault, R. Serizel, A. P. Shah, and J. Salamon, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” in Detection and Classification of Acoustic Scenes and Events (DCASE) Workshop , 2019, pp. 253–257
2019
Closest in time.