Fetching the paper…
Reading the bibliography…
Sound events often occur in unstructured environments where they exhibit wide variations in their frequency content and temporal structure.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
W.-H. Cheng, W.-T. Chu, and J.-L. Wu, “Semantic context detection based on hierarchical audio models,” in Proceedings of the 5th ACM SIGMM international workshop on Multimedia information retrieval , 2003, pp. 109–115
2003
Earlier work this paper cites.
L.-H. Cai, L. Lu, A. Hanjalic, H.-J. Zhang, and L.-H. Cai, “A flexible framework for key audio effects detection and auditory context inference,” IEEE Trans. on Audio, Speech, and Language Processing , vol. 14, no. 3, pp. 1026–1039, 2006
2006
Earlier work this paper cites.
A. Mesaros, T. Heittola, A. Eronen, and T. Virtanen, “Acoustic event detection in real life recordings,” in Proc. European Signal Processing Conference (EUSIPCO) , 2010, pp. 1267–1271
2010
Earlier work this paper cites.
T. Heittola, A. Mesaros, A. Eronen, and T. Virtanen, “Audio context recognition using audio event histograms,” in Proc. of the 18th European Signal Processing Conference (EUSIPCO) , 2010, pp. 1272–1276
2010
Earlier work this paper cites.
G. Forman and M. Scholz, “Apples-to-apples in cross-validation studies: pitfalls in classifier performance measurement,” ACM SIGKDD Explorations Newsletter , vol. 12, no. 1, pp. 49–57, 2010
2010
Earlier work this paper cites.
S. Goetze, J. Schroder, S. Gerlach, D. Hollosi, J.-E. Appell, and F. Wallhoff, “Acoustic monitoring and localization for social care,” Journal of Computing Science and Engineering , vol. 6, no. 1, pp. 40–50, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
Y. Bengio et al. , “Deep learning of representations for unsupervised and transfer learning.” ICML Unsupervised and Transfer Learning , vol. 27, pp. 17–36, 2012
2012
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in 2013 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , 2013, pp. 6645–6649
2013
Earlier work this paper cites.
T. Heittola, A. Mesaros, A. Eronen, and T. Virtanen, “Context-dependent sound event detection,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2013, no. 1, p. 1, 2013
2013
Earlier work this paper cites.
T. Heittola, A. Mesaros, T. Virtanen, and M. Gabbouj, “Supervised model training for overlapping sound events based on unsupervised source separation,” in Int. Conf. on Acoustics, Speech, and Signal Processing (ICASSP) , 2013, pp. 8677–8681
2013
Earlier work this paper cites.
O. Dikmen and A. Mesaros, “Sound event detection using non-negative dictionaries learned from annotated overlapping events,” in 2013 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics , 2013, pp. 1–4
2013
Earlier work this paper cites.
I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. C. Courville, and Y. Bengio, “Maxout networks.” ICML (3) , vol. 28, pp. 1319–1327, 2013
2013
Earlier work this paper cites.
J. W. Dennis, “Sound event recognition in unstructured environments using spectrogram image processing,” Nanyang Technological University, Singapore , 2014
2014
Earlier work this paper cites.
O. Gencoglu, T. Virtanen, and H. Huttunen, “Recognition of acoustic events using deep neural networks,” in Proc. European Signal Processing Conference (EUSIPCO) , 2014
2014
Earlier work this paper cites.
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014
2014
Earlier work this paper cites.
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio, “On the properties of neural machine translation: Encoder-decoder approaches,” Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation (SSST-8) , 2014
2014
Cited alongside, same era.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting.” Journal of Machine Learning Research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Cited alongside, same era.
K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep inside convolutional networks: Visualising image classification models and saliency maps,” ICLR Workshop , 2014
2014
Cited alongside, same era.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in Advances in neural information processing systems , 2014, pp. 3320–3328
2014
Cited alongside, same era.
B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in Proceedings of the 14th Python in Science Conference , 2015
2015
Later among the works it cites.
Y. Wang, L. Neves, and F. Metze, “Audio-based multimedia event detection using deep recurrent neural networks,” in 2016 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 2742–2746
2016
Later among the works it cites.
H. Phan, L. Hertel, M. Maass, and A. Mertins, “Robust audio event recognition with 1-max pooling convolutional neural networks,” Interspeech , 2016
2016
Later among the works it cites.
G. Parascandolo, H. Huttunen, and T. Virtanen, “Recurrent neural networks for polyphonic sound event detection in real life recordings,” in 2016 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , 2016, pp. 6440–6444
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Foggia, N. Petkov, A. Saggese, N. Strisciuglio, and M. Vento, “Reliable detection of audio events in highly noisy environments,” Pattern Recognition Letters , vol. 65, pp. 22–28, 2015
2015
Cited alongside, same era.
J. Salamon and J. P. Bello, “Feature learning with deep scattering for urban sound analysis,” in 2015 23rd European Signal Processing Conference (EUSIPCO) . IEEE, 2015, pp. 724–728
2015
Cited alongside, same era.
D. Stowell and D. Clayton, “Acoustic event detection for multiple overlapping similar sources,” in 2015 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2015, pp. 1–5
2015
Cited alongside, same era.
H. Zhang, I. McLoughlin, and Y. Song, “Robust sound event recognition using convolutional neural networks,” in 2015 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 559–563
2015
Cited alongside, same era.
K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in Int. Workshop on Machine Learning for Signal Processing (MLSP) , 2015, pp. 1–6
2015
Cited alongside, same era.
A. Mesaros, O. Dikmen, T. Heittola, and T. Virtanen, “Sound event detection in real life recordings using coupled matrix factorization of spectral representations and class activity annotations,” in 2015 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 151–155
2015
Cited alongside, same era.
E. Cakir, T. Heittola, H. Huttunen, and T. Virtanen, “Polyphonic sound event detection using multilabel deep neural networks,” in Int. Joint Conf. on Neural Networks (IJCNN) , 2015, pp. 1–7
2015
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Cited alongside, same era.
E. Cakir, E. Ozan, and T. Virtanen, “Filterbank learning for deep neural network based polyphonic sound event detection,” in Int. Joint Conf. on Neural Networks (IJCNN) , 2016
2016
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Later among the works it cites.
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in Proceedings of The 33rd International Conference on Machine Learning , 2016, pp. 173–182
2016
Later among the works it cites.
2016
Later among the works it cites.
S. Sigtia, E. Benetos, and S. Dixon, “An end-to-end neural network for polyphonic piano music transcription,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 5, pp. 927–939, 2016
2016
Later among the works it cites.
Y. Gal, “A theoretically grounded application of dropout in recurrent neural networks,” Advances in neural information processing systems , 2016
2016
Later among the works it cites.
A. Mesaros, T. Heittola, and T. Virtanen, “TUT database for acoustic scene classification and sound event detection,” in 24th European Signal Processing Conference (EUSIPCO) , 2016
2016
Later among the works it cites.
A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for polyphonic sound event detection,” Applied Sciences , vol. 6, no. 6, p. 162, 2016
2016
Later among the works it cites.
T. Heittola, A. Mesaros, and T. Virtanen. (2016) DCASE2016 baseline system. [Online]. Available: https://github.com/TUT-ARG/DCASE2016-baseline-system-python
2016
Later among the works it cites.
F. Chollet. (2016) Keras. [Online]. Available: https://github.com/fchollet/keras
2016
Later among the works it cites.
2016
Later among the works it cites.
T. Lidy and A. Schindler, “CQT-based convolutional neural networks for audio scene classification and domestic audio tagging,” DCASE2016 Challenge, Tech. Rep., September 2016
2016
Later among the works it cites.
E. Cakir, T. Heittola, and T. Virtanen, “Domestic audio tagging with convolutional neural networks,” DCASE2016 Challenge, Tech. Rep., September 2016
2016
Later among the works it cites.
S. Yun, S. Kim, S. Moon, J. Cho, and T. Kim, “Discriminative training of GMM parameters for audio scene classification,” DCASE2016 Challenge, Tech. Rep., September 2016
2016
Later among the works it cites.