J. P. Woodard, “Modeling and classification of natural sounds by product code hidden Markov models,” IEEE Transactions on Signal Processing , vol. 40, pp. 1833–1835, 1992
1992
Earlier work this paper cites.
D. P. W. Ellis, “Detecting alarm sounds,” https://academiccommons.columbia.edu/doi/10.7916/D8F19821/download , 2001
2001
Earlier work this paper cites.
D. Li, I. K. Sethi, N. Dimitrova, and T. McGee, “Classification of general audio data for content-based retrieval,” Pattern Recognition Letters , vol. 22, pp. 533–544, 2001
2001
Earlier work this paper cites.
G. Tzanetakis and P. Cook, “Musical genre classification of audio signals,” IEEE Transactions on Speech and Audio Processing , vol. 10, pp. 293–302, 2002
2002
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2009, pp. 248–255
2009
Earlier work this paper cites.
E. Law and L. Von Ahn, “Input-agreement: a new mechanism for collecting data using human computation games,” in Proceedings of the SIGCHI Conference on Human Factors in Computing Systems , 2009, pp. 1197–1206
2009
Earlier work this paper cites.
A. Mesaros, T. Heittola, A. Eronen, and T. Virtanen, “Acoustic event detection in real life recordings,” in European Signal Processing Conference (EUSIPCO) , 2010, pp. 1267–1271
2010
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in International Conference on Machine Learning (ICML) , 2010, pp. 807–814
2010
Earlier work this paper cites.
B. Uzkent, B. D. Barkana, and H. Cevikalp, “Non-speech environmental sound classification using SVMs with a new set of features,” International Journal of Innovative Computing, Information and Control , vol. 8, pp. 3511–3524, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems (NeurIPS) , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
L. Vuegen, B. Broeck, P. Karsmakers, J. F. Gemmeke, B. Vanrumste, and H. Hamme, “An MFCC-GMM approach for event detection and classification,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2013
2013
Earlier work this paper cites.
A. Van Den Oord, S. Dieleman, and B. Schrauwen, “Transfer learning by supervised pre-training for audio-based music classification,” in Conference of the International Society for Music Information Retrieval (ISMIR) , 2014, pp. 29–34
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” The Journal of Machine Learning Research , vol. 15, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
M. Lin, Q. Chen, and S. Yan, “Network in network,” in International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in Proceedings of the ACM International Conference on Multimedia , 2014, pp. 1041–1044
2014
Earlier work this paper cites.
E. Cakir, T. Heittola, H. Huttunen, and T. Virtanen, “Polyphonic sound event detection using multi label deep neural networks,” in International Joint Conference on Neural Networks (IJCNN) , 2015
2015
Earlier work this paper cites.
D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, and M. D. Plumbley, “Detection and classification of acoustic scenes and events,” IEEE Transactions on Multimedia , vol. 17, pp. 1733–1746, 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Machine Learning (ICML) , 2015, pp. 448–456
2015
Earlier work this paper cites.
B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in Proceedings of the Python in Science Conference , vol. 8, 2015, pp. 18–25
2015
Earlier work this paper cites.