D. M. Green, J. A. Swets et al. , Signal Detection Theory and Psychophysics . Wiley New York, 1966, vol. 1
1966
Earlier work this paper cites.
M. F. Porter et al. , “An algorithm for suffix stripping.” Program , vol. 14, no. 3, pp. 130–137, 1980
1980
Earlier work this paper cites.
P. Toth and S. Martello, Knapsack Problems: Algorithms and Computer Implementations . Wiley, 1990
1990
Earlier work this paper cites.
B. Whitman, G. Flake, and S. Lawrence, “Artist detection in music with minnowmatch,” in Neural Networks for Signal Processing XI: Proceedings of the 2001 IEEE Signal Processing Society Workshop (IEEE Cat. No. 01TH8584) . IEEE, 2001, pp. 559–568
2001
Earlier work this paper cites.
H. Liu and H. Motoda, “On issues of instance selection,” Data Mining and Knowledge Discovery , vol. 6, no. 2, p. 115, 2002
2002
Earlier work this paper cites.
M. I. Mandel and D. P. W. Ellis, “Song-level features and support vector machines for music classification,” in Proceedings of the International Society for Music Information Retrieval Conference (ISMIR 2005) , 2005
2005
Earlier work this paper cites.
L. Malfait, J. Berger, and M. Kastner, “P.563—The ITU-T standard for single-ended speech quality assessment,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 6, pp. 1924–1934, 2006
2006
Earlier work this paper cites.
J. Davis and M. Goadrich, “The relationship between Precision-Recall and ROC curves,” in Proceedings of the 23rd International Conference on Machine Learning , 2006, pp. 233–240
2006
Earlier work this paper cites.
D. A. Van Leeuwen and N. Brümmer, “An introduction to application-independent evaluation of speaker recognition systems,” in Speaker classification I . Springer, 2007, pp. 330–353
2007
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” University of Toronto, Tech. Rep., 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Flexer and D. Schnitzer, “Album and artist effects for audio similarity at the scale of the web,” Children , vol. 15, no. 15.95, pp. 4–07, 2009
2009
Earlier work this paper cites.
A. Mesaros, T. Heittola, A. Eronen, and T. Virtanen, “Acoustic event detection in real life recordings,” in 2010 18th European Signal Processing Conference . IEEE, 2010, pp. 1267–1271
2010
Earlier work this paper cites.
C. V. Cotton and D. P. W. Ellis, “Spectral vs. spectro-temporal features for acoustic event detection,” in 2011 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2011, pp. 69–72
2011
Earlier work this paper cites.
K. Sechidis, G. Tsoumakas, and I. Vlahavas, “On the stratification of multi-label data,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 2011, pp. 145–158
2011
Earlier work this paper cites.
H. Lei, J. Choi, A. Janin, and G. Friedland, “User verification: Matching the uploaders of videos across accounts,” in 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2011, pp. 2404–2407
2011
Earlier work this paper cites.
S. G. Jantzen, B. J. Sutherland, D. R. Minkley, and B. F. Koop, “Go trimming: Systematically reducing redundancy in large gene ontology datasets,” BMC research notes , vol. 4, no. 1, p. 267, 2011
2011
Earlier work this paper cites.
T. Drugman, J. Urbain, N. Bauwens, R. Chessini, A.-S. Aubriot, P. Lebecque, and T. Dutoit, “Audio and contact microphones for cough detection,” in Thirteenth Annual Conference of the International Speech Communication Association , 2012
2012
Earlier work this paper cites.
F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in Proceedings of the 21st ACM International Conference on Multimedia . ACM, 2013, pp. 411–412
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in European Conference on Computer Vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in Proceedings of the ACM International Conference on Multimedia . ACM, 2014, pp. 1041–1044
2014
Earlier work this paper cites.
D. Stowell and M. Plumbley, “An open dataset for research on audio field recording archives: freefield1010,” in Audio Engineering Society Conference: 53rd International Conference: Semantic Audio . Audio Engineering Society, 2014
2014
Earlier work this paper cites.
EBU Recommendation R 128, “Loudness Normalisation and Permitted Maximum Level of Audio Signals,” European Broadcasting Union , 2014
2014
Earlier work this paper cites.
M. Sabou, K. Bontcheva, L. Derczynski, and A. Scharl, “Corpus annotation through crowdsourcing: Towards best practice guidelines.” in LREC , 2014, pp. 859–866
2014
Earlier work this paper cites.
P. Foggia, N. Petkov, A. Saggese, N. Strisciuglio, and M. Vento, “Reliable detection of audio events in highly noisy environments,” Pattern Recognition Letters , vol. 65, pp. 22–28, 2015
2015
Earlier work this paper cites.
E. Cakir, T. Heittola, H. Huttunen, and T. Virtanen, “Polyphonic sound event detection using multi label deep neural networks,” in 2015 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2015, pp. 1–7
2015
Earlier work this paper cites.
K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in 2015 IEEE 25th International Workshop on Machine Learning for Signal Processing (MLSP) . IEEE, 2015, pp. 1–6
2015
Earlier work this paper cites.
P. Foster, S. Sigtia, S. Krstulovic, J. Barker, and M. D. Plumbley, “CHiME-home: A dataset for sound source recognition in a domestic environment,” in 2015 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2015
2015
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proceedings of the ACM International Conference on Multimedia . ACM, 2015, pp. 1015–1018
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “ImageNet large scale visual recognition challenge,” International Journal of Computer Vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
Martín Abadi et al., “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: https://www.tensorflow.org/
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in 3rd International Conference on Learning Representations, ICLR , 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning , 2015, pp. 448–456
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in 3rd International Conference on Learning Representations, ICLR , 2015
2015
Earlier work this paper cites.
T. N. Sainath, O. Vinyals, A. Senior, and H. Sak, “Convolutional, long short-term memory, fully connected deep neural networks,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 4580–4584
2015
Earlier work this paper cites.
Y. Wang, L. Neves, and F. Metze, “Audio-based multimedia event detection using deep recurrent neural networks,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 2742–2746
2016
Earlier work this paper cites.
M. Crocco, M. Cristani, A. Trucco, and V. Murino, “Audio surveillance: A systematic review,” ACM Computing Surveys (CSUR) , vol. 48, no. 4, pp. 1–46, 2016
2016
Earlier work this paper cites.
G. Parascandolo, H. Huttunen, and T. Virtanen, “Recurrent neural networks for polyphonic sound event detection in real life recordings,” in 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2016, pp. 6440–6444
2016
Earlier work this paper cites.
A. Mesaros, T. Heittola, and T. Virtanen, “TUT database for acoustic scene classification and sound event detection,” in 24th European Signal Processing Conference 2016 (EUSIPCO 2016) , Budapest, Hungary, 2016
2016
Earlier work this paper cites.
W. Han, E. Coutinho, H. Ruan, H. Li, B. Schuller, X. Yu, and X. Zhu, “Semi-supervised active learning for sound classification in hybrid learning environments,” PloS one , vol. 11, no. 9, p. e0162075, 2016
2016
Earlier work this paper cites.
E. Cakir, E. C. Ozan, and T. Virtanen, “Filterbank learning for deep neural network based polyphonic sound event detection,” in 2016 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2016, pp. 3399–3406
2016
Earlier work this paper cites.
D. Bogdanov, A. Porter, H. Boyer, X. Serra et al. , “Cross-collection evaluation for music classification tasks,” in Proceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR) . ISMIR, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.