Fetching the paper…
Reading the bibliography…
Environmental Sound Classification (ESC) is an active research area in the audio domain and has seen a lot of progress in the past years.
1912
Earlier work this paper cites.
J. Volkmann, S. S. Stevens, and E. B. Newman, “A scale for the measurement of the psychological magnitude pitch,” The Journal of the Acoustical Society of America , vol. 8, no. 3, pp. 208–208, 1937. [Online]. Available: https://doi.org/10.1121/1.1901999
1937
Earlier work this paper cites.
R. N. Shepard, “Circularity in judgments of relative pitch,” The Journal of the Acoustical Society of America , vol. 36, no. 12, pp. 2346–2353, 1964. [Online]. Available: https://doi.org/10.1121/1.1919362
1964
Earlier work this paper cites.
J. W. Cooley and J. W. Tukey, “An algorithm for the machine calculation of complex fourier series,” Mathematics of computation , vol. 19, no. 90, pp. 297–301, 1965
1965
Earlier work this paper cites.
J. Allen, “Short term spectral analysis, synthesis, and modification by discrete fourier transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 25, no. 3, pp. 235–238, June 1977
1977
Earlier work this paper cites.
F. J. Harris, “On the use of windows for harmonic analysis with the discrete fourier transform,” Proceedings of the IEEE , vol. 66, no. 1, pp. 51–83, 1978
1978
Earlier work this paper cites.
J. F. Kaiser, “Some useful properties of teager’s energy operators,” in 1993 IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 3, 1993, pp. 149–152 vol.3
1993
Earlier work this paper cites.
M. Slaney et al. , “An efficient implementation of the patterson-holdsworth auditory filter bank,” Apple Computer, Perception Group, Tech. Rep , vol. 35, no. 8, 1993
1993
Earlier work this paper cites.
B. Logan et al. , “Mel frequency cepstral coefficients for music modeling.” in Proceeding of the International Symposium on Music Information Retrieval (ISMIR) , Plymouth, USA, October 2000
2000
Earlier work this paper cites.
N. Marwan, N. Wessel, U. Meyerfeldt, A. Schirdewan, and J. Kurths, “Recurrence-plot-based measures of complexity and their application to heart-rate-variability data,” Physical review E , vol. 66, no. 2, p. 026702, 2002
2002
Earlier work this paper cites.
D.-N. Jiang, L. Lu, H.-J. Zhang, J.-H. Tao, and L.-H. Cai, “Music type classification by spectral contrast feature,” in Proceedings. IEEE International Conference on Multimedia and Expo , vol. 1. IEEE, 2002, pp. 113–116
2002
Earlier work this paper cites.
G. Heinzel, A. Rüdiger, and R. Schilling, “Spectrum and spectral density estimation by the discrete fourier transform (dft), including a comprehensive list of window functions and some new flat-top windows,” Max-Planck-Institut für Gravitationsphysik, Tech. Rep., 2002
2002
Earlier work this paper cites.
M. Raimbault and D. Dubois, “Urban soundscapes: Experiences and knowledge,” Cities , vol. 22, no. 5, pp. 339–350, 2005
2005
Earlier work this paper cites.
C. Harte, M. Sandler, and M. Gasser, “Detecting harmonic change in musical audio,” in Proceedings of the 1st ACM workshop on Audio and music computing multimedia , 2006, pp. 21–26
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR09 , 2009
2009
Cited alongside, same era.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research , vol. 12, pp. 2825–2830, 2011
2011
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Cited alongside, same era.
F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in ACM International Conference on Multimedia (MM’13) , ACM. Barcelona, Spain: ACM, 21/10/2013 2013, pp. 411–412
2013
Cited alongside, same era.
J. Salamon and J. P. Bello, “Deep convolutional neural networks and data augmentation for environmental sound classification,” IEEE Signal Processing Letters , vol. 24, no. 3, pp. 279–283, 2017
2017
Later among the works it cites.
V. Boddapati, A. Petef, J. Rasmusson, and L. Lundberg, “Classifying environmental sounds using image recognition networks,” Procedia computer science , vol. 112, pp. 2048–2056, 2017
2017
Later among the works it cites.
D. M. Agrawal, H. B. Sailor, M. H. Soni, and H. A. Patil, “Novel teo-based gammatone features for environmental sound classification,” in 2017 25th European Signal Processing Conference (EUSIPCO) . IEEE, 2017, pp. 1809–1813
2017
Later among the works it cites.
R. N. Tak, D. M. Agrawal, and H. A. Patil, “Novel phase encoded mel filterbank energies for environmental sound classification,” in International Conference on Pattern Recognition and Machine Intelligence . Springer, 2017, pp. 317–325
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in Proceedings of the 22nd ACM International Conference on Multimedia , ser. MM ’14. New York, NY, USA: Association for Computing Machinery, 2014, p. 1041–1044. [Online]. Available: https://doi.org/10.1145/2647868.2655045
2014
Cited alongside, same era.
2014
Cited alongside, same era.
K. J. Piczak, “Esc: Dataset for environmental sound classification,” in Proceedings of the 23rd ACM International Conference on Multimedia , ser. MM ’15. New York, NY, USA: Association for Computing Machinery, 2015, p. 1015–1018. [Online]. Available: https://doi.org/10.1145/2733373.2806390
2015
Cited alongside, same era.
K. J. Piczak, “Environmental sound classification with convolutional neural networks,” in 2015 IEEE 25th International Workshop on Machine Learning for Signal Processing (MLSP) , Sep. 2015, pp. 1–6
2015
Cited alongside, same era.
G. Koch, R. Zemel, and R. Salakhutdinov, “Siamese neural networks for one-shot image recognition,” in ICML deep learning workshop , vol. 2. Lille, 2015
2015
Cited alongside, same era.
Q. Jin and J. Liang, “Video description generation using audio and visual cues,” in Proceedings of the 2016 ACM on International Conference on Multimedia Retrieval , ser. ICMR ’16. New York, NY, USA: Association for Computing Machinery, 2016, p. 239–242. [Online]. Available: https://doi.org/10.1145/2911996.2912043
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2016
2016
Cited alongside, same era.
Y. Tokozume and T. Harada, “Learning environmental sounds with end-to-end convolutional neural network,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , March 2017, pp. 2721–2725
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Later among the works it cites.
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 1251–1258
2017
Later among the works it cites.
2017
Later among the works it cites.
B. Zhu, K. Xu, D. Wang, L. Zhang, B. Li, and Y. Peng, “Environmental sound classification based on multi-temporal resolution convolutional neural network combining with multi-level features,” in Pacific Rim Conference on Multimedia . Springer, 2018, pp. 528–537
2018
Later among the works it cites.
Z. Zhang, S. Xu, S. Cao, and S. Zhang, “Deep convolutional neural network with mixup for environmental sound classification,” in Chinese Conference on Pattern Recognition and Computer Vision (PRCV) . Springer, 2018, pp. 356–367
2018
Later among the works it cites.
S. Abdoli, P. Cardinal, and A. L. Koerich, “End-to-end environmental sound classification using a 1d convolutional neural network,” Expert Systems with Applications , vol. 136, pp. 252–263, 2019
2019
Later among the works it cites.
Z. Zhang, S. Xu, S. Zhang, T. Qiao, and S. Cao, “Learning attentive representations for environmental sound classification,” IEEE Access , vol. 7, pp. 130 327–130 339, 2019
2019
Later among the works it cites.
Y. Su, K. Zhang, J. Wang, and K. Madani, “Environment sound classification using a two-stream cnn based on decision-level fusion,” Sensors , vol. 19, no. 7, p. 1733, Apr 2019. [Online]. Available: http://dx.doi.org/10.3390/s19071733
2019
Later among the works it cites.
B. McFee, V. Lostanlen, M. McVicar, A. Metsai, S. Balke, C. Thome, C. Raffel, A. Malek, D. Lee, F. Zalkow, K. Lee, O. Nieto, J. Mason, D. Ellis, R. Yamamoto, S. Seyfarth, E. Battenberg, V. Morozov, R. Bittner, K. Choi, J. Moore, Z. Wei, S. Hidaka, nullmightybofo, P. Friesch, F.-R. Stoter, D. Herenu, T. Kim, M. Vollrath, and A. Weiss, “librosa/librosa: 0.7.2,” Jan. 2020. [Online]. Available: https://doi.org/10.5281/zenodo.3606573
2020
Closest in time.