Fetching the paper…
Reading the bibliography…
Speech, Music and Noise classification/segmentation is an important preprocessing step for audio processing/indexing.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum mean-square error log-spectral amplitude estimator,” IEEE transactions on acoustics, speech, and signal processing , vol. 33, no. 2, pp. 443–445, 1985
1985
Earlier work this paper cites.
Y. LeCun and Y. Bengio, “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks , vol. 3361, no. 10, p. 1995, 1995
1995
Earlier work this paper cites.
B. Logan, “Mel Frequency Cepstral Coefficients for Music Modeling,” in ISMIR , vol. 270, 2000, pp. 1–11
2000
Earlier work this paper cites.
J. Wolfe, “Speech and music, acoustics and coding, and what music might be ‘for’,” in Proc. 7th International Conference on Music Perception and Cognition , 2002, pp. 10–13
2002
Earlier work this paper cites.
G. Tzanetakis and P. Cook, “Musical genre classification of audio signals,” IEEE Transactions on speech and audio processing , vol. 10, no. 5, pp. 293–302, 2002
2002
Earlier work this paper cites.
Y. Bengio, “Learning deep architectures for AI,” Foundations and trends® in Machine Learning , vol. 2, no. 1, pp. 1–127, 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
J. Kola, C. Espy-Wilson, and T. Pruthi, “Voice activity detection,” Merit Bien , pp. 1–6, 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
M. Oquab, L. Bottou, I. Laptev, and J. Sivic, “Learning and transferring mid-level image representations using convolutional neural networks,” in Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on . IEEE, 2014, pp. 1717–1724
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, p. 436, 2015
2015
Cited alongside, same era.
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” arXiv preprint , 2016
2016
Later among the works it cites.
T. Sercu, C. Puhrsch, B. Kingsbury, and Y. LeCun, “Very deep multilingual convolutional neural networks for LVCSR,” in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on . IEEE, 2016, pp. 4955–4959
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1–9
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Cited alongside, same era.
A. van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, and A. Graves, “Conditional image generation with pixelcnn decoders,” in Advances in Neural Information Processing Systems , 2016, pp. 4790–4798
2016
Cited alongside, same era.
“MUSAN - OpenSLR [Online].” [Online]. Available: http://www.openslr.org/17/
Cited in the paper.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, and M. Isard, “Tensorflow: a system for large-scale machine learning,” in OSDI , vol. 16, 2016, pp. 265–283
2016
Later among the works it cites.
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, and B. Seybold, “CNN architectures for large-scale audio classification,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 131–135
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter, “Self-normalizing neural networks,” in Advances in Neural Information Processing Systems , 2017, pp. 971–980
2017
Later among the works it cites.