Fetching the paper…
Reading the bibliography…
Recently, the end-to-end approach that learns hierarchical representations from raw data using deep convolutional neural networks has been successfully explored in the image, text and speech domains.
E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “Evaluation of algorithms using games: The case of music tagging,” in ISMIR , 2009, pp. 387–392
2009
Earlier work this paper cites.
D. Erhan, Y. Bengio, A. Courville, and P. Vincent, “Visualizing higher-layer features of a deep network,” University of Montreal , vol. 1341, p. 3, 2009
2009
Earlier work this paper cites.
P. Hamel, S. Lemieux, Y. Bengio, and D. Eck, “Temporal pooling and multiscale learning for automatic annotation and ranking of music audio,” in Proceedings of the 12th International Conference on Music Information Retrieval (ISMIR) , 2011
2011
Earlier work this paper cites.
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere, “The million song dataset,” in Proceedings of the 12th International Conference on Music Information Retrieval (ISMIR) , vol. 2, no. 9, 2011, pp. 591–596
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Advances in neural information processing systems , 2013, pp. 3111–3119
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
S. Dieleman and B. Schrauwen, “End-to-end learning for music audio,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 6964–6968
2014
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in European conference on computer vision . Springer, 2014, pp. 818–833
2014
Cited alongside, same era.
X. Zhang, J. Zhao, and Y. LeCun, “Character-level convolutional networks for text classification,” in Advances in neural information processing systems , 2015, pp. 649–657
2015
Cited alongside, same era.
2015
Cited alongside, same era.
D. Palaz, M. M. Doss, and R. Collobert, “Convolutional neural networks-based continuous speech recognition using raw speech signal,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 4295–4299
2015
Cited alongside, same era.
2016
Later among the works it cites.
2016
Later among the works it cites.
J. Pons, T. Lidy, and X. Serra, “Experimenting with musically motivated convolutional neural networks,” in IEEE International Workshop on Content-Based Multimedia Indexing (CBMI) , 2016, pp. 1–6
2016
Later among the works it cites.
K. Choi, G. Fazekas, and M. Sandler, “Automatic tagging using deep convolutional neural networks,” in Proceedings of the 17th International Conference on Music Information Retrieval (ISMIR) , 2016, pp. 805–811
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Palaz, R. Collobert et al. , “Analysis of cnn-based speech recognition system using raw speech as input,” Idiap, Tech. Rep., 2015
2015
Cited alongside, same era.
T. N. Sainath, R. J. Weiss, A. W. Senior, K. W. Wilson, and O. Vinyals, “Learning the speech front-end with raw waveform cldnns.” in INTERSPEECH , 2015, pp. 1–5
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
D. Ardila, C. Resnick, A. Roberts, and D. Eck, “Audio deepdream: Optimizing raw audio with convolutional networks.”
Cited in the paper.
2016
Later among the works it cites.
2016
Later among the works it cites.