Fetching the paper…
Reading the bibliography…
We present a supervised neural network model for polyphonic piano music transcription.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Cognitive Modeling , vol. 5, p. 3, 1988
1988
Earlier work this paper cites.
L. R. Rabiner, “A tutorial on hidden Markov models and selected applications in speech recognition,” Proceedings of the IEEE , vol. 77, no. 2, pp. 257–286, 1989
1989
Earlier work this paper cites.
P. J. Werbos, “Backpropagation through time: what it does and how to do it,” Proceedings of the IEEE , vol. 78, no. 10, pp. 1550–1560, 1990
1990
Earlier work this paper cites.
J. C. Brown, “Calculation of a constant q spectral transform,” The Journal of the Acoustical Society of America , vol. 89, no. 1, pp. 425–434, 1991
1991
Earlier work this paper cites.
R. M. Neal, “Connectionist learning of belief networks,” Artificial Intelligence , vol. 56, no. 1, pp. 71–113, 1992
1992
Earlier work this paper cites.
L. Rabiner and B.-H. Juang, “Fundamentals of speech recognition,” 1993
1993
Earlier work this paper cites.
T. H. Cormen, C. E. Leiserson, R. L. Rivest, and C. Stein, Introduction to algorithms . MIT press Cambridge, 2001, vol. 2
2001
Earlier work this paper cites.
D. Eck and J. Schmidhuber, “Finding temporal structure in music: Blues improvisation with lstm recurrent networks,” in Neural Networks for Signal Processing, 2002. Proceedings of the 2002 12th IEEE Workshop on . IEEE, 2002, pp. 747–756
2002
Earlier work this paper cites.
A. P. Klapuri, “Multiple fundamental frequency estimation based on harmonicity and spectral smoothness,” IEEE Transactions on Speech and Audio Processing , vol. 11, no. 6, pp. 804–816, 2003
2003
Earlier work this paper cites.
P. Smaragdis and J. C. Brown, “Non-negative matrix factorization for polyphonic music transcription,” in 2003 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. IEEE, 2003, pp. 177–180
2003
Earlier work this paper cites.
S. A. Abdallah and M. D. Plumbley, “Polyphonic music transcription by non-negative sparse coding of power spectra,” in Proceedings of the 5th International Society for Music Information Retrieval Conference (ISMIR) , 2004, pp. 318–325
2004
Earlier work this paper cites.
E. Vincent and X. Rodet, “Music transcription with ISA and HMM,” in Independent Component Analysis and Blind Signal Separation . Springer, 2004, pp. 1197–1204
2004
Earlier work this paper cites.
P. Smaragdis, B. Raj, and M. Shashanka, “A probabilistic latent variable model for acoustic modeling.” Advances in models for acoustic processing, NIPS , vol. 148, 2006
2006
Earlier work this paper cites.
J. Bergstra, N. Casagrande, D. Erhan, D. Eck, and B. Kégl, “Aggregate features and adaboost for music classification,” Machine learning , vol. 65, no. 2-3, pp. 473–484, 2006
2006
Earlier work this paper cites.
A. Klapuri and M. Davy, Signal processing methods for music transcription . Springer Science & Business Media, 2007
2007
Earlier work this paper cites.
G. E. Poliner and D. P. Ellis, “A discriminative model for polyphonic piano transcription,” EURASIP Journal on Applied Signal Processing , vol. 2007, no. 1, pp. 154–154, 2007
2007
Earlier work this paper cites.
V. Emiya, R. Badeau, and B. David, “Automatic transcription of piano music based on hmm tracking of jointly-estimated pitches,” in 16th European on Signal Processing Conference . IEEE, 2008, pp. 1–5
2008
Earlier work this paper cites.
M. Bay, A. F. Ehmann, and J. S. Downie, “Evaluation of multiple-F0 estimation and tracking systems.” in Proceedings of the 9th International Society for Music Information Retrieval Conference (ISMIR) , 2009, pp. 315–320
2009
Earlier work this paper cites.
E. Vincent, N. Bertin, and R. Badeau, “Adaptive harmonic spectral decomposition for multiple pitch estimation,” IEEE Transactions on Audio, Speech, and Language Processing. , vol. 18, no. 3, pp. 528–537, 2010
2010
Cited alongside, same era.
N. Bertin, R. Badeau, and E. Vincent, “Enforcing harmonicity and smoothness in Bayesian non-negative matrix factorization applied to polyphonic music transcription,” IEEE Transactions on Audio, Speech, and Language Processing. , vol. 18, no. 3, pp. 538–549, 2010
2010
Cited alongside, same era.
G. C. Grindlay and D. P. Ellis, “A probabilistic subspace model for multi-instrument polyphonic transcription,” in Proceedings of the 11th International Society for Music Information Retrieval Conference (ISMIR) . International Society for Music Information Retrieval, 2010, pp. 21–26
2010
Cited alongside, same era.
V. Emiya, R. Badeau, and B. David, “Multipitch estimation of piano sounds using a new probabilistic spectral smoothness principle,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 6, pp. 1643–1654, 2010
S. Raczynski, E. Vincent, and S. Sagayama, “Dynamic Bayesian networks for symbolic polyphonic pitch modeling,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 21, no. 9, pp. 1830–1840, 2013
2013
Later among the works it cites.
N. Boulanger-Lewandowski, Y. Bengio, and P. Vincent, “High-dimensional sequence transduction,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2013, pp. 3178–3182
2013
Later among the works it cites.
O. Abdel-Hamid, L. Deng, and D. Yu, “Exploring convolutional neural network structures and optimization techniques for speech recognition.” in INTERSPEECH , 2013, pp. 3366–3370
2013
Later among the works it cites.
E. J. Humphrey, J. P. Bello, and Y. LeCun, “Feature learning and deep architectures: new directions for music informatics,” Journal of Intelligent Information Systems , vol. 41, no. 3, pp. 461–481, 2013
2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2010
Cited alongside, same era.
J. Nam, J. Ngiam, H. Lee, and M. Slaney, “A classification-based polyphonic piano transcription approach using learned feature representations.” in Proceedings of the 12th International Society for Music Information Retrieval Conference (ISMIR) , 2011, pp. 175–180
2011
Cited alongside, same era.
H. Larochelle and I. Murray, “The neural autoregressive distribution estimator,” in International Conference on Artificial Intelligence and Statistics , 2011, pp. 29–37
2011
Cited alongside, same era.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks,” in International Conference on Artificial Intelligence and Statistics , 2011, pp. 315–323
2011
Cited alongside, same era.
E. Benetos and S. Dixon, “A shift-invariant latent variable model for automatic music transcription,” Computer Music Journal , vol. 36, no. 4, pp. 81–94, 2012
2012
Cited alongside, same era.
S. Böck and M. Schedl, “Polyphonic piano note transcription with recurrent neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2012, pp. 121–124
2012
Cited alongside, same era.
O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, and G. Penn, “Applying convolutional neural networks concepts to hybrid NN-HMM model for speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2012, pp. 4277–4280
2012
Cited alongside, same era.
E. J. Humphrey and J. P. Bello, “Rethinking automatic chord recognition with convolutional neural networks,” in Machine Learning and Applications (ICMLA), 2012 11th International Conference on , vol. 2. IEEE, 2012, pp. 357–362
2012
Cited alongside, same era.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, and Q. V. Le, “Large scale distributed deep networks,” in Advances in Neural Information Processing Systems (NIPS) , 2012, pp. 1223–1231
2012
Cited alongside, same era.
N. Boulanger-Lewandowski, Y. Bengio, and P. Vincent, “Audio chord recognition with recurrent neural networks.” in Proceedings of the 13th International Society for Music Information Retrieval Conference (ISMIR) , 2013, pp. 335–340
2013
Later among the works it cites.
Y. Bengio, N. Boulanger-Lewandowski, and R. Pascanu, “Advances in optimizing recurrent networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2013, pp. 8624–8628
2013
Later among the works it cites.
T. Berg-Kirkpatrick, J. Andreas, and D. Klein, “Unsupervised transcription of piano music,” in Advances in Neural Information Processing Systems (NIPS) , 2014, pp. 1538–1546
2014
Later among the works it cites.
S. Sigtia, E. Benetos, S. Cherla, T. Weyde, A. S. d. Garcez, and S. Dixon, “An RNN-based music language model for improving automatic music transcription,” in Proceedings of the 15th International Society for Music Information Retrieval Conference (ISMIR) , 2014
2014
Later among the works it cites.
J. Schlüter and S. Böck, “Improved musical onset detection with convolutional neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2014, pp. 6979–6983
2014
Later among the works it cites.
N. Boulanger-Lewandowski, J. Droppo, M. Seltzer, and D. Yu, “Phone sequence modeling with recurrent neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 5417–5421
2014
Later among the works it cites.
K. O’Hanlon and M. D. Plumbley, “Polyphonic piano transcription using non-negative matrix factorisation with group sparsity,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2014, pp. 3112–3116
2014
Later among the works it cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” The Journal of Machine Learning Research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Later among the works it cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Closest in time.
S. Sigtia, N. Boulanger-Lewandowski, and S. Dixon, “Audio chord recognition with a hybrid neural network,” in Proceedings of the 16th International Society for Music Information Retrieval Conference (ISMIR) , 2015
2015
Closest in time.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled sampling for sequence prediction with recurrent neural networks,” in Advances in Neural Information Processing Systems , 2015, pp. 1171–1179
2015
Closest in time.
2015
Closest in time.
S. Sigtia, E. Benetos, N. Boulanger-Lewandowski, T. Weyde, A. S. d’Avila Garcez, and S. Dixon, “A Hybrid Recurrent Neural Network for Music Transcription,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , Brisbane, Australia, April 2015, pp. 2061–2065
2065
Closest in time.