Fetching the paper…
Reading the bibliography…
Automatic Music Transcription has seen significant progress in recent years by training custom deep neural networks on large datasets.
A. Friberg and J. Sundberg, “Perception of just-noticeable time displacement of a tone presented in a metrical sequence at different tempos,” Journal of The Acoustical Society of America , vol. 94, no. 3, pp. 1859–1859, 1993
1993
Earlier work this paper cites.
S. Handel, Listening: An introduction to the perception of auditory events. MIT Press, 1993
1993
Earlier work this paper cites.
MIDI Manufacturers Association and others, “The complete midi 1.0 detailed specification,” Los Angeles, CA, The MIDI Manufacturers Association , 1996
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
V. Emiya, R. Badeau, and B. David, “Multipitch estimation of piano sounds using a new probabilistic spectral smoothness principle,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 6, pp. 1643–1654, 2009
2009
Earlier work this paper cites.
J. Nam, J. Ngiam, H. Lee, and M. Slaney, “A classification-based polyphonic piano transcription approach using learned feature representations,” in ISMIR , 2011
2011
Earlier work this paper cites.
N. Boulanger-Lewandowski, Y. Bengio, and P. Vincent, “Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription,” in ICML , 2012
2012
Earlier work this paper cites.
S. Böck and M. Schedl, “Polyphonic piano note transcription with recurrent neural networks,” in ICASSP , 2012
2012
Earlier work this paper cites.
C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, D. P. Ellis, and C. C. Raffel, “mir_eval: A transparent implementation of common MIR metrics,” in ISMIR , 2014
2014
Earlier work this paper cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015. [Online]. Available: https://www.tensorflow.org/
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
K. Ullrich and E. van der Wel, “Music transcription with convolutional sequence-to-sequence models,” in ISMIR , 2017
2017
Earlier work this paper cites.
C. Hawthorne, E. Elsen, J. Song, A. Roberts, I. Simon, C. Raffel, J. Engel, S. Oore, and D. Eck, “Onsets and Frames: Dual-objective piano transcription,” in ISMIR , 2018
2018
Cited alongside, same era.
L. Dong, S. Xu, and B. Xu, “Speech-Transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in ICASSP , 2018
2018
Cited alongside, same era.
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang, “JAX: composable transformations of Python+NumPy programs,” 2018. [Online]. Available: http://github.com/google/jax
2018
Cited alongside, same era.
N. Shazeer and M. Stern, “Adafactor: Adaptive learning rates with sublinear memory cost,” in International Conference on Machine Learning . PMLR, 2018, pp. 4596–4604
2018
Cited alongside, same era.
T. Kwon, D. Jeong, and J. Nam, “Polyphonic piano transcription using autoregressive multi-state note model,” in ISMIR , 2020
2020
Later among the works it cites.
A. Elowsson, “Polyphonic pitch tracking with deep layered learning,” Journal of the Acoustical Society of America , vol. 148, no. 1, pp. 446–468, 2020
2020
Later among the works it cites.
J. Engel, R. Swavely, L. H. Hantrakul, A. Roberts, and C. Hawthorne, “Self-supervised pitch detection by inverse audio synthesis,” in ICML Workshop on Self-Supervision in Audio and Speech , 2020
2020
Later among the works it cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in NeurIPS , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.-Z. A. Huang, S. Dieleman, E. Elsen, J. Engel, and D. Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=r1lYRjC9F7
2019
Cited alongside, same era.
N.-Q. Pham, T.-S. Nguyen, J. Niehues, M. Müller, S. Stüker, and A. Waibel, “Very deep self-attention networks for end-to-end speech recognition,” in Interspeech , 2019
2019
Cited alongside, same era.
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, “Neural speech synthesis with Transformer network,” in AAAI , 2019
2019
Cited alongside, same era.
J. W. Kim and J. P. Bello, “Adversarial learning for improved onsets and frames music transcription,” in ISMIR , 2019
2019
Cited alongside, same era.
R. Kelz, S. Böck, and G. Widmer, “Deep polyphonic ADSR piano note transcription,” in ICASSP , 2019
2019
Cited alongside, same era.
M. Awiszus, “Automatic music transcription using sequence to sequence learning,” Master’s thesis, Karlsruhe Institute of Technology, 2019
2019
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text Transformer,” Journal of Machine Learning Research , vol. 21, no. 140, pp. 1–67, 2020
2020
Cited alongside, same era.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with Transformers,” in ECCV , 2020
2020
Cited alongside, same era.
2020
Later among the works it cites.
S. Oore, I. Simon, S. Dieleman, D. Eck, and K. Simonyan, “This time with feeling: Learning expressive musical performance,” Neural Computing and Applications , vol. 32, no. 4, pp. 955–967, 2020
2020
Later among the works it cites.
A. Ycart, L. Liu, E. Benetos, and M. Pearce, “Investigating the perceptual validity of evaluation metrics for automatic piano music transcription,” Transactions of the International Society for Music Information Retrieval , 2020
2020
Later among the works it cites.
N. Shazeer, “Glu variants improve transformer,” arXiv preprint arXiv:2002.05202 , 2020
2020
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=YicbFdNTTy
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
K. Lin, L. Wang, and Z. Liu, “End-to-end human pose and mesh reconstruction with Transformers,” in CVPR , 2021
2021
Closest in time.