Fetching the paper…
Reading the bibliography…
Automatic Music Transcription (AMT), inferring musical notes from raw audio, is a challenging task at the core of music understanding.
The complete MIDI 1.0 detailed specification
MIDI Manufacturers Association · 1996
Earlier work this paper cites.
A discriminative model for polyphonic piano transcription
Graham E Poliner and Daniel PW Ellis · 2006
Earlier work this paper cites.
Polyphonic piano note transcription with recurrent neural networks
Sebastian Böck and Markus Schedl · 2012
Earlier work this paper cites.
JAMS: A JSON annotated music specification for reproducible MIR research
Eric J Humphrey, Justin Salamon, Oriol Nieto, Jon Forsyth, Rachel M Bittner, and Juan Pablo Bello · 2014
Earlier work this paper cites.
pYIN: A fundamental frequency estimator using probabilistic threshold distributions
Matthias Mauch and Simon Dixon · 2014
Earlier work this paper cites.
mir_eval: A transparent implementation of common MIR metrics
Colin Raffel, Brian McFee, Eric J Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, Daniel PW Ellis, and C Colin Raffel · 2014
Earlier work this paper cites.
LibriSpeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
On the potential of simple framewise approaches to piano transcription
Rainer Kelz, Matthias Dorfer, Filip Korzeniowski, Sebastian Böck, Andreas Arzt, and Gerhard Widmer · 2016
Earlier work this paper cites.
Learning-based methods for comparing sequences, with applications to audio-to-MIDI alignment and matching
Colin Raffel · 2016
Earlier work this paper cites.
Learning features of music from scratch
John Thickstun, Zaid Harchaoui, and Sham Kakade · 2016
Earlier work this paper cites.
Onsets and Frames: Dual-objective piano transcription
Curtis Hawthorne, Erich Elsen, Jialin Song, Adam Roberts, Ian Simon, Colin Raffel, Jesse Engel, Sageev Oore, and Douglas Eck · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Increasing drum transcription vocabulary using data synthesis
Mark Cartwright and Juan Pablo Bello · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Enabling factorized piano music modeling and generation with the MAESTRO dataset
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck · 2018
Earlier work this paper cites.
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck · 2018
Earlier work this paper cites.
Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications
Bochen Li, Xinzhao Liu, Karthik Dinesh, Zhiyao Duan, and Gaurav Sharma · 2018
Cited alongside, same era.
Evaluating automatic polyphonic music transcription
Andrew McLeod and Mark Steedman · 2018
Cited alongside, same era.
GuitarSet: A dataset for guitar transcription
Qingyang Xi, Rachel M Bittner, Johan Pauwels, Xuzhou Ye, and Juan Pablo Bello · 2018
Cited alongside, same era.
Automatic music transcription and ethnomusicology: A user study
Andre Holzapfel, Emmanouil Benetos, et al · 2019
Cited alongside, same era.
Deep polyphonic ADSR piano note transcription
Rainer Kelz, Sebastian Böck, and Gerhard Widmer · 2019
Cited alongside, same era.
Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity
Simultaneous separation and transcription of mixtures with multiple polyphonic and percussive instruments
Ethan Manilow, Prem Seetharaman, and Bryan Pardo · 2020
Later among the works it cites.
Multi-instrument music transcription based on deep spherical clustering of spectrograms and pitchgrams
Keitaro Tanaka, Takayuki Nakatsuka, Ryo Nishikimi, Kazuyoshi Yoshii, and Shigeo Morishima · 2020
Later among the works it cites.
State-based transcription of components of carnatic music
Venkata Subramanian Viraraghavan, Arpan Pal, Hema Murthy, and Rangarajan Aravind · 2020
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2020
Later among the works it cites.
Codified audio language modeling learns useful representations for music information retrieval
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ethan Manilow, Gordon Wichern, Prem Seetharaman, and Jonathan Le Roux · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Cited alongside, same era.
Common Voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber · 2020
Cited alongside, same era.
JAX: composable transformations of Python + NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Improving perceptual quality of drum transcription with the expanded Groove MIDI Dataset
Lee Callender, Curtis Hawthorne, and Jesse Engel · 2020
Cited alongside, same era.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Cited alongside, same era.
Rodrigo Castellon, Chris Donahue, and Percy Liang · 2021
Closest in time.
ReconVAT: A semi-supervised automatic music transcription framework for low-resource real-world data
Kin Wai Cheuk, Dorien Herremans, and Li Su · 2021
Closest in time.
Variable-rate discrete representation learning
Sander Dieleman, Charlie Nash, Jesse Engel, and Karen Simonyan · 2021
Closest in time.
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Closest in time.
Sequence-to-sequence piano transcription with Transformers
Curtis Hawthorne, Ian Simon, Rigel Swavely, Ethan Manilow, and Jesse Engel · 2021
Closest in time.
Yuma Koizumi, Shigeki Karita, Scott Wisdom, Hakan Erdogan, John R Hershey, Llion Jones, and Michiel Bacchiani · 2021
Closest in time.
A unified model for zero-shot music source separation, transcription and synthesis
Liwei Lin, Qiuqiang Kong, Junyan Jiang, and Gus Xia · 2021
Closest in time.
Xinhao Mei, Xubo Liu, Qiushi Huang, Mark D Plumbley, and Wenwu Wang · 2021
Closest in time.
Attention is all you need in speech separation
Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, and Jianyuan Zhong · 2021
Closest in time.
Prateek Verma and Jonathan Berger · 2021
Closest in time.
A generative model for raw audio using Transformer architectures
Prateek Verma and Chris Chafe · 2021
Closest in time.