Fetching the paper…
Reading the bibliography…
Melody extraction is a vital music information retrieval task among music researchers for its potential applications in education pedagogy and the music industry.
J. F. Schouten, R. Ritsma, B. L. Cardozo, Pitch of the residue, The Journal of the Acoustical Society of America 34 (9B) (1962) 1418–1424
1962
Earlier work this paper cites.
E. D. Scheirer, Machine-listening systems, Unpublished Ph. D. Thesis, Massachusetts Institute of Technology (2000)
2000
Earlier work this paper cites.
S. Shamma, D. Klein, The case of the missing pitch templates: how harmonic templates emerge in the early auditory system, The Journal of the Acoustical Society of America 107 (5) (2000) 2631–2644
2000
Earlier work this paper cites.
J. Laroche, Time and pitch scale modification of audio signals, in: Applications of digital signal processing to audio and acoustics, Springer, 2002, pp. 279–309
2002
Earlier work this paper cites.
M. Goto, A real-time music-scene-description system: Predominant-f0 estimation for detecting melody and bass lines in real-world audio signals, Speech Communication 43 (4) (2004) 311–329
2004
Earlier work this paper cites.
T. Chi, P. Ru, S. A. Shamma, Multiresolution spectrotemporal analysis of complex sounds, The Journal of the Acoustical Society of America 118 (2) (2005) 887–906
2005
Earlier work this paper cites.
G. E. Poliner, D. P. Ellis, A. F. Ehmann, E. Gómez, S. Streich, B. Ong, Melody transcription from music audio: Approaches and evaluation, IEEE Transactions on Audio, Speech, and Language Processing 15 (4) (2007) 1247–1256
2007
Earlier work this paper cites.
J.-C. Chen, J.-S. R. Jang, Trues: Tone recognition using extended segments, ACM Transactions on Asian Language Information Processing (TALIP) 7 (3) (2008) 1–23
2008
Earlier work this paper cites.
W. A. Yost, Pitch perception, Attention, Perception, & Psychophysics 71 (8) (2009) 1701–1715
2009
Earlier work this paper cites.
J.-L. Durrieu, G. Richard, B. David, C. Févotte, Source/filter model for unsupervised main melody extraction from polyphonic audio signals, IEEE transactions on audio, speech, and language processing 18 (3) (2010) 564–575
2010
Earlier work this paper cites.
H. Tachibana, T. Ono, N. Ono, S. Sagayama, Melody line estimation in homophonic music audio signals based on temporal-variability of melodic source, in: 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, IEEE, 2010, pp. 425–428
2010
Earlier work this paper cites.
V. Rao, P. Rao, Vocal melody extraction in the presence of pitched accompaniment in polyphonic music, IEEE Transactions on Audio, Speech, and Language Processing 18 (8) (2010) 2145–2154
2010
Earlier work this paper cites.
C. Schörkhuber, A. Klapuri, Constant-q transform toolbox for music processing, in: 7th sound and music computing conference, Barcelona, Spain, 2010, pp. 3–64
2010
Earlier work this paper cites.
D. FitzGerald, M. Gainza, Single channel vocal separation using median filtering and factorisation techniques (2010)
2010
Earlier work this paper cites.
J.-L. Durrieu, B. David, G. Richard, A musically motivated mid-level representation for pitch estimation and musical audio source separation, IEEE Journal of Selected Topics in Signal Processing 5 (6) (2011) 1180–1191
2011
Earlier work this paper cites.
P. S. Huang, S. D. Chen, P. Smaragdis, M. Hasegawa-Johnson, Singing-voice separation from monaural recordings using robust principal component analysis, in: Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2012, pp. 57–60
2012
Earlier work this paper cites.
J. Salamon, E. Gómez, Melody extraction from polyphonic music signals using pitch contour characteristics, IEEE Transactions on Audio, Speech, and Language Processing 20 (6) (2012) 1759–1770
2012
Earlier work this paper cites.
Z. Rafii, B. Pardo, Repeating pattern extraction technique (repet): A simple method for music/voice separation, IEEE Transactions on Audio, Speech, and Language Processing 21 (1) (2013) 73–84
2013
Earlier work this paper cites.
J. Salamon, Melody extraction from polyphonic music signals, Ph.D. thesis, Department of Information and Communication Technologies Universitat Pompeu Fabra, Barcelona, Spain (2013)
2013
Earlier work this paper cites.
M. Mauch, S. Ewert, et al., The audio degradation toolbox and its application to robustness evaluation (2013)
2013
Earlier work this paper cites.
J. Salamon, E. Gomez, D. P. Ellis, G. Richard, Melody extraction from polyphonic music signals: Approaches, applications, and challenges, IEEE Signal Processing Magazine 31 (2) (2014) 118–134
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, D. P. Ellis, C. C. Raffel, mir_eval: A transparent implementation of common mir metrics, in: In Proceedings of the 15th International Society for Music Information Retrieval Conference, ISMIR, Citeseer, 2014
2014
Cited alongside, same era.
S. Kum, C. Oh, J. Nam, Melody extraction on vocal segments using multi-column deep neural networks., in: ISMIR, 2016, pp. 819–825
2016
Cited alongside, same era.
F. Rigaud, M. Radenen, Singing voice melody transcription using deep neural networks., in: ISMIR, 2016, pp. 737–743
2016
Cited alongside, same era.
Z.-C. Fan, J.-S. R. Jang, C.-L. Lu, Singing voice separation and pitch extraction from monaural polyphonic audio music via dnn and adaptive pitch tracking, in: 2016 IEEE Second International Conference on Multimedia Big Data (BigMM), IEEE, 2016, pp. 178–185
2016
Cited alongside, same era.
M.-T. Chen, B.-J. Li, T.-S. Chi, Cnn based two-stage multi-resolution end-to-end model for singing melody extraction, in: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2019, pp. 1005–1009
2019
Later among the works it cites.
Y. Gao, B. Zhu, W. Li, K. Li, Y. Wu, F. Huang, Vocal melody extraction via dnn-based pitch estimation and salience-based pitch refinement, in: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2019, pp. 1000–1004
2019
Later among the works it cites.
T. Nakano, K. Yoshii, Y. Wu, R. Nishikimi, K. W. E. Lin, M. Goto, Joint singing pitch estimation and voice separation based on a neural harmonic structure renderer, in: 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), IEEE, 2019, pp. 160–164
2019
Later among the works it cites.
A. Jansson, R. M. Bittner, S. Ewert, T. Weyde, Joint singing voice separation and f0 estimation with deep u-net architectures, in: 2019 27th European Signal Processing Conference (EUSIPCO), IEEE, 2019, pp. 1–5
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, J. Sun, Identity mappings in deep residual networks, in: European conference on computer vision, Springer, 2016, pp. 630–645
2016
Cited alongside, same era.
R. M. Bittner, B. McFee, J. Salamon, P. Li, J. P. Bello, Deep salience representations for f0 estimation in polyphonic music., in: ISMIR, 2017, pp. 63–70
2017
Cited alongside, same era.
H. Park, C. D. Yoo, Melody extraction and detection through lstm-rnn with harmonic sum loss, in: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2017, pp. 2766–2770
2017
Cited alongside, same era.
S. Kum, J. Nam, Classification-based singing melody extraction using deep convolutional neural networks (2017)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
V. Badrinarayanan, A. Kendall, R. Cipolla, Segnet: A deep convolutional encoder-decoder architecture for image segmentation, IEEE transactions on pattern analysis and machine intelligence 39 (12) (2017) 2481–2495
2017
Cited alongside, same era.
G. Huang, Z. Liu, L. Van Der Maaten, K. Q. Weinberger, Densely connected convolutional networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 4700–4708
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, in: Advances in neural information processing systems, 2017, pp. 5998–6008
2017
Cited alongside, same era.
2019
Later among the works it cites.
S. Kum, J. Nam, Joint detection and classification of singing voice melody using convolutional recurrent neural networks, Applied Sciences 9 (7) (2019) 1324
2019
Later among the works it cites.
2019
Later among the works it cites.
K. Sun, B. Xiao, D. Liu, J. Wang, Deep high-resolution representation learning for human pose estimation, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 5693–5703
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
X. Li, W. Wang, X. Hu, J. Yang, Selective kernel networks, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 510–519
2019
Later among the works it cites.
S. He, L. Schomaker, Deepotsu: Document enhancement and binarization using iterative deep learning, Pattern recognition 91 (2019) 379–390
2019
Later among the works it cites.
P. Gao, C.-Y. You, T.-S. Chi, A multi-dilation and multi-resolution fully convolutional network for singing melody extraction, in: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 551–555
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y. Mu, M. Tan, X. Wang, et al., Deep high-resolution representation learning for visual recognition, IEEE transactions on pattern analysis and machine intelligence (2020)
2020
Later among the works it cites.
Y. Wang, Q. Yao, J. T. Kwok, L. M. Ni, Generalizing from a few examples: A survey on few-shot learning, ACM Computing Surveys (CSUR) 53 (3) (2020) 1–34
2020
Later among the works it cites.
X. Du, B. Zhu, Q. Kong, Z. Ma, Singing melody extraction from polyphonic music based on spectral correlation modeling, in: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2021, pp. 241–245
2021
Later among the works it cites.
Y. Gao, X. Du, B. Zhu, X. Sun, W. Li, Z. Ma, An hrnet-blstm model with two-stage training for singing melody extraction, in: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2021, pp. 56–60
2021
Later among the works it cites.
S. Yu, X. Sun, Y. Yu, W. Li, Frequency-temporal attention network for singing melody extraction, in: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2021, pp. 251–255
2021
Later among the works it cites.
S. Yu, Y. Yu, X. Chen, W. Li, Hanme: Hierarchical attention network for singing melody extraction, IEEE Signal Processing Letters 28 (2021) 1006–1010
2021
Later among the works it cites.
Y. Gao, X. Zhang, W. Li, Vocal melody extraction via hrnet-based singing voice separation and encoder-decoder-based f0 estimation, Electronics 10 (3) (2021) 298
2021
Later among the works it cites.
Musicinformationretrieval.com/, https://musicinformationretrieval.com/ , accessed: 2022-19-01
2022
Closest in time.