Fetching the paper…
Reading the bibliography…
Automatic singing voice understanding tasks, such as singer identification, singing voice transcription, and singing technique classification, benefit from data-driven approaches that utilize deep learning techniques.
A. Ghias, J. Logan, D. Chamberlin, and B. C. Smith, “Query by humming: Musical information retrieval in an audio database,” in Proceedings of the third ACM international conference on Multimedia , 1995, pp. 231–236
1995
Earlier work this paper cites.
T. Nakano, M. Goto, and Y. Hiraga, “Mirusinger: A singing skill visualization interface using real-time feedback and music cd recordings as referential data,” in Ninth IEEE International Symposium on Multimedia Workshops (ISMW 2007) . IEEE, 2007, pp. 75–76
2007
Earlier work this paper cites.
D. P. Ellis, “Classifying music audio with timbral and chroma features,” in The 8th International Conference for Music Information Retrieval Conference (ISMIR) , 2007
2007
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
E. Molina, A. M. Barbancho-Perez, L. J. Tardon-Garcia, I. Barbancho-Perez et al. , “Evaluation framework for automatic singing transcription,” in The 15th International Society for Music Information Retrieval Conference (ISMIR) , 2014
2014
Earlier work this paper cites.
J. Schlüter and T. Grill, “Exploring data augmentation for improved singing voice detection with neural networks.” in The 16th International Society for Music Information Retrieval Conference (ISMIR) , 2015, pp. 121–126
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
M. Panteli, R. Bittner, J. P. Bello, and S. Dixon, “Towards the characterization of singing styles in world music,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 636–640
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. J. Weiss, and K. Wilson, “Cnn architectures for large-scale audio classification,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2017, pp. 131–135
2017
Earlier work this paper cites.
E. J. Humphrey, S. Reddy, P. Seetharaman, A. Kumar, R. M. Bittner, A. Demetriou, S. Gulati, A. Jansson, T. Jehan, B. Lehner, A. Kruspe, and L. Yang, “An introduction to signal processing for singing-voice analysis: High notes in the effort to automate the understanding of vocals in music,” IEEE Signal Processing Magazine , vol. 36, no. 1, pp. 82–94, 2018
2018
Earlier work this paper cites.
J. Wilkins, P. Seetharaman, A. Wahl, and B. A. Pardo, “Vocalset: A singing voice dataset,” in The Proceedings of the 19th International Society for Music Information Retrieval Conference (ISMIR) , 2018, pp. 468–474
2018
Earlier work this paper cites.
Z.-S. Fu and L. Su, “Hierarchical classification networks for singing voice segmentation and transcription,” in The 20th International Society for Music Information Retrieval Conference (ISMIR) , 2019
2019
Earlier work this paper cites.
Z. Nasrullah and Y. Zhao, “Music artist classification with convolutional recurrent neural networks,” in 2019 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2019, pp. 1–8
2019
Earlier work this paper cites.
M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International conference on machine learning . PMLR, 2019, pp. 6105–6114
2019
Earlier work this paper cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems , vol. 33, pp. 12 449–12 460, 2020
2020
Earlier work this paper cites.
S. Kum, J.-H. Lin, L. Su, and J. Nam, “Semi-supervised learning using teacher-student models for vocal melody extraction,” in The 21st International Society for Music Information Retrieval Conference (ISMIR) , 2020
2020
Earlier work this paper cites.
T.-H. Hsieh, K.-H. Cheng, Z.-C. Fan, Y.-C. Yang, and Y.-H. Yang, “Addressing the confounds of accompaniments in singer identification,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2020, pp. 1–5
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
E. Demirel, S. Ahlbäck, and S. Dixon, “Automatic lyrics transcription using dilated convolutional neural networks with self-attention,” in 2020 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2020, pp. 1–8
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar et al. , “Bootstrap your own latent-a new approach to self-supervised learning,” Advances in neural information processing systems , vol. 33, pp. 21 271–21 284, 2020
2020
2022
Later among the works it cites.
L. Ou, X. Gu, and Y. Wang, “Transfer learning of wav2vec 2.0 for automatic lyric transcription,” in The 23rd International Society for Music Information Retrieval Conference (ISMIR) , 2022
2022
Later among the works it cites.
M. Heydari and Z. Duan, “Singing beat tracking with self-supervised front-end and linear transformers,” in The 23rd International Society for Music Information Retrieval Conference (ISMIR) , 2022
2022
Later among the works it cites.
C. Donahue, J. Thickstun, and P. Liang, “Melody transcription via generative pre-training,” in The 23rd International Society for Music Information Retrieval Conference (ISMIR) , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2020
Cited alongside, same era.
B. Kang, S. Xie, M. Rohrbach, Z. Yan, A. Gordo, J. Feng, and Y. Kalantidis, “Decoupling representation and classifier for long-tailed recognition,” in International Conference on Learning Representations (ICLR) , 2020
2020
Cited alongside, same era.
Y. Yamamoto, J. Nam, H. Terasawa, and Y. Hiraga, “Investigating time-frequency representations for audio feature extraction in singing technique classification,” in 2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) . IEEE, 2021, pp. 890–896
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
J.-Y. Hsu and L. Su, “Vocano: A note transcription framework for singing voice in polyphonic music.” in The 22nd International Society for Music Information Retrieval Conference (ISMIR) , 2021, pp. 293–300
2021
Cited alongside, same era.
S. Basak, S. Agarwal, S. Ganapathy, and N. Takahashi, “End-to-end lyrics recognition with voice to singing style transfer,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2021, pp. 266–270
2021
Cited alongside, same era.
——, “Mstre-net: Multistreaming acoustic modeling for automatic lyrics transcription,” in The 22nd International Society for Music Information Retrieval Conference (ISMIR) , 2021
2021
Cited alongside, same era.
R. Castellon, C. Donahue, and P. Liang, “Codified audio language modeling learns useful representations for music information retrieval,” in The 22nd International Society for Music Information Retrieval Conference (ISMIR) , 2021
2021
Cited alongside, same era.
H. Yakura, K. Watanabe, and M. Goto, “Self-supervised contrastive learning for singing voices,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1614–1623, 2022
2022
Later among the works it cites.
C. Zhang, J. Yu, L. Chang, X. Tan, J. Chen, T. Qin, and K. Zhang, “Pdaugment: Data augmentation by pitch and duration adjustments for automatic lyrics transcription,” in The 23rd International Society for Music Information Retrieval Conference (ISMIR) , 2022
2022
Later among the works it cites.
H.-J. Chang, S.-w. Yang, and H.-y. Lee, “Distilhubert: Speech representation learning by layer-wise distillation of hidden-unit bert,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2022, pp. 7087–7091
2022
Later among the works it cites.
N. Vaessen and D. A. Van Leeuwen, “Fine-tuning wav2vec2 for speaker recognition,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2022, pp. 7967–7971
2022
Later among the works it cites.
S. Kum, J. Lee, K. L. Kim, T. Kim, and J. Nam, “Pseudo-label transfer from frame-level to note-level in a teacher-student framework for singing transcription from polyphonic music,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2022, pp. 796–800
2022
Later among the works it cites.
Y. Yamamoto, J. Nam, and H. Terasawa, “Deformable CNN and Imbalance-Aware Feature Learning for Singing Technique Classification,” in Proceedings of the 21st Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2022, pp. 2778–2782
2022
Later among the works it cites.
Z. Chen, S. Chen, Y. Wu, Y. Qian, C. Wang, S. Liu, Y. Qian, and M. Zeng, “Large-scale self-supervised speech representation learning for automatic speaker verification,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2022, pp. 6147–6151
2022
Later among the works it cites.
T. Deng, E. Nakamura, and K. Yoshii, “End-to-end lyrics transcription informed by pitch and onset estimation,” in The 23rd International Society for Music Information Retrieval Conference (ISMIR) , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Gao, “Automatic lyrics transcription of polyphonic music,” Ph.D. dissertation, National University of Singapore (Singapore), 2022
2022
Later among the works it cites.
H. Suda, D. Saito, S. Fukayama, T. Nakano, and M. Goto, “Singer diarization for polyphonic music with unison singing,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1531–1545, 2022
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
G.-T. Lin, C.-L. Feng, W.-P. Huang, Y. Tseng, T.-H. Lin, C.-A. Li, H.-y. Lee, and N. G. Ward, “On the utility of self-supervised models for prosody-related tasks,” in 2022 IEEE Spoken Language Technology Workshop (SLT) , 2023, pp. 1104–1111
2023
Closest in time.
S. Rouard, F. Massa, and A. Défossez, “Hybrid transformers for music source separation,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2023
2023
Closest in time.
S. Yong, L. Su, and J. Nam, “A phoneme-informed neural network model for note-level singing transcription,” in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 2023
2023
Closest in time.