Fetching the paper…
Reading the bibliography…
We introduce LyricWhiz, a robust, multilingual, and zero-shot automatic lyrics transcription method achieving state-of-the-art performance on various lyrics transcription datasets, even in challenging genres such as rock and metal.
V. I. Levenshtein et al. , “Binary codes capable of correcting deletions, insertions, and reversals,” in Soviet physics doklady , vol. 10, no. 8. Soviet Union, 1966, pp. 707–710
1966
Earlier work this paper cites.
T. Hosoya, M. Suzuki, A. Ito, S. Makino, L. A. Smith, D. Bainbridge, and I. H. Witten, “Lyrics recognition from a singing voice based on finite state automaton for music information retrieval.” in ISMIR , 2005, pp. 532–535
2005
Earlier work this paper cites.
H. Fujihara, M. Goto, and J. Ogata, “Hyperlinking Lyrics: A method for creating hyperlinks between phrases in song lyrics.” in ISMIR , 2008, pp. 281–286
2008
Earlier work this paper cites.
J. K. Hansen and I. Fraunhofer, “Recognition of phonemes in a-cappella recordings using temporal patterns and mel frequency cepstral coefficients,” in 9th Sound and Music Computing Conference (SMC) , 2012, pp. 494–499
2012
Earlier work this paper cites.
P. Knees and M. Schedl, “A survey of music similarity and recommendation from music context data,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) , vol. 10, no. 1, pp. 1–21, 2013
2013
Earlier work this paper cites.
A. M. Kruspe and I. Fraunhofer, “Training phoneme models for singing with" songified" speech data.” in ISMIR , 2015, pp. 336–342
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
E. Çano and M. Morisio, “MoodyLyrics: A sentiment annotated lyrics dataset,” in Proceedings of the 2017 international conference on intelligent systems, metaheuristics & swarm intelligence , 2017, pp. 118–124
2017
Earlier work this paper cites.
G. Meseguer-Brocal, A. Cohen-Hadria, and G. Peeters, “DALI: a large dataset of synchronized audio, lyrics and notes, automatically created using teacher-student machine learning paradigm.” in 19th International Society for Music Information Retrieval Conference , ISMIR, Ed., September 2018
2018
Earlier work this paper cites.
C. Gupta, H. Li, and Y. Wang, “Automatic pronunciation evaluation of singing.” in Interspeech , 2018, pp. 1507–1511
2018
Earlier work this paper cites.
G. R. Dabike and J. Barker, “Automatic lyric transcription from karaoke vocal tracks: Resources and a baseline system.” in Interspeech , 2019, pp. 579–583
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. Stoller, S. Durand, and S. Ewert, “End-to-end lyrics alignment for polyphonic music using an audio-to-character recognition model,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 181–185
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Earlier work this paper cites.
C. Gupta, E. Yılmaz, and H. Li, “Automatic lyrics alignment and transcription in polyphonic music: Does background music help?” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 496–500
2020
Earlier work this paper cites.
G. Meseguer-Brocal, A. Cohen-Hadria, and G. Peeters, “Creating DALI, a large dataset of synchronized audio, lyrics, and notes,” Transactions of the International Society for Music Information Retrieval , vol. 3, no. 1, 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems , vol. 33, pp. 12 449–12 460, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
E. Demirel, S. Ahlbäck, and S. Dixon, “Automatic lyrics transcription using dilated convolutional neural networks with self-attention,” in 2020 International Joint Conference on Neural Networks (IJCNN) . IEEE, 2020, pp. 1–8
2020
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
2021
Cited alongside, same era.
K. Schulze-Forster, C. S. Doire, G. Richard, and R. Badeau, “Phoneme level lyrics alignment and text-informed singing voice separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2382–2395, 2021
2021
Cited alongside, same era.
S. Basak, S. Agarwal, S. Ganapathy, and N. Takahashi, “End-to-end lyrics recognition with voice to singing style transfer,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 266–270
2021
Cited alongside, same era.
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
E. Demirel, S. Ahlbäck, and S. Dixon, “Low resource audio-to-lyrics alignment from polyphonic music recordings,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 586–590
2021
Cited alongside, same era.
X. Gao, C. Gupta, and H. Li, “Genre-conditioned acoustic models for automatic lyrics transcription of polyphonic music,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 791–795
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys , vol. 55, no. 9, pp. 1–35, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.