Fetching the paper…
Reading the bibliography…
Automatic lyrics transcription (ALT), which can be regarded as automatic speech recognition (ASR) on singing voice, is an interesting and practical topic in academia and industry.
Specaugment: A simple data augmentation method for automatic speech recognition
Park, D. S.; Chan, W.; Zhang, Y.; Chiu, C.-C.; Zoph, B.; Cubuk, E. D.; and Le, Q. V. 2019 · 1904
Earlier work this paper cites.
Meseguer-Brocal, G.; Cohen-Hadria, A.; and Peeters, G. 2019 · 1906
Earlier work this paper cites.
Low-delay singing voice alignment to text
Loscos, A.; Cano, P.; and Bonada, J. 1999 · 1999
Earlier work this paper cites.
Conformer: Convolution-augmented Transformer for Speech Recognition
Gulati, A.; Qin, J.; Chiu, C.-C.; Parmar, N.; Zhang, Y.; Yu, J.; Han, W.; Wang, S.; Zhang, Z.; Wu, Y.; et al. 2020 · 2005
Earlier work this paper cites.
Lyrics Recognition from a Singing Voice Based on Finite State Automaton for Music Information Retrieval
Hosoya, T.; Suzuki, M.; Ito, A.; Makino, S.; Smith, L. A.; Bainbridge, D.; and Witten, I. H. 2005 · 2005
Earlier work this paper cites.
Songs and emotions: are lyrics and melodies equal partners?
Ali, S. O.; and Peynircioğlu, Z. F. 2006 · 2006
Earlier work this paper cites.
Automatic synchronization between lyrics and music CD recordings based on Viterbi alignment of segregated vocal signals
Fujihara, H.; Goto, M.; Ogata, J.; Komatani, K.; Ogata, T.; and Okuno, H. G. 2006 · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A.; Fernández, S.; Gomez, F.; and Schmidhuber, J. 2006 · 2006
Earlier work this paper cites.
UWSpeech: Speech to Speech Translation for Unwritten Languages
Zhang, C.; Tan, X.; Ren, Y.; Qin, T.; Zhang, K.; and Liu, T.-Y. 2020b · 2006
Earlier work this paper cites.
Automatic alignment of music audio and lyrics
Mesaros, A.; and Virtanen, T. 2008 · 2008
Earlier work this paper cites.
Adaptation of a speech recognizer for singing voice
Mesaros, A.; and Virtanen, T. 2009 · 2009
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Graves, A. 2012 · 2012
Earlier work this paper cites.
Denoising Text to Speech with Frame-Level Noise Modeling
Zhang, C.; Ren, Y.; Tan, X.; Liu, J.; Zhang, K.; Qin, T.; Zhao, S.; and Liu, T.-Y. 2020a · 2012
Cited alongside, same era.
Singing voice identification and lyrics transcription for music information retrieval invited paper
Mesaros, A. 2013 · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Cited alongside, same era.
Keyword Spotting in A-capella Singing
Kruspe, A. M.; and Fraunhofer, I. 2014 · 2014
Cited alongside, same era.
Training Phoneme Models for Singing with” Songified” Speech Data
Kruspe, A. M.; and Fraunhofer, I. 2015 · 2015
Cited alongside, same era.
Automatic Pronunciation Evaluation of Singing
Gupta, C.; Li, H.; and Wang, Y. 2018 · 2018
Later among the works it cites.
Transcribing lyrics from commercial song audio: the first step towards singing content processing
Tsai, C.-P.; Tuan, Y.-L.; and Lee, L.-s. 2018 · 2018
Later among the works it cites.
Espnet: End-to-end speech processing toolkit
Watanabe, S.; Hori, T.; Karita, S.; Hayashi, T.; Nishitoba, J.; Unno, Y.; Soplin, N. E. Y.; Heymann, J.; Wiesner, M.; Chen, N.; et al. 2018 · 2018
Later among the works it cites.
Automatic Lyric Transcription from Karaoke Vocal Tracks: Resources and a Baseline System
Dabike, G. R.; and Barker, J. 2019 · 2019
Later among the works it cites.
Automatic Lyrics Transcription using Dilated Convolutional Neural Networks with Self-Attention
Demirel, E.; Ahlbäck, S.; and Dixon, S. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Panayotov, V.; Chen, G.; Povey, D.; and Khudanpur, S. 2015 · 2015
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Amodei, D.; Ananthanarayanan, S.; Anubhai, R.; Bai, J.; Battenberg, E.; Case, C.; Casper, J.; Catanzaro, B.; Cheng, Q.; Chen, G.; et al. 2016 · 2016
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, W.; Jaitly, N.; Le, Q.; and Vinyals, O. 2016 · 2016
Cited alongside, same era.
Bootstrapping a System for Phoneme Recognition and Keyword Spotting in Unaccompanied Singing
Kruspe, A. M.; and Fraunhofer, I. 2016 · 2016
Cited alongside, same era.
WORLD: a vocoder-based high-quality speech synthesis system for real-time applications
Morise, M.; Yokomori, F.; and Ozawa, K. 2016 · 2016
Cited alongside, same era.
Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi
McAuliffe, M.; Socolof, M.; Mihuc, S.; Wagner, M.; and Sonderegger, M. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Automatic Lyrics Alignment and Transcription in Polyphonic Music: Does Background Music Help?
Gupta, C.; Yılmaz, E.; and Li, H. 2020 · 2020
Later among the works it cites.
Does English Have Useful Syllable Division Patterns?
Kearns, D. M. 2020 · 2020
Later among the works it cites.
Towards fast and accurate streaming end-to-end ASR
Li, B.; Chang, S.-y.; Sainath, T. N.; Pang, R.; He, Y.; Strohman, T.; and Wu, Y. 2020 · 2020
Later among the works it cites.
Semantic Mask for Transformer Based End-to-End Speech Recognition
Wang, C.; Wu, Y.; Du, Y.; Li, J.; Liu, S.; Lu, L.; Ren, S.; Ye, G.; Zhao, S.; and Zhou, M. 2020 · 2020
Later among the works it cites.
Lrspeech: Extremely low-resource speech synthesis and recognition
Xu, J.; Tan, X.; Ren, Y.; Qin, T.; Li, J.; Zhao, S.; and Liu, T.-Y. 2020 · 2020
Later among the works it cites.
Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss
Zhang, Q.; Lu, H.; Sak, H.; Tripathi, A.; McDermott, E.; Koo, S.; and Kumar, S. 2020c · 2020
Later among the works it cites.
End-to-end lyrics Recognition with Voice to Singing Style Transfer
Basak, S.; Agarwal, S.; Ganapathy, S.; and Takahashi, N. 2021 · 2021
Closest in time.