Fetching the paper…
Reading the bibliography…
Automatic transcription of monophonic/polyphonic music is a challenging task due to the lack of availability of large amounts of transcribed data.
“Lyrics recognition from a singing voice based on finite state automaton for music information retrieval.,”
T. Hosoya, M. Suzuki, A. Ito, S. Makino, L. A. Smith, D. Bainbridge, and I. H. Witten, · 2005
Earlier work this paper cites.
“Songs and emotions: are lyrics and melodies equal partners?,”
S. O. Ali and Z. F. Peynircioğlu, · 2006
Earlier work this paper cites.
“Phoneme recognition in popular music.,”
M. Gruhne, C. Dittmar, and K. Schmidt, · 2007
Earlier work this paper cites.
“Speech-to-singing synthesis: Converting speaking voices to singing voices by controlling acoustic features unique to singing voices,”
T. Saitou, M. Goto, M. Unoki, and M. Akagi, · 2007
Earlier work this paper cites.
“Hyperlinking lyrics: A method for creating hyperlinks between phrases in song lyrics.,”
H. Fujihara, M. Goto, and J. Ogata, · 2008
Earlier work this paper cites.
“Fast and reliable f0 estimation method based on the period extraction of vocal fold vibration of singing voice and speech,”
M. Morise, H. Kawahara, and H. Katayose, · 2009
Earlier work this paper cites.
“Automatic recognition of lyrics in singing,”
A. Mesaros and T. Virtanen, · 2010
Earlier work this paper cites.
“Lyricsynchronizer: Automatic synchronization system between musical audio signals and lyrics,”
H. Fujihara, M. Goto, J. Ogata, and H. G. Okuno, · 2011
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al., · 2012
Earlier work this paper cites.
“Platinum: A method to extract excitation signals for voice synthesis system,”
M. Morise, · 2012
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
A. Graves and N. Jaitly, · 2014
Earlier work this paper cites.
“The efficacy of singing in foreign-language learning,”
A. J. Good, F. A. Russo, and J. Sullivan, · 2015
Cited alongside, same era.
“Cheaptrick, a spectral envelope estimator for high-quality speech synthesis,”
M. Morise, · 2015
Cited alongside, same era.
“Librispeech: An asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
“Audio augmentation for speech recognition,”
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, · 2016
Cited alongside, same era.
“World: A vocoder-based high-quality speech synthesis system for real-time applications,”
M. MORISE, F. YOKOMORI, and K. OZAWA, · 2016
“A comparative study on transformer vs rnn in speech applications,”
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang, et al., · 2019
Later among the works it cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Later among the works it cites.
“The dali dataset,” Feb. 2019
G. Meseguer Brocal, · 2019
Later among the works it cites.
“Automatic lyrics-to-audio alignment on polyphonic music using singing-adapted acoustic models,”
B. Sharma, C. Gupta, H. Li, and Y. Wang, · 2019
Later among the works it cites.
“End-to-end lyrics alignment for polyphonic music using an audio-to-character recognition model,”
D. Stoller, S. Durand, and S. Ewert, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Knowledge-based probabilistic modeling for tracking lyrics in music audio signals
G. Dzhambazov et al., · 2017
Cited alongside, same era.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, · 2017
Cited alongside, same era.
“Semi-supervised lyrics and solo-singing alignment,”
C. Gupta, R. Tong, H. Li, and Y. Wang, · 2018
Cited alongside, same era.
“Espnet: End-to-end speech processing toolkit,” 2018
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, · 2018
Cited alongside, same era.
“Acoustic modeling for automatic lyrics-to-audio alignment,”
C. Gupta, E. Yılmaz, and H. Li, · 2019
Cited alongside, same era.
“Acoustic modeling for automatic lyrics-to-audio alignment,”
C. Gupta, E. Yilmaz, and H. Li, · 2019
Later among the works it cites.
“Automatic lyrics alignment and transcription in polyphonic music: Does background music help?,”
C. Gupta, E. Yılmaz, and H. Li, · 2020
Later among the works it cites.
“Speech-to-singing conversion in an encoder-decoder framework,”
J. Parekh, P. Rao, and Y.-H. Yang, · 2020
Later among the works it cites.
“D3net: Densely connected multidilated densenet for music source separation,”
N. Takahashi and yuki Mitsufuji, · 2020
Later among the works it cites.
“Improving Voice Separation by Incorporating End-To-End Speech Recognition,”
N. Takahashi, M. K. Singh, S. Basak, P. Sudarsanam, S. Ganapathy, and Y. Mitsufuji, · 2020
Later among the works it cites.