doi:10.21437/Interspeech.2019-1925
B. Sharma, R. K. Das, H. Li, On the Importance of Audio-Source Separation for Singer Identification in Polyphonic Music, in: INTERSPEECH, 2019, pp. 2020–2024 · 1925
Earlier work this paper cites.
T. Saitou, N. Tsuji, M. Unoki, M. Akagi, Analysis of acoustic features affecting “singing-ness" and its application to singing-voice synthesis from speaking-voice, in: INTERSPEECH, 2004, pp. 1925–1928
1928
Earlier work this paper cites.
J. Sundberg, The level of the ‘singing formant’ and the source spectra of professional bass singers, Quarterly Progress and Status Report: STL-QPSR 11 (4) (1970) 21–39
1970
Earlier work this paper cites.
P. Oncley, Frequency, amplitude, and waveform modulation in the vocal vibrato, The Journal of the Acoustical Society of America 49 (1A) (1971) 136–136
1971
Earlier work this paper cites.
L. Rabiner, On the use of autocorrelation analysis for pitch detection, IEEE transactions on acoustics, speech, and signal processing 25 (1) (1977) 24–33
1977
Earlier work this paper cites.
J. M. Grey, J. W. Gordon, Perceptual effects of spectral modifications on musical timbres, The Journal of the Acoustical Society of America 63 (5) (1978) 1493–1500
1978
Earlier work this paper cites.
H. Fujisaki, Dynamic characteristics of voice fundamental frequency in speech and singing, in: The production of speech, Springer, 1983, pp. 39–55
1983
Earlier work this paper cites.
R. Leonard, A database for speaker-independent digit recognition, in: International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Vol. 9, IEEE, 1984, pp. 328–331
1984
Earlier work this paper cites.
J. Sundberg, T. D. Rossing, The science of singing voice, Vol. 87, ASA, 1990
1990
Earlier work this paper cites.
H. Sakoe, S. Chiba, A. Waibel, K. Lee, Dynamic programming algorithm optimization for spoken word recognition, Readings in speech recognition 159 (1990) 224
1990
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, J. McDaniel, SWITCHBOARD: Telephone speech corpus for research and development, in: International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Vol. 1, 1992, pp. 517–520
1992
Earlier work this paper cites.
D. B. Paul, J. M. Baker, The design for the Wall Street Journal-based CSR corpus, in: Proceedings of the workshop on Speech and Natural Language, Association for Computational Linguistics, 1992, pp. 357–362
1992
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, N. LDahlgren, V. Zue, TIMIT acoustic phonetic continuous speech corpus, Linguistic Data Consortium (1993)
1993
Earlier work this paper cites.
J. Sundberg, I. R. Titze, R. Scherer, Phonatory control in male singing: A study of the effects of subglottal pressure, fundamental frequency, and mode of phonation on the voice source, Journal of Voice 7 (1) (1993) 15 – 29
1993
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald, CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit, University of Edinburgh. The Centre for Speech Technology Research (CSTR) doi:http://dx.doi.org/10.7488/ds/1994
1994
Earlier work this paper cites.
S. Werner, E. Keller, Fundamentals of speech synthesis and speech recognition: Basic concepts, state of the art, and future challenges, chapter prosodic aspects of speech, E. Keller (1994) 23–40
1994
Earlier work this paper cites.
H.-G. Hirsch, D. Pearce, The Aurora experimental framework for the performance evaluation of speech recognition systems under noisy conditions, in: ASR2000-Automatic Speech Recognition: Challenges for the new Millenium ISCA Tutorial and Research Workshop (ITRW), 2000
2000
Earlier work this paper cites.
K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, T. Kitamura, Speech parameter generation algorithms for HMM-based speech synthesis, in: International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Vol. 3, 2000, pp. 1315–1318
2000
Earlier work this paper cites.
G. Tzanetakis, A. Ermolinskyi, P. Cook, Pitch histograms in audio and symbolic music information retrieval, Journal of New Music Research 32 (2) (2003) 143–152
2003
Earlier work this paper cites.
J. Kominek, A. W. Black, The CMU Arctic speech databases, in: Fifth ISCA workshop on speech synthesis (SSW5), 2004, pp. 223–224
2004
Earlier work this paper cites.
P. Knees, E. Pampalk, G. Widmer, Artist classification with web-based data, in: International Society for Music Information Retrieval Conference (ISMIR), 2004, pp. 517–524
2004
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, et al., The AMI meeting corpus: A pre-announcement, in: International workshop on machine learning for multimodal interaction, Springer, 2005, pp. 28–39
2005
Earlier work this paper cites.
M. Schedl, P. Knees, G. Widmer, Investigating web-based approaches to revealing prototypical music artists in genre taxonomies, in: 1st International Conference on Digital Information Management, 2006, pp. 519–524
2006
Earlier work this paper cites.
C. McKay, D. McEnnis, I. Fujinaga, A large publicly accessible prototype audio database for music research, in: International Society for Music Information Retrieval Conference (ISMIR), 2006, pp. 160–163
2006
Earlier work this paper cites.
Y. Ohishi, M. Goto, K. Itou, K. Takeda, On the human capability and acoustic cues for discriminating the singing and the speaking voices, 2006, pp. 1831–1837
2006
Earlier work this paper cites.
D. P. Ellis, Classifying music audio with timbral and chroma features, in: International Society for Music Information Retrieval Conference (ISMIR), 2007, pp. 339–340
2007
Earlier work this paper cites.
A. F. Martin, C. S. Greenberg, NIST 2008 speaker recognition evaluation: Performance across telephone and room microphone channels, in: INTERSPEECH, 2009, pp. 2579–2582
2009
Earlier work this paper cites.
C. L. Hsu, J. S. R. Jang, On the improvement of singing voice separation for monaural recordings using the MIR-1K dataset, IEEE Transactions on Audio, Speech, and Language Processing 18 (2) (2009) 310–319
2009
Earlier work this paper cites.