Fetching the paper…
Reading the bibliography…
Converting time domain waveforms to frequency domain spectrograms is typically considered to be a prepossessing step done before model training.
B. Murauer and G. Specht, “Detecting music genre using extreme gradient boosting,” in Companion Proceedings of the The Web Conference 2018 , 2018, pp. 1923–1927
1927
Earlier work this paper cites.
S. S. Stevens, J. E. Volkmann, and E. B. Newman, “A scale for the measurement of the psychological magnitude pitch,” The Journal of the Acoustical Society of America , 1937
1937
Earlier work this paper cites.
S. S. Stevens and J. Volkmann, “The relation of pitch to frequency: A revised scale,” The American Journal of Psychology , vol. 53, no. 3, pp. 329–353, 1940. [Online]. Available: http://www.jstor.org/stable/1417526
1940
Earlier work this paper cites.
G. Fant, Analys av de svenska konsonantljuden: talets allmänna svängningsstruktur . LM Ericsson, 1949
1949
Earlier work this paper cites.
W. Koening, “A new frequency scala for acoustic measurements,” Bell Lab Rec. , pp. 299–301, 1949
1949
Earlier work this paper cites.
J. Youngberg and S. Boll, “Constant-Q signal analysis and synthesis,” in International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , vol. 3. IEEE, 1978, pp. 375–378
1978
Earlier work this paper cites.
S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE transactions on acoustics, speech, and signal processing , vol. 28, no. 4, pp. 357–366, 1980
1980
Earlier work this paper cites.
S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE transactions on acoustics, speech, and signal processing , vol. 28, no. 4, pp. 357–366, 1980
1980
Earlier work this paper cites.
S. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences,” IEEE transactions on acoustics, speech, and signal processing , vol. 28, no. 4, pp. 357–366, 1980
1980
Earlier work this paper cites.
Z. Wang, “Fast algorithms for the discrete w transform and for the discrete fourier transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 32, no. 4, pp. 803–816, 1984
1984
Earlier work this paper cites.
S. H. Nawab and T. F. Quatieri, “Advanced topics in signal processing,” J. S. Lim and A. V. Oppenheim, Eds. Upper Saddle River, NJ, USA: Prentice-Hall, Inc., 1987, ch. Short-time Fourier Transform, pp. 289–337. [Online]. Available: http://dl.acm.org/citation.cfm?id=42739.42745
1987
Earlier work this paper cites.
D. O’shaughnessy, Speech communication: human and machine . Universities press, 1987
1987
Earlier work this paper cites.
G. Haines and A. G. Jones, “Logarithmic fourier transformation,” Geophysical Journal International , vol. 92, no. 1, pp. 171–178, 1988
1988
Earlier work this paper cites.
M. J. Palakal and M. J. Zoran, “Feature extraction from speech spectrograms using multi-layered network models,” in International Workshop on Tools for Artificial Intelligence . IEEE, 1989, pp. 224–230
1989
Earlier work this paper cites.
K. Hatazaki, Y. Komori, T. Kawabata, and K. Shikano, “Phoneme segmentation using spectrogram reading knowledge,” in International Conference on Acoustics, Speech, and Signal Processing (ICASSP) . IEEE, 1989, pp. 393–396
1989
Earlier work this paper cites.
A. V. Oppenheim and R. W. Schafer, “Fourier transform theorems,” in Discrete-time signal processing , 1989, p. 60
1989
Earlier work this paper cites.
J. C. Brown, “Calculation of a constant q spectral transform,” The Journal of the Acoustical Society of America , vol. 89, no. 1, pp. 425–434, 1991
1991
Earlier work this paper cites.
J. C. Brown and M. S. Puckette, “An efficient algorithm for the calculation of a constant q transform,” The Journal of the Acoustical Society of America , vol. 92, no. 5, pp. 2698–2701, 1992
1992
Earlier work this paper cites.
S. Furui, “Speaker-independent isolated word recognition based on emphasized spectral dynamics,” in International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , vol. 11. IEEE, 1986, pp. 1991–1994
1994
Earlier work this paper cites.
Y. LeCun, Y. Bengio et al. , “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks , vol. 3361, no. 10, p. 1995, 1995
1995
Earlier work this paper cites.
M. C. Recchione and A. P. Russo, “Feedforward neural network system for the detection and characterization of sonar signals with characteristic spectrogram textures,” Mar. 26 1996, uS Patent 5,502,688
1996
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
M. Slaney, “A matlab toolbox for auditory modeling work,” Interval Research Corporation , 1998
1998
Earlier work this paper cites.
S. Umesh, L. Cohen, and D. Nelson, “Fitting the mel scale,” in 1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings. ICASSP99 (Cat. No. 99CH36258) , vol. 1, 1999, pp. 217–220
1999
Cited alongside, same era.
S. Young, G. Evermann, M. Gales, T. Hain, D. Kershaw, X. Liu, G. Moore, J. Odell, D. Ollason, D. Povey et al. , “The htk book,” Cambridge university engineering department , vol. 3, p. 175, 2002
2002
Cited alongside, same era.
J. O. S. III, Introduction to Digital Filters with Audio Applications . Center for Computer Research in Music and Acoustics (CCRMA), Stanford University, 2007
2007
Cited alongside, same era.
C. Schörkhuber and A. Klapuri, “Constant-q transform toolbox for music processing,” in 7th Sound and Music Computing Conference, Barcelona, Spain , 2010, pp. 3–64
2010
Cited alongside, same era.
2018
Later among the works it cites.
J. Thickstun, Z. Harchaoui, D. P. Foster, and S. M. Kakade, “Invariances and data augmentation for supervised music transcription,” in International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2018
2018
Later among the works it cites.
T. Kim, J. Lee, and J. Nam, “Comparison and analysis of samplecnn architectures for audio classification,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, no. 2, pp. 285–297, 2019
2019
Closest in time.
H. Purwins, B. Li, T. Virtanen, J. Schlx00FCter, S.-Y. Chang, and T. Sainath, “Deep learning for audio signal processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 13, pp. 206–219, 2019
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Emiya, N. Bertin, B. David, and R. Badeau, “Maps-a piano database for multipitch estimation and automatic transcription of music,” 2010
2010
Cited alongside, same era.
L. R. Rabiner and R. W. Schafer, Theory and applications of digital speech processing . Pearson Upper Saddle River, NJ, 2011, vol. 64
2011
Cited alongside, same era.
S. S. Rajput and D. S. Bhadauria, “Implementation of fir filter using efficient window function and its application in filtering a speech signal,” International Journal of Electrical, Electronics and Mechanical Controls , vol. 1, no. 1, 2012
2012
Cited alongside, same era.
E. Benetos, S. Dixon, D. Giannoulis, H. Kirchhoff, and A. Klapuri, “Automatic music transcription: challenges and future directions,” Journal of Intelligent Information Systems , vol. 41, no. 3, pp. 407–434, 2013
2013
Cited alongside, same era.
B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and music signal analysis in python,” in Proceedings of the 14th python in science conference , vol. 8, 2015
2015
Cited alongside, same era.
M. Abadi, A. Agarwal, M. Abadi et al. , “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: http://tensorflow.org/
2015
Cited alongside, same era.
2016
Cited alongside, same era.
R. Kelz, M. Dorfer, F. Korzeniowski, S. Böck, A. Arzt, and G. Widmer, “On the potential of simple framewise approaches to piano transcription,” in Proceedings of the 17th International Society for Music Information Retrieval Conference, ISMIR, New York City, United States , 2016, pp. 475–481
2016
Cited alongside, same era.
C. Kim, S. Kim, K. Kim, M. Kumar, J. Kim, K. Lee, C. Han, A. Garg, E. Kim, M. Shin, S. Singh, L. Heck, and D. Gowda, “End-to-end training of a large vocabulary end-to-end speech recognition system,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2019, pp. 562–569
2019
Closest in time.
J. Zhao, X. Mao, and L. Chen, “Speech emotion recognition using deep 1d & 2d cnn lstm networks,” Biomedical Signal Processing and Control , vol. 47, pp. 312–323, 2019
2019
Closest in time.
A. Tjandra, S. Sakti, and S. Nakamura, “Speech-to-speech translation between untranscribed unknown languages,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2019, pp. 593–600
2019
Closest in time.
S. A. Shahriyar, M. Akhand, N. Siddique, and T. Shimamura, “Speech enhancement using convolutional denoising autoencoder,” in International Conference on Electrical, Computer and Communication Engineering (ECCE) . IEEE, 2019, pp. 1–5
2019
Closest in time.
J. Choi, J. Lee, J. Park, and J. Nam, “Zero-shot learning for audio-based music classification and tagging,” in Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR, Delft, The Netherlands , A. Flexer, G. Peeters, J. Urbano, and A. Volk, Eds., 2019, pp. 67–74
2019
Closest in time.
G. Doras and G. Peeters, “Cover detection using dominant melody embeddings,” in Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR, Delft, The Netherlands , A. Flexer, G. Peeters, J. Urbano, and A. Volk, Eds., 2019, pp. 107–114
2019
Closest in time.
G. Doras, P. Esling, and G. Peeters, “On the use of u-net for dominant melody estimation in polyphonic music,” in 2019 International Workshop on Multilayer Music Representation and Processing (MMRP) . IEEE, 2019, pp. 66–70
2019
Closest in time.
R. Kelz and G. Widmer, “Towards interpretable polyphonic transcription with invertible neural networks,” in Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR, Delft, The Netherlands , 2019, pp. 376–383
2019
Closest in time.
K. Qu. (2019) some problems when installing torchaudio on mac. https://stackoverflow.com/q/56659166/
2019
Closest in time.
L. Ericson. (2019) How to install torch audio on windows 10 conda? https://stackoverflow.com/q/54872876
2019
Closest in time.
K. W. Cheuk, K. Agres, and D. Herremans, “nnAudio: A pytorch audio processing tool using 1D convolution neural networks,” in ISMIR–Late breaking demo , Delft, The Netherlands, 2019
2019
Closest in time.
B. Balamurali, K. E. Lin, S. Lui, J.-M. Chen, and D. Herremans, “Toward robust audio spoofing detection: A detailed comparison of traditional and learned features,” IEEE Access , vol. 7, pp. 84 229–84 241, 2019
2019
Closest in time.
A. V. Gayer, Y. S. Chernyshova, and A. V. Sheshkus, “Effective real-time augmentation of training dataset for the neural networks learning,” in International Conference on Machine Vision , 2019
2019
Closest in time.
A. Holzapfel and E. Benetos, “Automatic music transcription and ethnomusicology: a user study,” in Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR, Delft, The Netherlands, November , A. Flexer, G. Peeters, J. Urbano, and A. Volk, Eds., 2019, pp. 678–684
2019
Closest in time.
K. W. E. Lin, B. Balamurali, E. Koh, S. Lui, and D. Herremans, “Singing voice separation using a deep convolutional neural network trained by ideal binary mask and cross entropy,” Neural Computing and Applications , vol. 32, no. 4, pp. 1037–1050, 2020
2020
Closest in time.
Y.-J. Luo, C.-C. Hsu, K. Agres, and D. Herremans, “Singing voice conversion with disentangled representations of singer and vocal technique using variational autoencoders,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 3277–3281
2020
Closest in time.
2020
Closest in time.