Fetching the paper…
Reading the bibliography…
Given the recent surge in developments of deep learning, this article provides a review of the state-of-the-art deep learning techniques for audio signal processing.
F. Rosenblatt, “The perceptron: A probabilistic model for information storage and organization in the brain,” Psychological Review , vol. 65, no. 6, p. 386, 1958
1958
Earlier work this paper cites.
S. Davis and P. Mermelstein, “Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences,” IEEE Transactions on ASSP , vol. 28, no. 4, pp. 357 – 366, 1980
1980
Earlier work this paper cites.
J. S. L. Daniel W. Griffin, “Signal Estimation from Modified Short-Time Fourier Transform,” IEEE Transactions on ASSP , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature , vol. 323, no. 6088, p. 533, 1986
1986
Earlier work this paper cites.
S. Furui, “Speaker-independent isolated word recognition based on emphasized spectral dynamics,” in ICASSP , 1986
1986
Earlier work this paper cites.
Y. LeCun, B. Boser et al. , “Backpropagation applied to handwritten zip code recognition,” Neural Computation , vol. 1, no. 4, pp. 541–551, 1989
1989
Earlier work this paper cites.
M. Holschneider, R. Kronland-Martinet, J. Morlet, and P. Tchamitchian, “Wavelets, time-frequency methods and phase space,” Springer , pp. 289–297, 1989
1989
Earlier work this paper cites.
J. L. Elman, “Finding structure in time,” Cognitive science , vol. 14, no. 2, pp. 179–211, 1990
1990
Earlier work this paper cites.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,” in Readings in speech recognition . Elsevier, 1990, pp. 393–404
1990
Earlier work this paper cites.
H. A. Bourlard and N. Morgan, Connectionist speech recognition: a hybrid approach . Springer Science & Business Media, 1994, vol. 247
1994
Earlier work this paper cites.
T. Robinson, M. Hochberg, and S. Renals, “The use of recurrent neural networks in continuous speech recognition,” in Automatic speech and speaker recognition . Springer, 1996, pp. 233–258
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
H. Purwins, B. Blankertz, and K. Obermayer, “A new method for tracking modulations in tonal music in audio data format,” in IJCNN , 2000
2000
Earlier work this paper cites.
B.-H. Juang and L. R. Rabiner, “Automatic speech recognition–a brief history of the technology development,” Georgia Institute of Technology. Atlanta Rutgers University and the University of California. Santa Barbara , vol. 1, p. 67, 2005
2005
Earlier work this paper cites.
A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber, “Connectionist Temporal Classification: Labeling Unsegmented Seuqnece Data with Recurrent Neural Networks,” in ICML , 2006
2006
Earlier work this paper cites.
A. Lacoste and D. Eck, “A supervised classification algorithm for note onset detection,” EURASIP Journal on Advances in Signal Processing , vol. 2007, no. 1, p. 043745, 2006
2006
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Transactions on ASLP , vol. 14, 2006
2006
Earlier work this paper cites.
A. Graves, S. Fernandez, and J. Schmidhuber, “Multi-Dimensional Recurrent Neural Networks,” in ICANN , 2007
2007
Earlier work this paper cites.
J. Chen, J. Benesty, Y. A. Huang, and E. J. Diethorn, “Fundamentals of noise reduction,” in Springer Handbook of Speech Processing . Springer, 2008, pp. 843–872
2008
Earlier work this paper cites.
S. Che, M. Boyer, J. Meng, D. Tarjan, J. W. Sheaffer, and K. Skadron, “A performance study of general-purpose applications on graphics processors using cuda,” Journal of parallel and distributed computing , vol. 68, no. 10, pp. 1370–1380, 2008
2008
Earlier work this paper cites.
A.-R. Mohamed, G. Dahl, and G. Hinton, “Deep belief networks for phone recognition,” in NIPS workshop on deep learning for speech recognition and related applications , vol. 1, no. 9, 2009, pp. 39–47
2009
Earlier work this paper cites.
S. Watanabe, Algebraic geometry and statistical learning theory . Cambridge University Press, 2009, vol. 25
2009
Earlier work this paper cites.
F. Eyben, S. Böck, B. Schuller, and A. Graves, “Universal Onset Detection with Bidirectional Long Short-Term Memory Neural Networks,” in ISMIR , 2010
2010
Earlier work this paper cites.
N. Jaitly and G. Hinton, “Learning a Better Representation of Speech Soundwaves using Restricted Boltzmann Machines,” in ICASSP , 2011
2011
Earlier work this paper cites.
G. Hinton, L. Deng et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NIPS , 2012
2012
Earlier work this paper cites.
A. Mohamed, G. Hinton, and G. Penn, “Understanding how Deep Belief Networks Perform Acoustic Modelling,” in ICASSP , 2012
2012
Earlier work this paper cites.
A. Graves, “Sequence Transduction with Recurrent Neural Networks,” ICML Representation Learning Workshop , 2012
2012
Earlier work this paper cites.
E. J. Humphrey and J. P. Bello, “Rethinking Automatic Chord Recognition with Convolutional Neural Networks,” in ICMLA , 2012
2012
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury et al. , “Learning filter banks within a deep neural network framework,” in ASRU , 2013
2013
Earlier work this paper cites.
N. Jaitly and G. E. Hinton, “Vocal tract length perturbation (VTLP) improves speech recognition,” in ICML Workshop on Deep Learning for Audio, Speech, and Language Processing , vol. 117, 2013
2013
Earlier work this paper cites.
N. Kanda, R. Takeda, and Y. Obuchi, “Elastic spectral distortion for low resource speech recognition with deep neural networks,” in ASRU , 2013
2013
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury, A. Mohamed, G. Dahl, G. Saon, H. Soltau, T. Beran, A. Aravkin, and B. Ramabhadran, “Improvements to Deep Convolutional Neural Networks for LVCSR,” in ASRU , 2013
2013
Earlier work this paper cites.
A. Narayanan and D. Wang, “Ideal ratio mask estimation using deep neural networks for robust speech recognition,” in ICASSP , 2013
2013
Earlier work this paper cites.
X. Lu, Y. Tsao, S. Matsuda, and C. Hori, “Speech enhancement based on deep denoising autoencoder.” in Interspeech , 2013
2013
Earlier work this paper cites.
J. Schlüter and S. Böck, “Improved Musical Onset Detection with Convolutional Neural Networks,” in ICASSP , 2014
2014
Earlier work this paper cites.
J. Chen, Y. Wang, and D. Wang, “A feature study for classification-based speech separation at low signal-to-noise ratios,” IEEE/ACM TASLP , vol. 22, no. 12, pp. 1993–2002, 2014
2014
Earlier work this paper cites.
D. Palaz, R. Collobert, and M. Doss, “Estimating Phoneme Class Conditional Probabilities From Raw Speech Signal using Convolutional Neural Networks,” in Interspeech , 2014
2014
Earlier work this paper cites.
Z. Tüske, P. Golik, R. Schlüter, and H. Ney, “Acoustic Modeling with Deep Neural Networks using Raw Time Signal for LVCSR,” in Interspeech , 2014
2014
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long Short-Term Memory Recurrent Neural Network Architectures for Large Scale Acoustic Modeling,” in Interspeech , 2014
2014
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards End-to-End Speech Recognition with Recurrent Neural Networks,” in ICML , 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, and et al
2014
Earlier work this paper cites.
H. Sak, O. Vinyals, G. Heigold, A. Senior, E. McDermott, R. Monga, and M. Mao, “Sequence discriminative distributed training of long short-term memory recurrent neural networks,” in Interspeech , 2014
2014
Earlier work this paper cites.
I. Lopez-Moreno, J. Gonzalez-Dominguez et al. , “Automatic language identification using deep neural networks,” in ICASSP , 2014
2014
Earlier work this paper cites.
B. McFee and D. P. W. Ellis, “Better beat tracking through robust onset aggregation,” in ICASSP , 2014
2014
Earlier work this paper cites.
K. Ullrich, J. Schlüter, and T. Grill, “Boundary Detection in Music Structure Analysis using Convolutional Neural Networks,” in ISMIR , 2014
2014
Earlier work this paper cites.
S. Dieleman and B. Schrauwen, “End-to-end learning for music audio,” in ICASSP , 2014
2014
Earlier work this paper cites.
Y. Wang, A. Narayanan, and D. Wang, “On Training Targets for Supervised Speech Separation,” IEEE/ACM Transactions on ASLP , vol. 22, no. 12, pp. 1849–1858, 2014
2014
Earlier work this paper cites.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “An experimental study on speech enhancement based on deep neural networks,” IEEE Signal Processing Letters , vol. 21, no. 1, pp. 65–68, 2014
2014
Earlier work this paper cites.
X. Feng, Y. Zhang, and J. Glass, “Speech feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognition,” in ICASSP , 2014
2014
Earlier work this paper cites.
B. Li and K. C. Sim, “A spectral masking approach to noise-robust speech recognition using deep neural networks,” IEEE/ACM Transactions on ASLP , vol. 22, no. 8, pp. 1296–1305, 2014
2014
Earlier work this paper cites.
Y. Hoshen, R. Weiss, and K. Wilson, “Speech Acoustic Modeling from Raw Multichannel Waveforms,” in ICASSP , 2015
2015
Cited alongside, same era.
T. N. Sainath, R. J. Weiss, K. W. Wilson, A. Senior, and O. Vinyals, “Learning the Speech Front-end with Raw Waveform CLDNNs,” in Interspeech , 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in MICCAI , 2015, pp. 234–241
2015
Cited alongside, same era.
M. Mimura, S. Sakai, and T. Kawahara, “Cross-domain speech recognition using nonparallel corpora with cycle-consistent adversarial networks,” in ASRU , 2017
2017
Later among the works it cites.
M. Cuturi and M. Blondel, “Soft-DTW: a Differentiable Loss Function for Time-Series,” in ICML , 2017
2017
Later among the works it cites.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in ICML , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
J. Thickstun, Z. Harchaoui, and S. M. Kakade, “Learning features of music from scratch,” in ICLR , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
J. Li, A. Mohamed, G. Zweig, and Y. Gong, “LSTM Time and Frequency Recurrence for Automatic Speech Recognition,” in ASRU , 2015
2015
Cited alongside, same era.
T. N. Sainath, O. Vinyals, A. Senior, and H. Sak, “Convolutional, Long Short-Term Memory, Fully Connected Deep Neural Networks,” in ICASSP , 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-Based Models for Speech Recognition,” in NIPS , 2015
2015
Cited alongside, same era.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in ICLR , 2015
2015
Cited alongside, same era.
J. Schlüter and T. Grill, “Exploring data augmentation for improved singing voice detection with neural networks,” in ISMIR , 2015
2015
Cited alongside, same era.
B. McFee, E. J. Humphrey, and J. P. Bello, “A software framework for musical data augmentation,” in ISMIR , 2015, pp. 248–254
2015
Cited alongside, same era.
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. Sainath, and M. Bacchiani, “Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home,” in Interspeech , 2017
2017
Later among the works it cites.
S.-Y. Chang, B. Li, T. N. Sainath, G. Simko, and C. Parada, “Endpoint detection using grid long short-term memory networks for streaming speech recognition,” in Interspeech , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Durand, J. P. Bello, B. David, and G. Richard, “Tracking using an ensemble of convolutional networks,” IEEE/ACM Transactions on ASLP , vol. 25, no. 1, pp. 76–89, Jan. 2017
2017
Later among the works it cites.
B. McFee and J. P. Bello, “Structured Training for Large-Vocabulary Chord Recognition,” in ISMIR , 2017, pp. 188–194
2017
Later among the works it cites.
J. Lee, J. Park, K. L. Kim, and J. Nam, “Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms,” in Proc. Sound Music Comput. Conf. , 2017, pp. 220–226
2017
Later among the works it cites.
S. Chakrabarty and E. A. P. Habets, “Multi-speaker localization using convolutional neural network trained with noise,” in NIPS Workshop on Machine Learning for Audio Processing , 2017
2017
Later among the works it cites.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in ICASSP , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, and others, “Tacotron: Towards end-to-end speech synthesis,” in In Proc. Interspeech , 2017, pp. 4006–4010
2017
Later among the works it cites.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, D. Eck, K. Simonyan, and M. Norouzi, “Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders,” Proc. Int. Conf. Mach. Learn. , vol. 70, pp. 1068–1077, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Mishra, B. L. Sturm, and S. Dixon, “Local interpretable model-agnostic explanations for music content analysis,” in ISMIR , 2017
2017
Later among the works it cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient Neural Audio Synthesis,” in PMLR , vol. 80, 2018, pp. 2410–2419
2018
Later among the works it cites.
M. A. Román, A. Pertusa, and J. Calvo-Zaragoza, “An End-to-end Framework for Audio-to-Score Music Transcription on Monophonic Excerpts,” in ISMIR , 2018
2018
Later among the works it cites.
Y. C. Subakan and P. Smaragdis, “Generative adversarial source separation,” in ICASSP , 2018
2018
Later among the works it cites.
C. Donahue, B. Li, and R. Prabhavalkar, “Exploring speech enhancement with generative adversarial networks for robust speech recognition,” in ICASSP , 2018
2018
Later among the works it cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan et al. , “Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,” IEEE Int. Conf. Acoust., Speech Signal Process. , pp. 4779–4783, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
A. Narayanan, A. Misra et al. , “Toward domain-invariant speech recognition via large scale training,” in SLT , 2018, pp. 441–447
2018
Later among the works it cites.
Y. Tokozume, Y. Ushiku, and T. Harada, “Learning from between-class examples for deep sound recognition,” in ICLR , 2018
2018
Later among the works it cites.
C.-C. Chiu, T. Sainath et al. , “State-of-the-art speech recognition with sequence-to-sequence models,” in ICASSP , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. Bayle, “Deep learning for music,” 2018. [Online]. Available: https://github.com/ybayle/awesome-deep-learning-music
2018
Later among the works it cites.
M. Fuentes, B. McFee, H. C. Crayencour, S. Essid, and J. P. Bello, “Analysis of common design choices in deep learning systems for downbeat tracking,” in ISMIR , 2018
2018
Later among the works it cites.
H. Schreiber and M. Müller, “A Single-Step Approach to Musical Tempo Estimation using a Convolutional Neural Network,” in ISMIR , 2018
2018
Later among the works it cites.
A. Mesaros, T. Heittola et al. , “Detection and classification of acoustic scenes and events: Outcome of the DCASE 2016 challenge,” IEEE/ACM Transactions on ASLP , vol. 26, no. 2, pp. 379–393, 2018
2018
Later among the works it cites.
E. L. Ferguson, S. B. Williams, and C. T. Jin, “Sound Source Localization in a Multipath Environment Using Convolutional Neural Networks,” in ICASSP , 2018
2018
Later among the works it cites.
S. Adavanne, A. Politis, and T. Virtanen, “Direction of arrival estimation for multiple sound sources using convolutional recurrent neural network,” in ESPC , 2018
2018
Later among the works it cites.
A. Pandey and D. Wang, “A New Framework for Supervised Speech Enhancement in the Time Domain,” in Interspeech , 2018
2018
Later among the works it cites.
S.-W. Fu, T.-W. Wang, Y. Tsao, X. Lu, and H. Kawai, “End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 9, 2018, pp. 1570–1584
2018
Later among the works it cites.
M. Kolbæk, Z.-H. Tan, and J. Jensen, “Monaural speech enhancement using deep neural networks by maximizing a short-time objective intelligibility measure,” in ICASSP , 2018
2018
Later among the works it cites.
Q. Liu, Y. Xu, P. J. Jackson, W. Wang, and P. Coleman, “Iterative Deep Neural Networks for Speaker-Independent Binaural Blind Speech Separation,” in ICASSP , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. v. d. Driessche, E. Lockhart, L. C. Cobo, F. Stimberg et al. , “Parallel WaveNet: Fast high-fidelity speech synthesis,” PMLR , vol. 80, pp. 3918–3926, 2018
2018
Later among the works it cites.
K. Chen, B. Chen, J. Lai, and K. Yu, “High-quality voice conversion using spectrogram-based wavenet vocoder,” Interspeech , 2018
2018
Later among the works it cites.
S.-Y. Chang, B. Li, G. Simko, T. N. Sainath, A. Tripathi, A. van den Oord, and O. Vinyals, “Temporal modeling using dilated convolution and gating for voice-activity-detection,” in ICASSP , 2018
2018
Later among the works it cites.
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
“AudioSet: A large-scale dataset of manually annotated audio events,” https://research.google.com/audioset/ , accessed: 2019-01-15
2019
Closest in time.
S. Huang, Q. Li, C. Anil, X. Bao, S. Oore, and R. B. Grosse, “TimbreTron: A WaveNet(CycleGAN(CQT(Audio))) Pipeline for Musical Timbre Transfer,” Proc. ICLR , 2019
2019
Closest in time.
C. Donahue, I. Simon, and S. Dieleman, “Piano Genie,” in Proc. Int. Conf. Intelligent User Interfaces , 2019, pp. 160–164
2019
Closest in time.
“ImageNet,” http://www.image-net.org , accessed: 2019-01-15
2019
Closest in time.
“Linguistic Data Consortium,” https://catalog.ldc.upenn.edu , accessed: 2019-01-15
2019
Closest in time.
“Million Song Dataset,” https://labrosa.ee.columbia.edu/millionsong/ , accessed: 2019-01-15
2019
Closest in time.
“Reference Annotations: The Beatles,” http://isophonics.net/content/reference-annotations-beatles , accessed: 2019-01-15
2019
Closest in time.