Fetching the paper…
Reading the bibliography…
We present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly.
1988
Earlier work this paper cites.
M. Bosi and R. E. Goldberg, Introduction to Digital Audio Coding and Standards . Springer, 2002
2002
Earlier work this paper cites.
ITU-T, “Subjective test methodology for evaluating speech communication systems that include noise suppression algorithm,” ITU-T Recommendation P.835, Tech. Rep., 2003
2003
Earlier work this paper cites.
Y. Hu and P. C. Loizou, “Subjective comparison of speech enhancement algorithms,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2006
2006
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in International Conference on Artificial Intelligence and Statistics (AISTATS) , 2010
2010
Earlier work this paper cites.
Y. Wang and D. Wang, “Cocktail party processing via structured prediction,” in Neural Information Processing Systems (NIPS) , 2012
2012
Earlier work this paper cites.
P. C. Loizou, Speech Enhancement: Theory and Practice , 2nd ed. CRC Press, 2013
2013
Earlier work this paper cites.
X. Lu, Y. Tsao, S. Matsuda, , and C. Hori, “Speech enhancement based on deep denoising autoencoder,” in Interspeech , 2013
2013
Earlier work this paper cites.
A. Narayanan and D. Wang, “Ideal ratio mask estimation using deep neural networks for robust speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2013
2013
Earlier work this paper cites.
J. L. Roux and E. Vincent, “Consistent Wiener filtering for audio source separation,” IEEE Signal Processing Letters , vol. 20, no. 3, 2013
2013
Earlier work this paper cites.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in ICML Workshop on Deep Learning for Audio, Speech, and Language Processing , 2013
2013
Earlier work this paper cites.
P. Smaragdis, C. Fevotte, G. J. Mysore, N. Mohammadiha, and M. Hoffman, “Static and dynamic source separation using nonnegative factorizations: A unified view,” IEEE Signal Processing Magazine , vol. 31, no. 3, 2014
2014
Earlier work this paper cites.
F. Weninger, J. R. Hershey, J. L. Roux, and B. Schuller, “Discriminatively trained recurrent neural networks for single-channel speech separation,” in IEEE Global Conference on Signal and Information Processing , 2014
2014
Earlier work this paper cites.
Y. Xu, J. Du, L.-R. Dai, , and C.-H. Lee, “A regression approach to speech enhancement based on deep neural networks,” IEEE/ACM Transactions on Audio, Speech and Language Processing , vol. 23, no. 1, 2015
2015
Earlier work this paper cites.
T. Gerkmann, M. Krawczyk-Becker, and J. L. Roux, “Phase processing for single-channel speech enhancement: History and recent advances,” IEEE Signal Processing Magazine , vol. 32, no. 2, 2015
2015
Cited alongside, same era.
Y. Wang and D. Wang, “A deep neural network for time-domain signal reconstruction,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Cited alongside, same era.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. L. Roux, “Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International Conference on Machine Learning (ICML) , 2015
2015
Cited alongside, same era.
A. Mesaros, T. Heittola, and T. Virtanen, “TUT database for acoustic scene classification and sound event detection,” in European Signal Processing Conference (EUSIPCO) , 2016
2016
Later among the works it cites.
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise-robust text-to-speech,” in ISCA Speech Synthesis Workshop , 2016
2016
Later among the works it cites.
J. Chen and D. Wang, “Long short-term memory for speaker generalization in supervised speech separation,” Journal of the Acoustical Society of America , vol. 141, no. 6, 2017
2017
Later among the works it cites.
D. S. Williamson and D. Wang, “Time-frequency masking in the complex domain for speech dereverberation and denoising,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 7, 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in International Conference on Learning Representations (ICLR) , 2015
2015
Cited alongside, same era.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and F. Li, “ImageNet large scale visual recognition challenge,” International Journal on Computer Vision (IJCV) , vol. 115, no. 3, 2015
2015
Cited alongside, same era.
P. Foster, S. Sigtia, S. Krstulovic, J. Barker, and M. D. Plumbley, “CHiMe-Home: A dataset for sound source recognition in a domestic environment,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2015
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Cited alongside, same era.
2016
Cited alongside, same era.
X.-L. Zhang and D. Wang, “A deep ensemble learning method for monaural speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 5, 2016
2016
Cited alongside, same era.
F. G. Germain, G. J. Mysore, and T. Fujioka, “Equalization matching of speech recordings in real-world environments,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2016
2016
Cited alongside, same era.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in European Conference on Computer Vision (ECCV) , 2016
2016
Cited alongside, same era.
J. A. Moorer, “A note on the implementation of audio processing by short-term Fourier transform,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: Speech enhancement generative adversarial network,” in Interspeech , 2017
2017
Later among the works it cites.
K. Qian, Y. Zhang, S. Chang, X. Yang, D. Florencio, and M. Hasegawa-Johnson, “Speech enhancement using Bayesian WaveNet,” in Interspeech , 2017
2017
Later among the works it cites.
Q. Chen and V. Koltun, “Photographic image synthesis with cascaded refinement networks,” in International Conference on Computer Vision (ICCV) , 2017
2017
Later among the works it cites.
Q. Chen, J. Xu, and V. Koltun, “Fast image processing with fully-convolutional networks,” in International Conference on Computer Vision (ICCV) , 2017
2017
Later among the works it cites.
D. Rethage, J. Pons, and X. Serra, “A WaveNet for speech denoising,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2018
2018
Closest in time.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Computer Vision and Pattern Recognition (CVPR) , 2018
2018
Closest in time.
A. Mesaros, T. Heittola, E. Benetos, P. Foster, M. Lagrange, T. Virtanen, and M. D. Plumbley, “Detection and classification of acoustic scenes and events: Outcome of the DCASE 2016 challenge,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 2, 2018
2018
Closest in time.