Fetching the paper…
Reading the bibliography…
This paper proposes an end-to-end approach for single-channel speaker-independent multi-speaker speech separation, where time-frequency (T-F) masking, the short-time Fourier transform (STFT), and its inverse are represented as layers within a deep network.
D. W. Griffin and J. S. Lim, “Signal Estimation from Modified Short-Time Fourier Transform,”
1984
Earlier work this paper cites.
F. Bach and M. Jordan, “Learning Spectral Clustering, with Application to Speech Separation,”
2006
Earlier work this paper cites.
D. Wang and G. J. Brown,
2006
Earlier work this paper cites.
C. M. Bishop,
2006
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,”
2006
Earlier work this paper cites.
J. Le Roux, N. Ono, and S. Sagayama, “Explicit consistency constraints for STFT spectrograms and their application to phase reconstruction,” in
2008
Earlier work this paper cites.
J. R. Hershey, S. Rennie, P. A. Olsen, and T. T. Kristjansson, “Super-Human Multi-Talker Speech Recognition: A Graphical Modeling Approach,”
2010
Earlier work this paper cites.
D. Gunawan and D. Sen, “Iterative Phase Estimation for the Synthesis of Separated Sources from Single-Channel Mixtures,” in
2010
Earlier work this paper cites.
N. Sturmel and L. Daudet, “Signal Reconstruction from STFT Magnitude: A State of the Art,” in
2011
Earlier work this paper cites.
N. Sturmel and L. Daudet, “Informed Source Separation using Iterative Reconstruction,”
2013
Earlier work this paper cites.
J. Le Roux and E. Vincent, “Consistent Wiener Filtering for Audio Source Separation,”
2013
Earlier work this paper cites.
F. Weninger, J. R. Hershey, J. Le Roux, and B. Schuller, “Discriminatively Trained Recurrent Neural Networks for Single-channel Speech Separation,” in
2014
Earlier work this paper cites.
Y. Wang, A. Narayanan, and D. Wang, “On Training Targets for Supervised Speech Separation,”
2014
Earlier work this paper cites.
T. Gerkmann, M. Krawczyk-Becker, and J. Le Roux, “Phase Processing for Single-Channel Speech Enhancement: History and Recent Advances,”
2015
Earlier work this paper cites.
K. Han, Y. Wang, D. Wang, W. S. Woods, and I. Merks, “Learning Spectral Mapping for Speech Dereverberation and Denoising,”
2015
Cited alongside, same era.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Phase-Sensitive and Recognition-Boosted Speech Separation using Deep Recurrent Neural Networks,” in
2015
Cited alongside, same era.
Y. Wang and D. Wang, “A Deep Neural Network for Time-Domain Signal Reconstruction,” in
2015
Cited alongside, same era.
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. Le Roux, J. R. Hershey, and B. Schuller, “Speech Enhancement with LSTM Recurrent Neural Networks and its Application to Noise-Robust ASR,” in
2015
Cited alongside, same era.
J. R. Hershey, Z. Chen, and J. Le Roux, “Deep Clustering: Discriminative Embeddings for Segmentation and Separation,” in
2016
Cited alongside, same era.
S. Venkataramani and P. Smaragdis, “End-to-end source separation with adaptive front-ends,” in
2017
Later among the works it cites.
D. S. Williamson and D. Wang, “Time-Frequency Masking in the Complex Domain for Speech Dereverberation and Denoising,”
2017
Later among the works it cites.
D. S. Williamson and D. Wang, “Speech Dereverberation and Denoising using Complex Ratio Masks,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
——, “Raw Waveform-Based Speech Enhancement by Fully Convolutional Networks,” in
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Isik, J. Le Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-Channel Multi-Speaker Separation using Deep Clustering,” in
2016
Cited alongside, same era.
K. Li, B. Wu, and C.-H. Lee, “An Iterative Phase Recovery Framework with Phase Mask for Spectral Mapping with an Application to Speech Enhancement,” in
2016
Cited alongside, same era.
D. S. Williamson, Y. Wang, and D. Wang, “Complex Ratio Masking for Monaural Speech Separation,”
2016
Cited alongside, same era.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep Attractor Network for Single-Microphone Speaker Separation,” in
2017
Cited alongside, same era.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation Invariant Training of Deep Models for Speaker-Independent Multi-talker Speech Separation,” in
2017
Cited alongside, same era.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multi-Talker Speech Separation with Utterance-Level Permutation Invariant Training of Deep Recurrent Neural Networks,”
2017
Cited alongside, same era.
Y. Zhao, Z.-Q. Wang, and D. Wang, “A Two-stage Algorithm for Noisy and Reverberant Speech Enhancement,” in
2017
Cited alongside, same era.
Later among the works it cites.
K. Qian, Y. Zhang, S. Chang, X. Yang, M. Hasegawa-Johnson, D. Florencio, and M. Hasegawa-Johnson, “Speech Enhancement using Bayesian WaveNet,” in
2017
Later among the works it cites.
S. Pascual, A. Bonafonte, and J. Serra, “SEGAN: Speech Enhancement Generative Adversarial Network,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
D. Wang and J. Chen, “Supervised Speech Separation Based on Deep Learning: An Overview,” in
2017
Later among the works it cites.
Z.-Q. Wang and D. Wang, “Recurrent Deep Stacking Networks for Supervised Speech Separation,” in
2017
Later among the works it cites.
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Alternative Objective Functions for Deep Clustering,” in
2018
Closest in time.
Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-Independent Speech Separation with Deep Attractor Network,”
2018
Closest in time.
J. Le Roux, J. R. Hershey, S. T. Wisdom, and H. Erdogan, “SDR – half-baked or well done?” Mitsubishi Electric Research Laboratories (MERL), Cambridge, MA, USA, Tech. Rep., 2018
2018
Closest in time.