Fetching the paper…
Reading the bibliography…
Deep neural network (DNN) based end-to-end optimization in the complex time-frequency (T-F) domain or time domain has shown considerable potential in monaural speech separation.
D. W. Griffin and J. S. Lim, “Signal Estimation from Modified Short-Time Fourier Transform,” IEEE Trans. Audio, Speech, Signal Process. , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual Evaluation of Speech Quality (PESQ)-A New Method for Speech Quality Assessment of Telephone Networks and Codecs,” in Proc. ICASSP , vol. 2, 2001, pp. 749–752
2001
Earlier work this paper cites.
J. Le Roux, N. Ono, and S. Sagayama, “Explicit Consistency Constraints for STFT Spectrograms and Their Application to Phase Reconstruction,” Proceedings of SAPA , 2008
2008
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An Algorithm for Intelligibility Prediction of Time–Frequency Weighted Noisy Speech,” IEEE Trans. Audio, Speech, Lang. Process. , vol. 19, no. 7, pp. 2125–2136, sep 2011
2011
Earlier work this paper cites.
Y. Wang, A. Narayanan, and D. Wang, “On Training Targets for Supervised Speech Separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 22, pp. 1849–1858, 2014
2014
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Phase-Sensitive and Recognition-Boosted Speech Separation using Deep Recurrent Neural Networks,” in Proc. ICASSP , 2015, pp. 708–712
2015
Earlier work this paper cites.
K. Han, Y. Wang, D. Wang, W. S. Woods, I. Merks, and T. Zhang, “Learning Spectral Mapping for Speech Dereverberation and Denoising,” IEEE Trans. Audio, Speech, Lang. Process. , vol. 23, no. 6, pp. 982–992, 2015
2015
Earlier work this paper cites.
D. S. Williamson, Y. Wang, and D. Wang, “Complex Ratio Masking for Monaural Speech Separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , pp. 483–492, 2016
2016
Earlier work this paper cites.
A. Courville, I. Goodfellow, and Y. Bengio, Deep Learning . MIT Press, 2016. [Online]. Available: http://www.deeplearningbook.org
2016
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep Clustering: Discriminative Embeddings for Segmentation and Separation,” in Proc. ICASSP , 2016, pp. 31–35
2016
Earlier work this paper cites.
Y. Isik, J. Le Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” in Proc. Interspeech , Sep. 2016, pp. 545–549
2016
Earlier work this paper cites.
Z.-Q. Wang and D. Wang, “A Joint Training Framework for Robust Automatic Speech Recognition,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 24, no. 4, pp. 796–806, 2016
2016
Earlier work this paper cites.
S.-W. Fu, T.-Y. Hu, Y. Tsao, and X. Lu, “Complex Spectrogram Enhancement By Convolutional Neural Network with Multi-Metrics Learning,” in Proc. MLSP , 2017, pp. 1–6
2017
Earlier work this paper cites.
S. Pascual, A. Bonafonte, and J. Serr, “SEGAN : Speech Enhancement Generative Adversarial Network,” in Proc. Interspeech , 2017
2017
Earlier work this paper cites.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multi-Talker Speech Separation with Utterance-Level Permutation Invariant Training of Deep Recurrent Neural Networks,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 25, no. 10, pp. 1901–1913, 2017
2017
Earlier work this paper cites.
D. Wang and J. Chen, “Supervised Speech Separation Based on Deep Learning: An Overview,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 26, pp. 1702–1726, 2018
2018
Earlier work this paper cites.
Z.-Q. Wang, J. Le Roux, D. Wang, and J. R. Hershey, “End-to-End Speech Separation with Unfolded Iterative Phase Reconstruction,” in Proc. Interspeech , 2018, pp. 2708–2712
2018
Earlier work this paper cites.
D. Rethage, J. Pons, and X. Serra, “A WaveNet for Speech Denoising,” in Proc. ICASSP , 2018, pp. 5069–5073
2018
Cited alongside, same era.
D. Stoller, S. Ewert, and S. Dixon, “Wave-U-Net: A Multi-Scale Neural Network for End-to-End Audio Source Separation,” in Proc. ISMIR , 2018, pp. 334–340
2018
Cited alongside, same era.
Z.-Q. Wang, K. Tan, and D. Wang, “Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric Perspective,” in Proc. ICASSP , 2019, pp. 71–75
2019
Cited alongside, same era.
Y. Zhao, Z.-Q. Wang, and D. Wang, “Two-Stage Deep Learning for Noisy-Reverberant Speech Enhancement,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, no. 1, pp. 53–62, 2019
2019
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, no. 8, pp. 1256–1266, 2019
——, “Multi-Microphone Complex Spectral Mapping for Speech Dereverberation,” in Proc. ICASSP , 2020, pp. 486–490
2020
Later among the works it cites.
A. Défossez, G. Synnaeve, and Y. Adi, “Real Time Speech Enhancement in the Waveform Domain,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
U. Isik, R. Giri, N. Phansalkar, J. M. Valin, K. Helwani, and A. Krishnaswamy, “PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased Loss,” in Proc. Interspeech , 2020, pp. 2487–2491
2020
Later among the works it cites.
M. Kolbæk, Z.-H. Tan, S. H. Jensen, and J. Jensen, “On Loss Functions for Supervised Monaural Time-Domain Speech Enhancement,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 28, pp. 825–838, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
A. Pandey and D. Wang, “A New Framework for CNN-Based Speech Enhancement in the Time Domain,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, pp. 1179–1188, 2019
2019
Cited alongside, same era.
Y. Liu and D. Wang, “Divide and Conquer: A Deep CASA Approach to Talker-Independent Monaural Speaker Separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 27, no. 12, pp. 2092–2102, 2019
2019
Cited alongside, same era.
S. Wisdom, J. R. Hershey, K. Wilson, J. Thorpe, M. Chinen, B. Patton, and R. A. Saurous, “Differentiable Consistency Constraints for Improved Deep Speech Enhancement,” in Proc. ICASSP , 2019, pp. 900–904
2019
Cited alongside, same era.
F. G. Germain, Q. Chen, and V. Koltun, “Speech Denoising with Deep Feature Losses,” in Proc. Interspeech , 2019, pp. 2723–2727
2019
Cited alongside, same era.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR – Half-Baked or Well Done?” in Proc. ICASSP , 2019, pp. 626–630
2019
Cited alongside, same era.
P. Wang and D. Wang, “Enhanced Spectral Features for Distortion-Independent Acoustic Modeling,” in Proc. Interspeech , 2019
2019
Cited alongside, same era.
Y. Luo, E. Ceolini, C. Han, S.-C. Liu, and N. Mesgarani, “FaSNet: Low-latency Adaptive Beamforming for Multi-Microphone Audio Processing,” in Proc. WASPAA , 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
P. Manuel, “https://github.com/mpariente/pystoi,” 2020
2020
Later among the works it cites.
Ludlows, “https://github.com/ludlows/python-pesq,” 2020
2020
Later among the works it cites.
N. Turpault, S. Wisdom, H. Erdogan, J. Hershey, R. Serizel, E. Fonseca, P. Seetharaman, and J. Salamon, “Improving Sound Event Detection in Domestic Environments using Sound Separation,” in Proc. DCASE , 2020
2020
Later among the works it cites.
Z.-Q. Wang, P. Wang, and D. Wang, “Multi-Microphone Complex Spectral Mapping for Utterance-Wise and Continuous Speaker Separation,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , 2021
2021
Closest in time.
A. Li, W. Liu, C. Zheng, C. Fan, and X. Li, “Two Heads Are Better Than One: A Two-Stage Complex Spectral Mapping Approach for Monaural Speech Enhancement,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , 2021
2021
Closest in time.
S. Braun, H. Gamper, C. K. A. Reddy, and I. Tashev, “Towards Efficient Models for Real-Time Deep Noise Suppression,” in Proc. ICASSP , 2021, pp. 656–660
2021
Closest in time.
A. Aroudi and S. Braun, “DBnet: Doa-Driven Beamforming Network for End-to-End Reverberant Sound Source Separation,” in Proc. ICASSP , 2021, pp. 211–215
2021
Closest in time.
R. Sawata, S. Uhlich, S. Takahashi, and Y. Mitsufuji, “All For One And One For All: Improving Music Separation By Bridging Networks,” in Proc. ICASSP , 2021, pp. 51–55
2021
Closest in time.
A. Li, C. Zheng, R. Peng, and X. Li, “On The Importance of Power Compression and Phase Estimation in Monaural Speech Dereverberation,” JASA Express Letters , vol. 1, no. 1, p. 014802, 2021
2021
Closest in time.
P. Manocha, Z. Jin, R. Zhang, and A. Finkelstein, “CDPAM: Contrastive Learning for Perceptual Audio Similarity,” in Proc. ICASSP , 2021, pp. 196–200
2021
Closest in time.
S. Wisdom, H. Erdogan, D. P. W. Ellis, R. Serizel, N. Turpault, E. Fonseca, J. Salamon, P. Seetharaman, and J. R. Hershey, “What’s All The Fuss About Free Universal Sound Separation Data?” in Proc. ICASSP , 2021, pp. 186–190
2021
Closest in time.