Fetching the paper…
Reading the bibliography…
Many deep learning techniques are available to perform source separation and reduce background noise.
“Image method for efficiently simulating small-room acoustics,”
J. B. Allen and D. A. Berkley, · 1979
Earlier work this paper cites.
“Beamforming: a versatile approach to spatial filtering,”
B. D. van Veen and K. M. Buckley, · 1988
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
Microphone Arrays Signal Processing Techniques and Applications
Michael Brandstein and Darren Ward, · 2001
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, · 2001
Earlier work this paper cites.
“Adam: A method for stochastic optimization,” arXiv preprint, 2014
Diederik Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Multi-talker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
M. Kolbæk, D. Yu, Z. Tan, and J. Jensen, · 2017
Earlier work this paper cites.
“Deep attractor network for single-microphone speaker separation,”
Z. Chen, Y. Luo, and N. Mesgarani, · 2017
Earlier work this paper cites.
“Listening to each speaker one by one with recurrent selective hearing networks,”
K. Kinoshita, L. Drude, M. Delcroix, and T. Nakatani, · 2018
Earlier work this paper cites.
“Tasnet: Time-domain audio separation network for real-time, single-channel speech separation,”
Y. Luo and N. Mesgarani, · 2018
Cited alongside, same era.
“Phase-sensitive joint learning algorithms for deep learning-based speech enhancement,”
J. Lee, J. Skoglund, T. Shabestary, and H. Kang, · 2018
Cited alongside, same era.
“Exploring tradeoffs in models for low-latency speech enhancement,”
K. Wilson, M. Chinen, J. Thorpe, B. Patton, J. Hershey, R. A. Saurous, J. Skoglund, and R. F. Lyon, · 2018
Cited alongside, same era.
“The 2018 signal separation evaluation campaign,”
Fabian-Robert Stöter, Antoine Liutkus, and Nobutaka Ito, · 2018
Cited alongside, same era.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Y. Luo and N. Mesgarani, · 2019
Cited alongside, same era.
“Multi-speaker doa estimation using deep convolutional networks trained with noise signals,”
S. Chakrabarty and E. A. P. Habets, · 2019
Later among the works it cites.
“Real-time speech enhancement using an efficient convolutional recurrent network for dual-microphone mobile phones in close-talk scenarios,”
K. Tan, X. Zhang, and D. Wang, · 2019
Later among the works it cites.
“SDR – half-baked or well done?,”
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, · 2019
Later among the works it cites.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92), [sound],” The Centre for Speech Technology Research, University of Edinburgh, 2019, https://github.com/microsoft/DNS-Challenge
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald, · 2019
Later among the works it cites.
“Whamr!: Noisy and reverberant single-channel speech separation,”
M. Maciejewski, G. Wichern, E. McQuinn, and J. L. Roux, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ziqiang Shi, Huibin Lin, Liu Liu, Rujie Liu, Shoji Hayakawa, Shouji Harada, and Jiqing Han, · 2019
Cited alongside, same era.
“Integration of neural networks and probabilistic spatial models for acoustic blind source separation,”
L. Drude and R. Haeb-Umbach, · 2019
Cited alongside, same era.
“Low-latency speaker-independent continuous speech separation,”
T. Yoshioka, Z. Chen, C. Liu, X. Xiao, H. Erdogan, and D. Dimitriadis, · 2019
Cited alongside, same era.
“Fasnet: Low-latency adaptive beamforming for multi-microphone audio processing,”
Y. Luo, C. Han, N. Mesgarani, E. Ceolini, and S. Liu, · 2019
Cited alongside, same era.
“Device and produced speech (DAPS) dataset,” https://ccrma.stanford.edu/~gautham/Site/daps.html
Cited in the paper.
“Beam-TasNet: Time-domain audio separation network meets frequency-domain beamformer,”
T. Ochiai, M. Delcroix, R. Ikeshita, K. Kinoshita, T. Nakatani, and S. Araki, · 2020
Closest in time.
“A consolidated view of loss functions for supervised deep learning-based speech enhancement,” arXiv preprint, 2020
Sebastian Braun and Ivan Tashev, · 2020
Closest in time.
“IEEE ICASSP 2021 Deep Noise Suppression (DNS) Challenge,” https://github.com/microsoft/DNS-Challenge
2021
Closest in time.