Fetching the paper…
Reading the bibliography…
In recent years time domain speech separation has excelled over frequency domain separation in single channel scenarios and noise-free environments.
“CSR-I (WSJ0) complete LDC93S6A,”
J. Garofolo, D. Graff, D. Paul, and D. Pallett, · 1993
Earlier work this paper cites.
“The Kaldi speech recognition toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, . Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, · 2011
Earlier work this paper cites.
“mir_eval: a transparent implementation of common MIR metrics,”
C. Raffel, B. Mcfee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, D. P. W. Ellis, C. C. Raffel, B. Mcfee, and E. J. Humphrey, · 2014
Earlier work this paper cites.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, · 2016
Earlier work this paper cites.
“Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise,”
T. Higuchi, N. Ito, T. Yoshioka, and T. Nakatani, · 2016
Earlier work this paper cites.
“Deep attractor network for single-microphone speaker separation,”
Z. Chen, Y. Luo, and N. Mesgarani, · 2017
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
D. Yu, M. Kolbaek, Z. Tan, and J. Jensen, · 2017
Earlier work this paper cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
M. Kolbaek, D. Yu, Z. Tan, and J. Jensen, · 2017
Earlier work this paper cites.
“Speaker-aware neural network based beamformer for speaker extraction in speech mixtures,”
Katerina Zmolíková and Marc Delcroix and Keisuke Kinoshita and Takuya Higuchi and Atsunori Ogawa and Tomohiro Nakatani, · 2017
Earlier work this paper cites.
“Glottal model based speech beamforming for ad-hoc microphone arrays,”
Y. Zhang, D. A. F. Florêncio, and M. Hasegawa-Johnson, · 2017
Cited alongside, same era.
“Alternative objective functions for deep clustering,”
Z. Wang, J. Le Roux, and J. R. Hershey, · 2018
Cited alongside, same era.
“Listening to each speaker one by one with recurrent selective hearing networks,”
K. Kinoshita, L. Drude, M. Delcroix, and T. Nakatani, · 2018
Cited alongside, same era.
“End-to-end speech separation with unfolded iterative phase reconstruction,”
Z. Wang, J. Le Roux, D. Wang, and J. R. Hershey, · 2018
Cited alongside, same era.
“TasNet: Surpassing ideal time-frequency masking for speech separation,”
Y. Luo and N. Mesgarani, · 2018
Cited alongside, same era.
Z. Shi, H. Lin, L. Liu, R. Liu, and J. Han, · 2019
Closest in time.
“Recursive speech separation for unknown number of speakers,”
N. Takahashi, S. Parthasaarathy, N. Goswami, and Y. Mitsufuji, · 2019
Closest in time.
“A comprehensive study of speech separation: Spectrogram vs waveform separation,”
F. Bahmaninezhad, J. Wu, R. Gu, S.-X. Zhang, Y. Xu, M. Yu, and D. Yu, · 2019
Closest in time.
“Integration of neural networks and probabilistic spatial models for acoustic blind source separation,”
L. Drude and R. Haeb-Umbach, · 2019
Closest in time.
“FaSNet: Low-latency Adaptive Beamforming for Multi-microphone Audio Processing,”
Y. Luo, E. Ceolini, C. Han, S.-C. Liu, and N. Mesgarani, · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Yin, Z. Wang, R. Xia, J. Li, and Y. Yan, · 2018
Cited alongside, same era.
“Phasebook and friends: Leveraging discrete representations for source separation,”
J. Le Roux, G. Wichern, S. Watanabe, A. Sarroff, and J. Hershey, · 2019
Cited alongside, same era.
“Deep learning based phase reconstruction for speaker separation: A trigonometric perspective,”
Z. Wang, K. Tan, and D. Wang, · 2019
Cited alongside, same era.
“Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,”
Y. Luo and N. Mesgarani, · 2019
Cited alongside, same era.
Closest in time.
“WHAM!: Extending speech separation to noisy environments,”
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux, · 2019
Closest in time.
“End-to-end multi-channel speech separation,”
R. Gu, J. Wu, S. Zhang, L. Chen, Y. Xu, M. Yu, D. Su, Y. Zou, and D. Yu, · 2019
Closest in time.
L. Drude, J. Heitkaemper, C. Boeddeker, and R. Haeb-Umbach, · 2019
Closest in time.