Fetching the paper…
Reading the bibliography…
We present a multi-channel database of overlapping speech for training, evaluation, and detailed analysis of source separation and extraction algorithms: SMS-WSJ -- Spatialized Multi-Speaker Wall Street Journal.
“Improved MVDR beamforming using single-channel mask prediction networks,”
H. Erdogan, J. R. Hershey, S. Watanabe, M. I. Mandel, and J. Le Roux, · 1910
Earlier work this paper cites.
“Image method for efficiently simulating small-room acoustics,”
J. B. Allen and D. A. Berkley, · 1979
Earlier work this paper cites.
“The design for the Wall Street Journal-based CSR corpus,”
D. B. Paul and J. M. Baker, · 1992
Earlier work this paper cites.
“Learning the parts of objects by non-negative matrix factorization,”
D. D. Lee and H. S. Seung, · 1999
Earlier work this paper cites.
“The precedence effect,”
R. Y. Litovsky, H. S. Colburn, W. A. Yost, and S. J. Guzman, · 1999
Earlier work this paper cites.
Independent component analysis
A. Hyvärinen, J. Karhunen, and E. Oja, · 2001
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, · 2001
Earlier work this paper cites.
“Normalized observation vector clustering approach for sparse source separation,”
S. Araki, H. Sawada, R. Mukai, and S. Makino, · 2006
Earlier work this paper cites.
“Room impulse response generator,”
E. A. P. Habets, · 2006
Earlier work this paper cites.
“Performance measurement in blind audio source separation,”
E. Vincent, R. Gribonval, and C. Févotte, · 2006
Earlier work this paper cites.
“Measuring dependence of bin-wise separated signals for permutation alignment in frequency-domain BSS,”
H. Sawada, S. Araki, and S. Makino, · 2007
Earlier work this paper cites.
“Model-based expectation-maximization source separation and localization,”
M. I. Mandel, R. J. Weiss, and D. P. W. Ellis, · 2010
Cited alongside, same era.
“An EM approach to integrated multichannel speech separation and noise suppression,”
D. H. Tran Vu and R. Haeb-Umbach, · 2010
Cited alongside, same era.
“Evaluating source separation algorithms with reverberant speech,”
M. I. Mandel, S. Bressler, B. Shinn-Cunningham, and D. P. W. Ellis, · 2010
Cited alongside, same era.
“On optimal frequency-domain multichannel linear filtering for noise reduction,”
M. Souden, J. Benesty, and S. Affes, · 2010
Cited alongside, same era.
“The Kaldi speech recognition toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, · 2011
Cited alongside, same era.
“An algorithm for intelligibility prediction of time–frequency weighted noisy speech,”
“Deep clustering: discriminative embeddings for segmentation and separation,”
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, · 2016
Later among the works it cites.
“Complex angular central Gaussian mixture model for directional statistics in mask-based microphone array signal processing,”
N. Ito, S. Araki, and T. Nakatani, · 2016
Later among the works it cites.
“A generic neural acoustic beamforming architecture for robust multi-channel speech processing,”
J. Heymann, L. Drude, and R. Haeb-Umbach, · 2017
Later among the works it cites.
“Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation,”
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, · 2018
Later among the works it cites.
“The fifth ’CHiME’ speech separation and recognition challenge: Dataset, task and baselines,”
J. Barker, S. Watanabe, E. Vincent, and J. Trmal, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, · 2011
Cited alongside, same era.
“The REVERB challenge: A common evaluation framework for dereverberation and recognition of reverberant speech,”
K. Kinoshita, M. Delcroix, T. Yoshioka, T. Nakatani, A. Sehr, W. Kellermann, and R. Maas, · 2013
Cited alongside, same era.
“Permutation-free convolutive blind source separation via full-band clustering based on frequency-independent source presence priors,”
N. Ito, S. Araki, and T. Nakatani, · 2013
Cited alongside, same era.
“Deep neural network based speech separation for robust speech recognition,”
Y. Tu, J. Du, Y. Xu, L. Dai, and C.-H. Lee, · 2014
Cited alongside, same era.
“Multichannel audio database in various acoustic environments,”
E. Hadad, F. Heese, P. Vary, and S. Gannot, · 2014
Cited alongside, same era.
H. Seki, T. Hori, S. Watanabe, J. Le Roux, and J. R. Hershey, · 2018
Later among the works it cites.
“A review of blind source separation methods: two converging routes to ILRMA originating from ICA and NMF,”
H. Sawada, N. Ono, H. Kameoka, D. Kitamura, and H. Saruwatari, · 2019
Closest in time.
“SDR–half-baked or well done?,”
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, · 2019
Closest in time.
“MIMO-SPEECH: End-to-end multi-channel multi-speaker speech recognition,”
X. Chang, W. Zhang, Y. Qian, J. Le Roux, and S. Watanabe, · 2019
Closest in time.
“WHAM!: Extending speech separation to noisy environments,”
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux, · 2019
Closest in time.