Fetching the paper…
Reading the bibliography…
To date, the bulk of research on single-channel speech separation has been conducted using clean, near-field, read speech, which is not representative of many modern applications.
CSR-I (WSJ0) Complete LDC93S6A
J. Garofolo, D. Graff, D. Paul, and D. Pallett, · 1993
Earlier work this paper cites.
“Probabilistic Linear Discriminant Analysis,”
Sergey Ioffe, · 2006
Earlier work this paper cites.
“Performance Measurement in Blind Audio Source Separation,”
Emmanuel Vincent, Rémi Gribonval, and Cédric Févotte, · 2006
Earlier work this paper cites.
“Probabilistic Linear Discriminant Analysis for Inferences About Identity,”
Simon J. D. Prince and James H. Elder, · 2007
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukás̆ Burget, Ondr̆ej Glembek, Nagendra Goel, Mirko Hannermann, Petr Motlíc̆ek, Yanmin Qian, Petr Schwarz, Jan Silovský, Georg Stemmer, and Karel Veselý, · 2011
Earlier work this paper cites.
Mixer-6 Speech LDC2013S03
L. Brandschain, D. Graff, and K. Walker, · 2013
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“MUSAN: A Music, Speech, and Noise Corpus,” 2015, · 2015
Cited alongside, same era.
“Deep clustering: Discriminative embeddings for segmentation and separation,”
John R. Hershey, Jonathan Le Roux, Zhuo Chen, and Shinji Watanabe, · 2016
Cited alongside, same era.
“Single-channel multi-speaker separation using deep clustering,”
Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R. Hershey, · 2016
Cited alongside, same era.
“Acoustic modelling from the signal domain using cnns.,”
Pegah Ghahremani, Vimal Manohar, Daniel Povey, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
D. Yu, M. Kolbæk, Z. Tan, and J. Jensen, · 2017
Cited alongside, same era.
“Deep attractor network for single-microphone speaker separation,”
Z. Chen, Y. Luo, and N. Mesgarani, · 2017
Later among the works it cites.
“Speaker-independent speech separation with deep attractor network,”
Zhuo Chen, Yi Luo, and Nima Mesgarani, · 2017
Later among the works it cites.
“VoxCeleb: a large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Later among the works it cites.
“Listening to each speaker one by one with recurrent selective hearing networks,”
K. Kinoshita, L. Drude, M. Delcroix, and T. Nakatani, · 2018
Closest in time.
“The Fifth CHiME Speech Separation and Recognition Challenge: Dataset, task and baselines,”
Jon Barker, Shinji Watanabe, Emmanuel Vincent, and Jan Trmal, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Multi-talker Speech Separation with Utterance-level Permutation Invariant Training of Deep Recurrent Neural Networks,”
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, · 2017
Cited alongside, same era.
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Closest in time.