Fetching the paper…
Reading the bibliography…
The dominant speech separation models are based on complex recurrent or convolution neural network that model speech sequences indirectly conditioning on context, such as passing information through many intermediate states in recurrent neural network, leading to suboptimal separation performance.
J. Garofolo, D. Graff, D. Paul, and D. Pallett, “Csr-i (wsj0) complete ldc93s6a,”
1993
Earlier work this paper cites.
A. W. Bronkhorst, “The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,”
2000
Earlier work this paper cites.
S. Haykin and Z. Chen, “The cocktail party problem,”
2005
Earlier work this paper cites.
M. N. Schmidt and R. K. Olsson, “Single-channel speech separation using sparse non-negative matrix factorization,” 2006
2006
Earlier work this paper cites.
T. Virtanen, “Speech recognition using factorial hidden markov models for separation in the feature space,” in
2006
Earlier work this paper cites.
E. B. D. Wang, G. J. Brown, and C. Darwin, “Computational auditory scene analysis: Principles, algorithms and applications,”
2008
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2014
Earlier work this paper cites.
J. Le Roux, J. R. Hershey, and F. Weninger, “Deep nmf for speech separation,” in
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in
2016
Earlier work this paper cites.
Y. Isik, J. Le Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,”
2016
Earlier work this paper cites.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in
2017
Cited alongside, same era.
M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Cited alongside, same era.
J. Gou, Z. Yi, D. Zhang, Y. Zhan, X. Shen, and L. Du, “Sparsity and geometry preserving graph embedding for dimensionality reduction,”
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Tasnet: time-domain audio separation network for real-time, single-channel speech separation,” in
2018
G.-P. Yang, C.-I. Tuan, H.-Y. Lee, and L.-s. Lee, “Improved speech separation with time-and-frequency cross-domain joint embedding and clustering,” in
2019
Later among the works it cites.
Y. Luo and Mesgarani, “Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
2019
Later among the works it cites.
Z. Shi, H. Lin, L. Liu, R. Liu, J. Han, and A. Shi, “Deep attention gated dilated temporal convolutional networks with intra-parallel convolutional modules for end-to-end monaural speech separation,” in
2019
Later among the works it cites.
Z. Shi, H. Lin, L. Liu, R. Liu, S. Hayakawa, S. Harada, and J. Han, “End-to-end monaural speech separation with multi-scale dynamic weighted gated dilated convolutional pyramid network,” in
2019
Later among the works it cites.
N. Takahashi, S. Parthasaarathy, N. Goswami, and Y. Mitsufuji, “Recursive speech separation for unknown number of speakers,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
M. Sperber, J. Niehues, G. Neubig, S. Stüker, and A. Waibel, “Self-attentional acoustic models,”
2018
Cited alongside, same era.
Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-independent speech separation with deep attractor network,”
2018
Cited alongside, same era.
C. Xu, W. Rao, X. Xiao, E. S. Chng, and H. Li, “Single channel speech separation with constrained utterance level permutation invariant training using grid lstm,” in
2018
Cited alongside, same era.
C. Li, L. Zhu, S. Xu, P. Gao, and B. Xu, “Cbldnn-based speaker-independent speech separation via generative adversarial training,” in
2018
Cited alongside, same era.
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Alternative objective functions for deep clustering,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
E. N. N. Ocquaye, Q. Mao, H. Song, G. Xu, and Y. Xue, “Dual exclusive attentive transfer for unsupervised deep convolutional domain adaptation in speech emotion recognition,”
2019
Cited alongside, same era.
2019
Later among the works it cites.
D. Ditter and T. Gerkmann, “A multi-phase gammatone filterbank for speech separation via tasnet,”
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
Y. Liu and D. Wang, “Divide and conquer: A deep casa approach to talker-independent monaural speaker separation,”
2019
Later among the works it cites.