Fetching the paper…
Reading the bibliography…
Deep neural network with dual-path bi-directional long short-term memory (BiLSTM) block has been proved to be very effective in sequence modeling, especially in speech separation, e.g.
Rix, A.W., Beerends, J.G., Hollier, M.P., Hekstra, A.P.: Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs 2, 749–752 (2001)
2001
Earlier work this paper cites.
Févotte, C., Gribonval, R., Vincent, E.: Bss_eval toolbox user guide–revision 2.0 (2005)
2005
Earlier work this paper cites.
Shao, Y., Wang, D.: Model-based sequential organization in cochannel speech. IEEE Transactions on Audio, Speech, and Language Processing 14(1), 289–298 (2006)
2006
Earlier work this paper cites.
Vincent, E., Gribonval, R., Févotte, C.: Performance measurement in blind audio source separation. IEEE transactions on audio, speech, and language processing 14(4), 1462–1469 (2006)
2006
Earlier work this paper cites.
Virtanen, T.: Speech recognition using factorial hidden markov models for separation in the feature space. In: Ninth International Conference on Spoken Language Processing (2006)
2006
Earlier work this paper cites.
Wang, D., Brown, G.J.: Computational auditory scene analysis: Principles, algorithms, and applications. Wiley-IEEE press (2006)
2006
Earlier work this paper cites.
Smaragdis, P., et al.: Convolutive speech bases and their application to supervised speech separation. IEEE Transactions on audio speech and language processing 15(1), 1 (2007)
2007
Earlier work this paper cites.
Hu, K., Wang, D.: An unsupervised approach to cochannel speech separation. IEEE Transactions on audio, speech, and language processing 21(1), 122–131 (2013)
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
Le Roux, J., Weninger, F.J., Hershey, J.R.: Sparse nmf–half-baked or well done? Mitsubishi Electric Research Labs (MERL), Cambridge, MA, USA, Tech. Rep., no. TR2015-023 (2015)
2015
Earlier work this paper cites.
Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization. arXiv preprint arXiv:1607.06450 (2016)
2016
Earlier work this paper cites.
Hershey, J.R., Chen, Z., Le Roux, J., Watanabe, S.: Deep clustering: Discriminative embeddings for segmentation and separation. In: Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on. pp. 31–35. IEEE (2016)
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
Jensen, J., Taal, C.H.: An algorithm for predicting the intelligibility of speech masked by modulated noise maskers. IEEE Transactions on Audio, Speech, and Language Processing 24(11), 2009–2022 (2016)
2016
Cited alongside, same era.
Chen, Z., Luo, Y., Mesgarani, N.: Deep attractor network for single-microphone speaker separation. In: Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on. pp. 246–250. IEEE (2017)
2017
Cited alongside, same era.
Liu, J., Wang, G., Hu, P., Duan, L.Y., Kot, A.C.: Global context-aware attention lstm networks for 3d action recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 1647–1656 (2017)
2017
Cited alongside, same era.
2019
Later among the works it cites.
Liu, Y., Wang, D.: Divide and conquer: A deep casa approach to talker-independent monaural speaker separation. IEEE Transactions on Audio, Speech, and Language Processing 27(12), 2092–2102 (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Luo, Y., Mesgarani, N.: Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation. IEEE Transactions on Audio, Speech, and Language Processing 27(8), 1256–1266 (2019)
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Venkataramani, S., Casebeer, J., Smaragdis, P.: Adaptive front-ends for end-to-end source separation. In: Proc. NIPS (2017)
2017
Cited alongside, same era.
Yu, D., Kolbæk, M., Tan, Z.H., Jensen, J.: Permutation invariant training of deep models for speaker-independent multi-talker speech separation. In: Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on. pp. 241–245. IEEE (2017)
2017
Cited alongside, same era.
Luo, Y., Chen, Z., Mesgarani, N.: Speaker-independent speech separation with deep attractor network. IEEE/ACM Transactions on Audio, Speech, and Language Processing 26(4), 787–796 (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Wu, Y., He, K.: Group normalization. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 3–19 (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
Shi, Z., Lin, H., Liu, L., Liu, R., Han, J., Shi, A.: Deep attention gated dilated temporal convolutional networks with intra-parallel convolutional modules for end-to-end monaural speech separation. Proc. Interspeech 2019 pp. 3183–3187 (2019)
2019
Later among the works it cites.
Shi, Z., Lin, H., Liu, L., Liu, R., Hayakawa, S., Han, J.: Furcax: End-to-end monaural speech separation based on deep gated (de)convolutional neural networks with adversarial example training. In: Proc. ICASSP (2019)
2019
Later among the works it cites.
Shi, Z., Lin, H., Liu, L., Liu, R., Hayakawa, S., Harada, S., Han, J.: End-to-end monaural speech separation with multi-scale dynamic weighted gated dilated convolutional pyramid network. Proc. Interspeech 2019 pp. 4614–4618 (2019)
2019
Later among the works it cites.
Wang, Z.Q., Tan, K., Wang, D.: Deep learning based phase reconstruction for speaker separation: A trigonometric perspective. In: ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 71–75 (2019)
2019
Later among the works it cites.
2020
Closest in time.
2020
Closest in time.
Zhang, L., Shi, Z., Han, J., Shi, A., Ma, D.: Furcanext: End-to-end monaural speech separation with dynamic gated dilated temporal convolutional networks. In: International Conference on Multimedia Modeling. pp. 653–665. Springer (2020)
2020
Closest in time.