Fetching the paper…
Reading the bibliography…
Single-channel, speaker-independent speech separation methods have recently seen great progress.
R. R. Sokal, “A statistical method for evaluating systematic relationship,” University of Kansas science bulletin , vol. 28, pp. 1409–1438, 1958
1958
Earlier work this paper cites.
G. L. Romani, S. J. Williamson, and L. Kaufman, “Tonotopic organization of the human auditory cortex,” Science , vol. 216, no. 4552, pp. 1339–1340, 1982
1982
Earlier work this paper cites.
S. Imai, “Cepstral analysis synthesis on the mel frequency scale,” in Acoustics, Speech, and Signal Processing, IEEE International Conference on ICASSP’83. , vol. 8. IEEE, 1983, pp. 93–96
1983
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 32, no. 2, pp. 236–243, 1984
1984
Earlier work this paper cites.
C. Pantev, M. Hoke, B. Lutkenhoner, and K. Lehnertz, “Tonotopic organization of the auditory cortex: pitch versus frequency representation,” Science , vol. 246, no. 4929, pp. 486–488, 1989
1989
Earlier work this paper cites.
T.-W. Lee, M. S. Lewicki, M. Girolami, and T. J. Sejnowski, “Blind source separation of more sources than mixtures using overcomplete representations,” IEEE signal processing letters , vol. 6, no. 4, pp. 87–90, 1999
1999
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in Acoustics, Speech, and Signal Processing, 2001. Proceedings.(ICASSP’01). 2001 IEEE International Conference on , vol. 2. IEEE, 2001, pp. 749–752
2001
Earlier work this paper cites.
M. Zibulevsky and B. A. Pearlmutter, “Blind source separation by sparse decomposition in a signal dictionary,” Neural computation , vol. 13, no. 4, pp. 863–882, 2001
2001
Earlier work this paper cites.
C. J. Darwin, D. S. Brungart, and B. D. Simpson, “Effects of fundamental frequency and vocal-tract length changes on attention to one of two simultaneous talkers,” The Journal of the Acoustical Society of America , vol. 114, no. 5, pp. 2913–2922, 2003
2003
Earlier work this paper cites.
S. Choi, A. Cichocki, H.-M. Park, and S.-Y. Lee, “Blind source separation and independent component analysis: A review,” Neural Information Processing-Letters and Reviews , vol. 6, no. 1, pp. 1–57, 2005
2005
Earlier work this paper cites.
D. Wang, “On ideal binary mask as the computational goal of auditory scene analysis,” in Speech separation by humans and machines . Springer, 2005, pp. 181–197
2005
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE transactions on audio, speech, and language processing , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
ITU-T Rec. P.10, “Vocabulary for performance and quality of service,” 2006
2006
Earlier work this paper cites.
J. Le Roux, N. Ono, and S. Sagayama, “Explicit consistency constraints for stft spectrograms and their application to phase reconstruction.” in SAPA@ INTERSPEECH , 2008, pp. 23–28
2008
Earlier work this paper cites.
Y. Li and D. Wang, “On the optimality of ideal binary time–frequency masks,” Speech Communication , vol. 51, no. 3, pp. 230–239, 2009
2009
Earlier work this paper cites.
F.-Y. Wang, C.-Y. Chi, T.-H. Chan, and Y. Wang, “Nonnegative least-correlated component analysis for separation of dependent sources by volume maximization,” IEEE transactions on pattern analysis and machine intelligence , vol. 32, no. 5, pp. 875–888, 2010
2010
Earlier work this paper cites.
C. H. Ding, T. Li, and M. I. Jordan, “Convex and semi-nonnegative matrix factorizations,” IEEE transactions on pattern analysis and machine intelligence , vol. 32, no. 1, pp. 45–55, 2010
2010
Earlier work this paper cites.
J. R. Hershey, S. J. Rennie, P. A. Olsen, and T. T. Kristjansson, “Super-human multi-talker speech recognition: A graphical modeling approach,” Computer Speech & Language , vol. 24, no. 1, pp. 45–66, 2010
2010
Earlier work this paper cites.
X. Lu, Y. Tsao, S. Matsuda, and C. Hori, “Speech enhancement based on deep denoising autoencoder.” in Interspeech , 2013, pp. 436–440
2013
Earlier work this paper cites.
K. Yoshii, R. Tomioka, D. Mochihashi, and M. Goto, “Beyond nmf: Time-domain audio source separation without phase reconstruction.” in ISMIR , 2013, pp. 369–374
2013
Earlier work this paper cites.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “An experimental study on speech enhancement based on deep neural networks,” IEEE Signal processing letters , vol. 21, no. 1, pp. 65–68, 2014
2014
Earlier work this paper cites.
Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 22, no. 12, pp. 1849–1858, 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Y. Lei, N. Scheffer, L. Ferrer, and M. McLaren, “A novel scheme for speaker recognition using a phonetically-aware deep neural network,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on . IEEE, 2014, pp. 1695–1699
2014
Cited alongside, same era.
——, “A regression approach to speech enhancement based on deep neural networks,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 23, no. 1, pp. 7–19, 2015
2015
Cited alongside, same era.
2017
Later among the works it cites.
S. Pascual, A. Bonafonte, and J. Serrà, “Segan: Speech enhancement generative adversarial network,” Proc. Interspeech 2017 , pp. 3642–3646, 2017
2017
Later among the works it cites.
C. Lea, M. D. Flynn, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks for action segmentation and detection,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 156–165
2017
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention . Springer, 2015, pp. 234–241
2015
Cited alongside, same era.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on . IEEE, 2015, pp. 708–712
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 1026–1034
2015
Cited alongside, same era.
C. Weng, D. Yu, M. L. Seltzer, and J. Droppo, “Deep neural networks for single-channel multi-talker speech recognition,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 23, no. 10, pp. 1670–1679, 2015
2015
Cited alongside, same era.
M. McLaren, Y. Lei, and L. Ferrer, “Advances in deep neural network approaches to speaker recognition,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on . IEEE, 2015, pp. 4814–4818
2015
Cited alongside, same era.
Y. Isik, J. Le Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” Interspeech 2016 , pp. 545–549, 2016
2016
Cited alongside, same era.
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on . IEEE, 2016, pp. 31–35
2016
Cited alongside, same era.
C. Lea, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks: A unified approach to action segmentation,” in European Conference on Computer Vision . Springer, 2016, pp. 47–54
2016
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Gannot, E. Vincent, S. Markovich-Golan, A. Ozerov, S. Gannot, E. Vincent, S. Markovich-Golan, and A. Ozerov, “A consolidated perspective on multimicrophone speech enhancement and source separation,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 25, no. 4, pp. 692–730, 2017
2017
Later among the works it cites.
Z. Chen, J. Li, X. Xiao, T. Yoshioka, H. Wang, Z. Wang, and Y. Gong, “Cracking the cocktail party problem by multi-beam deep attractor network,” in Automatic Speech Recognition and Understanding Workshop (ASRU), 2017 IEEE . IEEE, 2017, pp. 437–444
2017
Later among the works it cites.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2018
2018
Closest in time.
Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-independent speech separation with deep attractor network,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 4, pp. 787–796, 2018. [Online]. Available: http://dx.doi.org/10.1109/TASLP.2018.2795749
2018
Closest in time.
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Alternative objective functions for deep clustering,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018
2018
Closest in time.
2018
Closest in time.
C. Li, L. Zhu, S. Xu, P. Gao, and B. Xu, “ CBLDNN
2018
Closest in time.
2018
Closest in time.
Y. Luo and N. Mesgarani, “Tasnet: time-domain audio separation network for real-time, single-channel speech separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2018 IEEE International Conference on . IEEE, 2018
2018
Closest in time.
S.-W. Fu, T.-W. Wang, Y. Tsao, X. Lu, and H. Kawai, “End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 26, no. 9, pp. 1570–1584, 2018
2018
Closest in time.
Y. Luo and N. Mesgarani, “Real-time single-channel dereverberation and separation with time-domain audio separation network,” Proc. Interspeech 2018 , pp. 342–346, 2018
2018
Closest in time.
2018
Closest in time.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 4510–4520
2018
Closest in time.
C. Xu, X. Xiao, and H. Li, “Single channel speech separation with constrained utterance level permutation invariant training using grid lstm,” in Acoustics, Speech and Signal Processing (ICASSP), 2018 IEEE International Conference on . IEEE, 2018
2018
Closest in time.
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2018 IEEE International Conference on . IEEE, 2018
2018
Closest in time.