Fetching the paper…
Reading the bibliography…
The performance of single channel source separation algorithms has improved greatly in recent times with the development and deployment of neural networks.
J. B. Allen and L. R. Rabiner, “A unified approach to short-time fourier analysis and synthesis,” Proceedings of the IEEE , vol. 65, no. 11, pp. 1558–1564, 1977
1977
Earlier work this paper cites.
C. Févotte, R. Gribonval, and E. Vincent, “Bss_eval toolbox user guide–revision 2.0,” 2005
2005
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on . IEEE, 2010, pp. 4214–4217
2010
Earlier work this paper cites.
C. Hummersone, T. Stokes, and T. Brookes, “On the ideal ratio mask as the goal of computational auditory scene analysis,” in Blind source separation . Springer, 2014, pp. 349–368
2014
Earlier work this paper cites.
P. Smaragdis, “Nmf? neural nets? it’s all the same.” 2015, speech and Audio in the Northeast (SANE). [Online]. Available: http://www.merl.com/events/sane2015
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention . Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis, “Joint optimization of masks and deep recurrent neural networks for monaural source separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 12, pp. 2136–2147, 2015
2015
Earlier work this paper cites.
Y. Isik, J. L. Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
A. A. Nugraha, A. Liutkus, and E. Vincent, “Multichannel audio source separation with deep neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 9, pp. 1652–1664, Sept 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , March 2017, pp. 246–250
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Tasnet: Time-domain audio separation network for real-time single-channel speech separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on . IEEE, 2018
2018
Closest in time.
D. Rethage, J. Pons, and X. Serra, “A wavenet for speech denoising,” 2018
2018
Closest in time.
2018
Closest in time.
S.-W. Fu, Y. Tsao, X. Lu, and H. Kawai, “End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks,” IEEE Transactions on Audio, Speech, and Language Processing , 2018
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Qian, Y. Zhang, S. Chang, X. Yang, D. Florêncio, and M. Hasegawa-Johnson, “Speech enhancement using bayesian wavenet,” 2017
2017
Cited alongside, same era.
Z.-Q. Wang and D. Wang, “Recurrent deep stacking networks for supervised speech separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 71–75
2017
Cited alongside, same era.
2017
Cited alongside, same era.
S. W. Fu, Y. Tsao, X. Lu, and H. Kawai, “Raw waveform-based speech enhancement by fully convolutional networks,” in 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , Dec 2017
2017
Cited alongside, same era.
P. Smaragdis and S. Venkataramani, “A neural network alternative to non-negative audio models,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on . IEEE, 2017, pp. 86–90
2017
Cited alongside, same era.
Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-independent speech separation with deep attractor network,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 4, pp. 787–796, April 2018
2018
Cited alongside, same era.
S. Venkataramani, J. Casebeer, and P. Smaragdis, “Adaptive front-ends for end-to-end source separation,” in Workshop Machine Learning for Audio Signal Processing at NIPS (ML4Audio@NIPS17)
Cited in the paper.
“Wall street journal 0 (wsj0) database,” https://catalog.ldc.upenn.edu/ldc93s6a
Cited in the paper.
2018
Closest in time.
S.-W. Fu, T.-W. Wang, Y. Tsao, X. Lu, and H. Kawai, “End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 26, no. 9, pp. 1570–1584, 2018
2018
Closest in time.
2018
Closest in time.
S. Venkataramani, R. Higa, and P. Smaragdis, “Performance based cost functions for end-to-end speech separation,” in 2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) , Dec 2018
2018
Closest in time.
Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Alternative objective functions for deep clustering,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018
2018
Closest in time.