Fetching the paper…
Reading the bibliography…
Monaural source separation is important for many real world applications.
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,” IEEE Transactions on Acoustics, Speech and Signal Processing , vol. 27, no. 2, pp. 113–120, Apr. 1979
1979
Earlier work this paper cites.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,” IEEE Transactions on Acoustics, Speech and Signal Processing , vol. 32, no. 6, pp. 1109–1121, Dec. 1984
1984
Earlier work this paper cites.
M. C. Mozer, “A focused back-propagation algorithm for temporal pattern recognition,” Complex Systems , vol. 3, no. 4, pp. 349–381, 1989
1989
Earlier work this paper cites.
P. J. Werbos, “Backpropagation through time: what it does and how to do it,” Proceedings of the IEEE , vol. 78, no. 10, pp. 1550–1560, Oct. 1990
1990
Earlier work this paper cites.
J. Garofolo, L. Lamel, W. Fisher, J. Fiscus, D. Pallett, N. Dahlgren, and V. Zue, “TIMIT: Acoustic-phonetic continuous speech corpus,” Linguistic Data Consortium, 1993
1993
Earlier work this paper cites.
R. J. Williams and D. Zipser, “Gradient-based learning algorithms for recurrent networks and their computational complexity,” in Backpropagation: Theory, Architectures, and Applications , 1995, pp. 433–486
1995
Earlier work this paper cites.
R. H. Byrd, P. Lu, J. Nocedal, and C. Zhu, “A limited memory algorithm for bound constrained optimization,” SIAM Journal on Scientific Computing , vol. 16, no. 5, pp. 1190–1208, Sep. 1995
1995
Earlier work this paper cites.
D. D. Lee and H. S. Seung, “Learning the parts of objects by non-negative matrix factorization,” Nature , vol. 401, no. 6755, pp. 788–791, Oct. 1999
1999
Earlier work this paper cites.
T. Hofmann, “Probabilistic latent semantic indexing,” in Proceedings of the 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval , 1999, pp. 50–57
1999
Earlier work this paper cites.
P. Kabal, “TSP speech database,” McGill University, Montreal, Quebec, Tech. Rep., 2002
2002
Earlier work this paper cites.
O. Yilmaz and S. Rickard, “Blind separation of speech mixtures via time-frequency masking,” IEEE Transactions on Signal Processing , vol. 52, no. 7, pp. 1830–1847, Jul. 2004
2004
Earlier work this paper cites.
P. Smaragdis, B. Raj, and M. Shashanka, “A probabilistic latent variable model for acoustic modeling,” in Proceedings of the Advances in Models for Acoustic Processing, Neural Information Processing Systems Workshop , vol. 148, 2006
2006
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Fevotte, “Performance measurement in blind audio source separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 14, no. 4, pp. 1462–1469, Jul. 2006
2006
Earlier work this paper cites.
A. Ozerov, P. Philippe, F. Bimbot, and R. Gribonval, “Adaptation of Bayesian models for single-channel source separation and its application to voice/music separation in popular songs,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 5, pp. 1564–1578, Jul. 2007
2007
Earlier work this paper cites.
D. Wang, “Time-frequency masking for speech separation and its potential for hearing aid design,” Trends in Amplification , vol. 12, pp. 332–353, 2008
2008
Cited alongside, same era.
C.-L. Hsu and J.-S. Jang, “On the improvement of singing voice separation for monaural recordings using the MIR-1K dataset,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 2, pp. 310–319, Feb. 2010
2010
Cited alongside, same era.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks,” in Proceedings of the 14th International Conference on Artificial Intelligence and Statistics (AISTATS) , vol. 15, 2011, pp. 315–323
2011
Cited alongside, same era.
C. Taal, R. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time-frequency weighted noisy speech,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19, no. 7, pp. 2125–2136, Sep. 2011
2011
Cited alongside, same era.
P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis, “Deep learning for monaural speech separation,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 1562–1566
2014
Later among the works it cites.
P.-S. Huang and M. Kim and M. Hasegawa-Johnson and P. Smaragdis, “Singing-voice separation from monaural recordings using deep recurrent neural networks,” in Proceedings of the 15th International Society for Music Information Retrieval (ISMIR) , 2014
2014
Later among the works it cites.
D. Liu, P. Smaragdis, and M. Kim, “Experiments on deep learning for speech denoising,” in Proceedings of the 15th Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2014, pp. 2685–2689
2014
Later among the works it cites.
F. Weninger, F. Eyben, and B. Schuller, “Single-channel speech separation with memory-enhanced recurrent neural networks,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 3709–3713
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P.-S. Huang, S. D. Chen, P. Smaragdis, and M. Hasegawa-Johnson, “Singing-voice separation from monaural recordings using robust principal component analysis,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2012, pp. 57–60
2012
Cited alongside, same era.
A. L. Maas, Q. V. Le, T. M. O’Neil, O. Vinyals, P. Nguyen, and A. Y. Ng, “Recurrent neural networks for noise reduction in robust ASR,” in Proceedings of the 13th Annual Conference of the International Speech Communication Association (INTERSPEECH) , 2012, pp. 22–25
2012
Cited alongside, same era.
P. Sprechmann, A. Bronstein, and G. Sapiro, “Real-time online singing voice separation from monaural recordings using robust low-rank modeling,” in Proceedings of the 13th International Society for Music Information Retrieval (ISMIR) , 2012
2012
Cited alongside, same era.
Y.-H. Yang, “On sparse and low-rank matrix decomposition for singing voice separation,” in Proceedings of the 20th ACM International Conference on Multimedia , 2012, pp. 757–760
2012
Cited alongside, same era.
J. Li, D. Yu, J.-T. Huang, and Y. Gong, “Improving wideband speech recognition using mixed-bandwidth training data in CD-DNN-HMM,” in Proceedings of the IEEE Spoken Language Technology Workshop (SLT) , 2012, pp. 131–136
2012
Cited alongside, same era.
Y.-H. Yang, “Low-rank representation of both singing voice and music accompaniment via learned dictionaries,” in Proceedings of the 14th International Society for Music Information Retrieval Conference (ISMIR) , 2013
2013
Cited alongside, same era.
A. Narayanan and D. Wang, “Ideal ratio mask estimation using deep neural networks for robust speech recognition,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2013, pp. 7092–7096
2013
Cited alongside, same era.
Y. Wang and D. Wang, “Towards scaling up classification-based speech separation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 21, no. 7, pp. 1381–1390, Jul. 2013
2013
Cited alongside, same era.
2014
Later among the works it cites.
S. Nie, H. Zhang, X. Zhang, and W. Liu, “Deep stacking networks with time series for speech separation,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 6667–6671
2014
Later among the works it cites.
Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 22, no. 12, pp. 1849–1858, Dec. 2014
2014
Later among the works it cites.
Y. Tu, J. Du, Y. Xu, L. Dai, and C.-H. Lee, “Deep neural network based speech separation for robust speech recognition,” in Proceedings of the International Symposium on Chinese Spoken Language Processing , 2014, pp. 532–536
2014
Later among the works it cites.
E. Grais, M. Sen, and H. Erdogan, “Deep neural networks for single channel source separation,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 3734–3738
2014
Later among the works it cites.
R. Pascanu, C. Gulcehre, K. Cho, and Y. Bengio, “How to construct deep recurrent neural networks,” in Proceedings of the International Conference on Learning Representations , 2014
2014
Later among the works it cites.
F. Weninger, J. R. Hershey, J. Le Roux, and B. Schuller, “Discriminatively trained recurrent neural networks for single-channel speech separation,” in Proceedings of the IEEE Global Conference on Signal and Information Processing (GlobalSIP) Symposium on Machine Learning Applications in Speech Processing , 2014, pp. 577–581
2014
Later among the works it cites.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “A regression approach to speech enhancement based on deep neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 1, pp. 7–19, Jan. 2015
2015
Closest in time.
J. Bruna, P. Sprechmann, and Y. Lecun, “Source separation with scattering non-negative matrix factorization,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Closest in time.