Fetching the paper…
Reading the bibliography…
We propose a multi-objective framework to learn both secondary targets not directly related to the intended task of speech enhancement (SE) and the primary target of the clean log-power spectra (LPS) features to be used directly for constructing the enhanced speech signals.
F. Itakura and S. Saito, “Analysis synthesis telephony based on the maximum likelihood method,” in
1968
Earlier work this paper cites.
N. Ahmed, T. Natarajan, and K. R. Rao, “Discrete cosine transform,”
1974
Earlier work this paper cites.
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,”
1979
Earlier work this paper cites.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,”
1984
Earlier work this paper cites.
——, “Speech enhancement using a minimum mean-square error log-spectral amplitude estimator,”
1985
Earlier work this paper cites.
J. S. Garofolo
1988
Earlier work this paper cites.
A. Varga and H. J. Steeneken, “Assessment for automatic speech recognition: Ii. noisex-92: A database and an experiment to study the effect of additive noise on speech recognition systems,”
1993
Earlier work this paper cites.
R. Caruna, “Multitask learning: A knowledge-based source of inductive bias,” in
1993
Earlier work this paper cites.
S. Kullback,
1997
Earlier work this paper cites.
R. Vergin, D. O’shaughnessy, and A. Farhat, “Generalized mel frequency cepstral coefficients for large-vocabulary speaker-independent continuous-speech recognition,”
1999
Earlier work this paper cites.
I. Cohen and B. Berdugo, “Speech enhancement for non-stationary noise environments,”
2001
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
D.-N. Jiang, L. Lu, H.-J. Zhang, J.-H. Tao, and L.-H. Cai, “Music type classification by spectral contrast feature,” in
2002
Earlier work this paper cites.
I. Cohen, “Noise spectrum estimation in adverse environments: Improved minima controlled recursive averaging,”
2003
Earlier work this paper cites.
G. Hu, “100 nogarofolo1988gettingnspeech environmental sounds,”
2004
Cited alongside, same era.
D. L. Wang and G. J. Brown,
2006
Cited alongside, same era.
K. S. R. Murty and B. Yegnanarayana, “Combining evidence from residual phase and mfcc features for speaker recognition,”
2006
Cited alongside, same era.
K. W. Wilson, B. Raj, and P. Smaragdis, “Regularized non-negative matrix factorization with temporal dependencies for speech denoising.” in
2008
Cited alongside, same era.
J. Du and Q. Huo, “A speech enhancement approach using piecewise linear approximation of an explicit model of environmental distortions.” in
2008
Cited alongside, same era.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,”
M. L. Seltzer and J. Droppo, “Multi-task learning in deep neural networks for improved phoneme recognition,” in
2013
Later among the works it cites.
Y. X. Wang, K. Han, and D. L. Wang, “Exploring monaural features for classification-based speech segregation,”
2013
Later among the works it cites.
G. E. Dahl, T. N. Sainath, and G. E. Hinton, “Improving deep neural networks for lvcsr using rectified linear units and dropout,” in
2013
Later among the works it cites.
2013
Later among the works it cites.
——, “An experimental study on speech enhancement based on deep neural networks,”
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath
2012
Cited alongside, same era.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,”
2012
Cited alongside, same era.
N. Mohammadiha, P. Smaragdis, and A. Leijon, “Supervised and unsupervised speech enhancement using nonnegative matrix factorization,”
2013
Cited alongside, same era.
X.-L. Zhang and J. Wu, “Denoising deep neural networks based voice activity detection,” in
2013
Cited alongside, same era.
X. Lu, Y. Tsao, S. Matsuda, and C. Hori, “Speech enhancement based on deep denoising autoencoder.” in
2013
Cited alongside, same era.
Y. X. Wang and D. L. Wang, “Towards scaling up classification-based speech separation,”
2013
Cited alongside, same era.
——, “Dynamic noise aware training for speech enhancement based on deep neural networks.” in
2014
Later among the works it cites.
B. Xia and C. Bao, “Wiener filtering based speech enhancement with weighted denoising auto-encoder and noise classification,”
2014
Later among the works it cites.
P. S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis, “Deep learning for monaural speech separation,” in
2014
Later among the works it cites.
D. Liu, P. Smaragdis, and M. Kim, “Experiments on deep learning for speech denoising,” in
2014
Later among the works it cites.
Y. X. Wang, A. Narayanan, and D. L. Wang, “On training targets for supervised speech separation,”
2014
Later among the works it cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,”
2014
Later among the works it cites.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “A regression approach to speech enhancement based on deep neural networks,”
2015
Later among the works it cites.
Z. Huang, J. Li, S. M. Siniscalchi, I.-F. Chen, J. Wu, and C.-H. Lee, “Rapid adaptation for deep neural networks through multi-task learning,” 2015, submitted to INTERSPEECH
2015
Later among the works it cites.