Fetching the paper…
Reading the bibliography…
Noisy situations cause huge problems for suffers of hearing loss as hearing aids often make the signal more audible but do not always restore the intelligibility.
U. Kjems, M. S. Pedersen, J. B. Boldt, T. Lunner, D. Wang, Speech intelligibility of ideal binary masked mixtures, in: 2010 18th European Signal Processing Conference, IEEE, 2010, pp. 1909–1913
1913
Earlier work this paper cites.
S. Boll, A spectral subtraction algorithm for suppression of acoustic noise in speech, in: Acoustics, Speech, and Signal Processing, IEEE International Conference on ICASSP’79., Vol. 4, IEEE, 1979, pp. 200–203
1979
Earlier work this paper cites.
D. Griffin, J. Lim, Signal estimation from modified short-time fourier transform, IEEE Transactions on Acoustics, Speech, and Signal Processing 32 (2) (1984) 236–243
1984
Earlier work this paper cites.
Y. Ephraim, D. Malah, Speech enhancement using a minimum mean-square error log-spectral amplitude estimator, IEEE transactions on acoustics, speech, and signal processing 33 (2) (1985) 443–445
1985
Earlier work this paper cites.
Q. Summerfield, Lipreading and audio-visual speech perception, Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences 335 (1273) (1992) 71–78
1992
Earlier work this paper cites.
A. Varga, H. J. Steeneken, Assessment for automatic speech recognition: Ii. noisex-92: A database and an experiment to study the effect of additive noise on speech recognition systems, Speech communication 12 (3) (1993) 247–251
1993
Earlier work this paper cites.
K. W. Grant, P.-F. Seitz, The use of visible speech cues for improving auditory detection of spoken sentences, The Journal of the Acoustical Society of America 108 (3) (2000) 1197–1208
2000
Earlier work this paper cites.
K. W. Grant, S. Greenberg, Speech intelligibility derived from asynchronous processing of auditory-visual information, in: AVSP 2001-International Conference on Auditory-Visual Speech Processing, 2001
2001
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, A. P. Hekstra, Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs, in: 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No. 01CH37221), Vol. 2, IEEE, 2001, pp. 749–752
2001
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, X. Shao, An audio-visual corpus for speech perception and automatic speech recognition, The Journal of the Acoustical Society of America 120 (5) (2006) 2421–2424
2006
Earlier work this paper cites.
Y. Hu, P. C. Loizou, Evaluation of objective quality measures for speech enhancement, IEEE Transactions on audio, speech, and language processing 16 (1) (2007) 229–238
2007
Earlier work this paper cites.
D. Wang, U. Kjems, M. S. Pedersen, J. B. Boldt, T. Lunner, Speech intelligibility in background noise with ideal binary time-frequency masking, The Journal of the Acoustical Society of America 125 (4) (2009) 2336–2347
2009
Earlier work this paper cites.
D. E. King, Dlib-ml: A machine learning toolkit, Journal of Machine Learning Research 10 (2009) 1755–1758
2009
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, J. Jensen, An algorithm for intelligibility prediction of time–frequency weighted noisy speech, IEEE Transactions on Audio, Speech, and Language Processing 19 (7) (2011) 2125–2136
2011
Cited alongside, same era.
E. Z. Golumbic, G. B. Cogan, C. E. Schroeder, D. Poeppel, Visual input enhances selective speech envelope tracking in auditory cortex at a “cocktail party”, Journal of Neuroscience 33 (4) (2013) 1417–1426
2013
Cited alongside, same era.
M. Ahmadi, V. L. Gross, D. G. Sinex, Perceptual learning for speech in noise after application of binary time-frequency masks, The Journal of the Acoustical Society of America 133 (3) (2013) 1687–1692
2013
Cited alongside, same era.
A. Narayanan, D. Wang, Investigation of speech separation as a front-end for noise robust speech recognition, IEEE/ACM Transactions on Audio, Speech, and Language Processing 22 (4) (2014) 826–835
2014
Cited alongside, same era.
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, M. Rubinstein, Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation, ACM Transactions on Graphics (TOG) 37 (4) (2018) 112
2018
Later among the works it cites.
M. Gogate, A. Adeel, R. Marxer, J. Barker, A. Hussain, Dnn driven speaker independent audio-visual mask estimation for speech separation, Proc. Interspeech 2018 (2018) 2723–2727
2018
Later among the works it cites.
A. Gabbay, A. Shamir, S. Peleg, Visual speech enhancement, in: Interspeech, ISCA, 2018, pp. 1170–1174
2018
Later among the works it cites.
J.-C. Hou, S.-S. Wang, Y.-H. Lai, Y. Tsao, H.-W. Chang, H.-M. Wang, Audio-visual speech enhancement using multimodal deep convolutional neural networks, IEEE Transactions on Emerging Topics in Computational Intelligence 2 (2) (2018) 117–128
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wang, A. Narayanan, D. Wang, On training targets for supervised speech separation, IEEE/ACM transactions on audio, speech, and language processing 22 (12) (2014) 1849–1858
2014
Cited alongside, same era.
H. Kayser, C. Spille, D. Marquardt, B. T. Meyer, Improving automatic speech recognition in spatially-aware hearing aids, in: Sixteenth Annual Conference of the International Speech Communication Association, 2015
2015
Cited alongside, same era.
J. Barker, R. Marxer, E. Vincent, S. Watanabe, The third ‘chime’speech separation and recognition challenge: Dataset, task and baselines, in: Automatic Speech Recognition and Understanding (ASRU), 2015 IEEE Workshop on, IEEE, 2015, pp. 504–511
2015
Cited alongside, same era.
doi:10.1109/TMM.2015.2407694
N. Harte, E. Gillen, Tcd-timit: An audio-visual corpus of continuous speech, IEEE Transactions on Multimedia 17 (5) (2015) 603–615 · 2015
Cited alongside, same era.
J. R. Hershey, Z. Chen, J. Le Roux, S. Watanabe, Deep clustering: Discriminative embeddings for segmentation and separation, in: 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2016, pp. 31–35
2016
Cited alongside, same era.
A. Adeel, M. Gogate, A. Hussain, Towards next-generation lip-reading driven hearing-aids: A preliminary prototype demo, in: International Workshop on Challenges in Hearing Assistive Technology (CHAT-2017), Stockholm University, August 19th, Collocated with Interspeech 2017, 2017
2017
Cited alongside, same era.
D. Yu, M. Kolbæk, Z.-H. Tan, J. Jensen, Permutation invariant training of deep models for speaker-independent multi-talker speech separation, in: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2017, pp. 241–245
2017
Cited alongside, same era.
S. Pascual, A. Bonafonte, J. Serrà, Segan: Speech enhancement generative adversarial network, Proc. Interspeech 2017 (2017) 3642–3646
2017
Cited alongside, same era.
D. Rethage, J. Pons, X. Serra, A wavenet for speech denoising, in: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2018, pp. 5069–5073
2018
Later among the works it cites.
A. Pandey, D. Wang, A new framework for supervised speech enhancement in the time domain., in: Interspeech, 2018, pp. 1136–1140
2018
Later among the works it cites.
S. Pascual, M. Park, J. Serrà, A. Bonafonte, K.-H. Ahn, Language and noise transfer in speech enhancement generative adversarial network, in: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2018, pp. 5019–5023
2018
Later among the works it cites.
A. Owens, A. A. Efros, Audio-visual scene analysis with self-supervised multisensory features, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 631–648
2018
Later among the works it cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, A. Torralba, The sound of pixels, in: The European Conference on Computer Vision (ECCV), 2018
2018
Later among the works it cites.
D. Wang, J. Chen, Supervised speech separation based on deep learning: An overview, IEEE/ACM Transactions on Audio, Speech, and Language Processing 26 (10) (2018) 1702–1726
2018
Later among the works it cites.
Y. Luo, N. Mesgarani, Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation, IEEE/ACM Transactions on Audio, Speech, and Language Processing 27 (8) (2019) 1256–1266
2019
Closest in time.
F. Developers, ffmpeg tool [software], http://ffmpeg.org/ (2000–2019)
2019
Closest in time.
J. Le Roux, S. Wisdom, H. Erdogan, J. R. Hershey, Sdr–half-baked or well done?, in: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2019, pp. 626–630
2019
Closest in time.