Fetching the paper…
Reading the bibliography…
When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” in
1941
Earlier work this paper cites.
J. Lim and A. Oppenheim, “All-pole modeling of degraded speech,”
1978
Earlier work this paper cites.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,”
1984
Earlier work this paper cites.
Y. Ephraim, “Statistical-model-based speech enhancement systems,”
1992
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “Darpa timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,”
1993
Earlier work this paper cites.
P. Scalart
1996
Earlier work this paper cites.
A. W. Bronkhorst, “The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,”
2000
Earlier work this paper cites.
L. Girin, J. L. Schwartz, and G. Feng, “Audio-visual enhancement of speech in noise,”
2001
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
S. Parveen and P. Green, “Speech enhancement with missing data techniques using recurrent neural networks,” in
2004
Earlier work this paper cites.
L.-P. Yang and Q.-J. Fu, “Spectral subtraction-based speech enhancement for cochlear implant patients in background noise,”
2005
Earlier work this paper cites.
F. L. Chong, I. V. McLoughlin, and K. Pawlikowski, “A methodology for improving PESQ accuracy for Chinese speech,” in
2005
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition,”
2006
Earlier work this paper cites.
D. Yu, L. Deng, J. Droppo, J. Wu, Y. Gong, and A. Acero, “A minimum-mean-square-error noise reduction algorithm on mel-frequency cepstra for robust speech recognition,” in
2008
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in
2011
Cited alongside, same era.
A. L. Maas, Q. V. Le, T. M. O’Neil, O. Vinyals, P. Nguyen, and A. Y. Ng, “Recurrent neural networks for noise reduction in robust asr,” in
2012
Cited alongside, same era.
P. C. Loizou,
2013
Cited alongside, same era.
F. Khan and B. Milner, “Speaker separation using visually-derived binary masks,” in
2013
Cited alongside, same era.
V. Kazemi and J. Sullivan, “One millisecond face alignment with an ensemble of regression trees,” in
2014
Cited alongside, same era.
N. Harte and E. Gillen, “Tcd-timit: An audio-visual corpus of continuous speech,”
2015
Cited alongside, same era.
T. L. Cornu and B. Milner, “Generating intelligible audio speech from visual speech,” in
2017
Closest in time.
A. Ephrat, T. Halperin, and S. Peleg, “Improved speech reconstruction from silent video,” in
2017
Closest in time.
A. Ephrat and S. Peleg, “Vid2speech: speech reconstruction from silent video,” in
2017
Closest in time.
Z. Chen, “Single channel auditory source separation with neural network,” Ph.D. dissertation, Columbia Univ., 2017
2017
Closest in time.
M. Kolbæk, Z.-H. Tan, and J. Jensen, “Speech intelligibility potential of general and specialized deep neural network based speech enhancement systems,”
2017
Closest in time.
S. Pascual, A. Bonafonte, and J. Serrà, “Segan: Speech enhancement generative adversarial network,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Noda, Y. Yamaguchi, K. Nakadai, H. G. Okuno, and T. Ogata, “Audio-visual speech recognition using deep learning,”
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Cited alongside, same era.
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas, “Lipnet: Sentence-level lipreading,”
2016
Cited alongside, same era.
Y. Isik, J. L. Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” in
2016
Cited alongside, same era.
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman, “Visually indicated sounds,” in
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in
2016
Cited alongside, same era.
2017
Closest in time.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, D. Eck, K. Simonyan, and M. Norouzi, “Neural audio synthesis of musical notes with wavenet autoencoders,” in
2017
Closest in time.
E. M. Grais and M. D. Plumbley, “Single channel audio source separation using convolutional denoising autoencoders,” in
2017
Closest in time.
J.-C. Hou, S.-S. Wang, Y.-H. Lai, J.-C. Lin, Y. Tsao, H.-W. Chang, and H.-M. Wang, “Audio-visual speech enhancement based on multimodal deep convolutional neural network,”
2018
Closest in time.
A. Gabbay, A. Ephrat, T. Halperin, and S. Peleg, “Seeing through noise: Speaker separation and enhancement using visually-derived speech,” in
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.