Fetching the paper…
Reading the bibliography…
Isolating the voice of a specific person while filtering out other voices or background noises is challenging when video is shot in noisy environments.
“Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,”
Yariv Ephraim and David Malah, · 1984
Earlier work this paper cites.
“Darpa timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,”
John S Garofolo, Lori F Lamel, William M Fisher, Jonathon G Fiscus, and David S Pallett, · 1993
Earlier work this paper cites.
“The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,”
Adelbert W Bronkhorst, · 2000
Earlier work this paper cites.
“Audio-visual enhancement of speech in noise,”
L Girin, J L Schwartz, and G Feng, · 2001
Earlier work this paper cites.
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, · 2001
Earlier work this paper cites.
“Robust real-time face detection,”
Paul Viola and Michael J Jones, · 2004
Earlier work this paper cites.
“Video assisted speech source separation,”
Wenwu Wang, Darren Cosker, Yulia Hicks, S Saneit, and Jonathon Chambers, · 2005
Earlier work this paper cites.
“An audio-visual corpus for speech perception and automatic speech recognition,”
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao, · 2006
Earlier work this paper cites.
“Performance measurement in blind audio source separation,”
Emmanuel Vincent, Rémi Gribonval, and Cédric Févotte, · 2006
Earlier work this paper cites.
“Soft mask methods for single-channel speaker separation,”
Aarthi M Reddy and Bhiksha Raj, · 2007
Cited alongside, same era.
“A supervised learning approach to monaural segregation of reverberant speech,”
Zhaozhang Jin and DeLiang Wang, · 2009
Cited alongside, same era.
“Multimodal deep learning,”
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng, · 2011
Cited alongside, same era.
“Speaker separation using visually-derived binary masks,”
Faheem Khan and Ben Milner, · 2013
Cited alongside, same era.
“Deep learning for monaural speech separation,”
Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, and Paris Smaragdis, · 2014
Cited alongside, same era.
“Tcd-timit: An audio-visual corpus of continuous speech,”
Naomi Harte and Eoin Gillen, · 2015
Cited alongside, same era.
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2016
Later among the works it cites.
Audio-visual speaker separation
Faheem Khan, · 2016
Later among the works it cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Later among the works it cites.
Single Channel auditory source separation with neural network
Zhuo Chen, · 2017
Closest in time.
“Vid2speech: speech reconstruction from silent video,”
Ariel Ephrat and Shmuel Peleg, · 2017
Closest in time.
“Improved speech reconstruction from silent video,”
Ariel Ephrat, Tavi Halperin, and Shmuel Peleg, · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R Hershey, · 2016
Cited alongside, same era.
“Lip reading sentences in the wild,”
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman, · 2016
Cited alongside, same era.
“Visually indicated sounds,”
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman, · 2016
Cited alongside, same era.
“Generating intelligible audio speech from visual speech,”
Thomas Le Cornu and Ben Milner, · 2017
Closest in time.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbaek, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Closest in time.
“Audio-visual speech enhancement based on multimodal deep convolutional neural network,”
Jen-Cheng Hou, Syu-Siang Wang, Ying-Hui Lai, Jen-Chun Lin, Yu Tsao, Hsiu-Wen Chang, and Hsin-Min Wang, · 2017
Closest in time.