Fetching the paper…
Reading the bibliography…
We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise.
Some experiments on the recognition of speech, with one and with two ears
E Colin Cherry. 1953 · 1953
Earlier work this paper cites.
TIMIT Acoustic-phonetic Continuous Speech Corpus
J S Garofolo, Lori Lamel, W M Fisher, Jonathan Fiscus, D S. Pallett, N L. Dahlgren, and V Zue. 1992 · 1992
Earlier work this paper cites.
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs. In Acoustics, Speech, and Signal Processing, 2001. Proceedings.(ICASSP’01). 2001 IEEE International Conference on
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra. 2001 · 2001
Earlier work this paper cites.
Audio-visual sound separation via hidden Markov models. In Advances in Neural Information Processing Systems
John R Hershey and Michael Casey. 2002 · 2002
Earlier work this paper cites.
Moving-Talker, Speaker-Independent Feature Study, and Baseline Results Using the CUAVE Multimodal Speech Corpus
Eric K. Patterson, Sabri Gurbuz, Zekeriya Tufekci, and John N. Gowdy. 2002 · 2002
Earlier work this paper cites.
Audio-visual graphical models for speech processing. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
John Hershey, Hagai Attias, Nebojsa Jojic, and Trausti Kristjansson. 2004 · 2004
Earlier work this paper cites.
Performance Measurement in Blind Audio Source Separation
E. Vincent, R. Gribonval, and C. Fevotte. 2006 · 2006
Earlier work this paper cites.
Object detectors emerge in deep scene cnns
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. 2014 · 2006
Earlier work this paper cites.
Lip-Reading Aids Word Recognition Most in Moderate Noise: A Bayesian Explanation Using High-Dimensional Feature Space
Wei Ji Ma, Xiang Zhou, Lars A. Ross, John J. Foxe, and Lucas C. Parra. 2009 · 2009
Earlier work this paper cites.
The cocktail party problem
Josh H McDermott. 2009 · 2009
Earlier work this paper cites.
Blind audiovisual source separation based on sparse redundant representations
Anna Llagostera Casanovas, Gianluca Monaci, Pierre Vandergheynst, and Rémi Gribonval. 2010 · 2010
Earlier work this paper cites.
Handbook of Blind Source Separation: Independent component analysis and applications
Pierre Comon and Christian Jutten. 2010 · 2010
Earlier work this paper cites.
A short-time objective intelligibility measure for time-frequency weighted noisy speech. In Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on
Cees H Taal, Richard C Hendriks, Richard Heusdens, and Jesper Jensen. 2010 · 2010
Earlier work this paper cites.
Speech Intelligibility Prediction Using a Neurogram Similarity Index Measure
Andrew Hines and Naomi Harte. 2012 · 2011
Earlier work this paper cites.
Towards real-time audiovisual speaker localization. In Signal Processing Conference, 2011 19th European
Gianluca Monaci. 2011 · 2011
Earlier work this paper cites.
Multimodal Deep Learning. In ICML
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng. 2011 · 2011
Earlier work this paper cites.
Visual input enhances selective speech envelope tracking in auditory cortex at a "cocktail party"
Elana Zion Golumbic, Gregory B Cogan, Charles E. Schroeder, and David Poeppel. 2013 · 2013
Earlier work this paper cites.
The second ’chime’ speech separation and recognition challenge: Datasets, tasks and baselines
Emmanuel Vincent, Jon Barker, Shinji Watanabe, Jonathan Le Roux, Francesco Nesta, and Marco Matassoni. 2013 · 2013
Earlier work this paper cites.
Audiovisual Speech Source Separation: An overview of key methodologies
Bertrand Rivet, Wenwu Wang, Syed M. Naqvi, and Jonathon A. Chambers. 2014 · 2014
Cited alongside, same era.
On training targets for supervised speech separation
Yuxuan Wang, Arun Narayanan, and DeLiang Wang. 2014 · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks. In European conference on computer vision
Matthew D Zeiler and Rob Fergus. 2014 · 2014
Cited alongside, same era.
Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks
Hakan Erdogan, John R. Hershey, Shinji Watanabe, and Jonathan Le Roux. 2015 · 2015
Cited alongside, same era.
TCD-TIMIT: An audio-visual corpus of continuous speech
Naomi Harte and Eoin Gillen. 2015 · 2015
Cited alongside, same era.
ViSQOLAudio: An objective audio quality metric for low bitrate codecs
Oracle performance investigation of the ideal masks. In IWAENC
Ziteng Wang, Xiaofei Wang, Xu Li, Qiang Fu, and Yonghong Yan. 2016 · 2016
Later among the works it cites.
Features for Masking-Based Monaural Speech Separation in Reverberant Conditions
Masood Delfarah and DeLiang Wang. 2017 · 2017
Later among the works it cites.
Improved Speech Reconstruction from Silent Video. In ICCV 2017 Workshop on Computer Vision for Audio-Visual Media
Ariel Ephrat, Tavi Halperin, and Shmuel Peleg. 2017 · 2017
Later among the works it cites.
Audio visual speech recognition with multimodal recurrent neural networks. In Neural Networks (IJCNN), 2017 International Joint Conference on
Weijiang Feng, Naiyang Guan, Yuan Li, Xiang Zhang, and Zhigang Luo. 2017 · 2017
Later among the works it cites.
Visual Speech Enhancement using Noise-Invariant Training
Aviv Gabbay, Asaph Shamir, and Shmuel Peleg. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrew Hines, Eoin Gillen, Damien Kelly, Jan Skoglund, Anil C. Kokaram, and Naomi Harte. 2015 · 2015
Cited alongside, same era.
Deep multimodal speaker naming. In Proceedings of the 23rd ACM international conference on Multimedia
Yongtao Hu, Jimmy SJ Ren, Jingwen Dai, Chang Yuan, Li Xu, and Wenping Wang. 2015 · 2015
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In ICML
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Cited alongside, same era.
Deep multimodal learning for audio-visual speech recognition. In Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on
Youssef Mroueh, Etienne Marcheret, and Vaibhava Goel. 2015 · 2015
Cited alongside, same era.
Speech Enhancement with LSTM Recurrent Neural Networks and its Application to Noise-Robust ASR. In LVA/ICA
Felix Weninger, Hakan Erdogan, Shinji Watanabe, Emmanuel Vincent, Jonathan Le Roux, John R. Hershey, and Björn W. Schuller. 2015 · 2015
Cited alongside, same era.
Lip Reading Sentences in the Wild
Joon Son Chung, Andrew W. Senior, Oriol Vinyals, and Andrew Zisserman. 2016 · 2016
Cited alongside, same era.
Synthesizing normalized faces from facial identity features. In CVPR’17
Forrester Cole, David Belanger, Dilip Krishnan, Aaron Sarna, Inbar Mosseri, and William T Freeman. 2016 · 2016
Cited alongside, same era.
Audio Set: An ontology and human-labeled dataset for audio events. In Proc. IEEE ICASSP 2017
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Later among the works it cites.
Putting a Face to the Voice: Fusing Audio and Visual Signals Across a Video to Determine Speakers
Ken Hoover, Sourish Chaudhuri, Caroline Pantofaru, Malcolm Slaney, and Ian Sturdy. 2017 · 2017
Later among the works it cites.
Audio-visual object localization and separation using low-rank and sparsity. In Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on
Jie Pu, Yannis Panagakis, Stavros Petridis, and Maja Pantic. 2017 · 2017
Later among the works it cites.
Multiple-target deep learning for LSTM-RNN based speech enhancement. In HSCMA
Lei Sun, Jun Du, Li-Rong Dai, and Chin-Hui Lee. 2017 · 2017
Later among the works it cites.
Supervised Speech Separation Based on Deep Learning: An Overview
DeLiang Wang and Jitong Chen. 2017 · 2017
Later among the works it cites.
Permutation invariant training of deep models for speaker-independent multi-talker speech separation
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen. 2017 · 2017
Later among the works it cites.
The Conversation: Deep Audio-Visual Speech Enhancement. In arXiv:1804.04121
T. Afouras, J. S. Chung, and A. Zisserman. 2018 · 2018
Closest in time.
Seeing Through Noise: Speaker Separation and Enhancement using Visually-derived Speech
Aviv Gabbay, Ariel Ephrat, Tavi Halperin, and Shmuel Peleg. 2018 · 2018
Closest in time.
Learning to Separate Object Sounds by Watching Unlabeled Video
R. Gao, R. Feris, and K. Grauman. 2018 · 2018
Closest in time.
Audio-Visual Speech Enhancement Using Multimodal Deep Convolutional Neural Networks
Jen-Cheng Hou, Syu-Siang Wang, Ying-Hui Lai, Jen-Chun Lin, Yu Tsao, Hsiu-Wen Chang, and Hsin-Min Wang. 2018 · 2018
Closest in time.
Audio-Visual Scene Analysis with Self-Supervised Multisensory Features
Andrew Owens and Alexei A Efros. 2018 · 2018
Closest in time.
The Sound of Pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. 2018 · 2018
Closest in time.