Fetching the paper…
Reading the bibliography…
Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos.
D. Griffin and J. S. Lim, “Signal estimation from modified short-time fourier transform,” in
1984
Earlier work this paper cites.
L. Girin, J.-L. Schwartz, and G. Feng, “Audio-visual enhancement of speech in noise,”
2001
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
S. Deligne, G. Potamianos, and C. Neti, “Audio-visual speech enhancement with avcdcn (audio-visual codebook dependent cepstral normalization),” in
2002
Earlier work this paper cites.
J. R. Hershey and M. Casey, “Audio-visual sound separation via hidden markov models,” in
2002
Earlier work this paper cites.
R. Goecke, G. Potamianos, and C. Neti, “Noisy audio feature enhancement using audio-visual speech data,” May 2002
2002
Earlier work this paper cites.
J. Hershey, H. Attias, N. Jojic, and T. Kristjansson, “Audio-visual graphical models for speech processing,” in
2004
Earlier work this paper cites.
W. Wang, D. Cosker, Y. Hicks, S. Saneit, and J. Chambers, “Video assisted speech source separation,” in
2005
Earlier work this paper cites.
C. Févotte, R. Gribonval, and E. Vincent, “BSS EVAL toolbox user guide,”
2005
Earlier work this paper cites.
A. M. Reddy and B. Raj, “Soft mask methods for single-channel speaker separation,”
2007
Earlier work this paper cites.
M. H. Radfar and R. M. Dansereau, “Single-channel speech separation using soft mask filtering,”
2007
Earlier work this paper cites.
S. Makino, T.-W. Lee, and H. Sawada,
2007
Earlier work this paper cites.
Z. Jin and D. Wang, “A supervised learning approach to monaural segregation of reverberant speech,”
2009
Earlier work this paper cites.
I. Almajai and B. P. Milner, “Effective visually-derived wiener filtering for audio-visual speech processing,” in
2009
Earlier work this paper cites.
M. Anusuya and S. K. Katti, “Speech recognition by machine, a review,”
2010
Earlier work this paper cites.
V. Emiya, E. Vincent, N. Harlander, and V. Hohmann, “Subjective and objective quality assessment of audio source separation,”
2011
Earlier work this paper cites.
C. Taal, R. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time-frequency weighted noisy speech,”
2011
Cited alongside, same era.
Q. Liu, W. Wang, P. J. Jackson, M. Barnard, J. Kittler, and J. Chambers, “Source separation of convolutive and noisy mixtures using audio-visual dictionary learning and probabilistic time-frequency masking,”
2013
Cited alongside, same era.
F. Khan and B. Milner, “Speaker separation using visually-derived binary masks,” in
2013
Cited alongside, same era.
B. Rivet, W. Wang, S. M. Naqvi, and J. A. Chambers, “Audiovisual speech source separation: An overview of key methodologies,”
2014
Cited alongside, same era.
P. Mowlaee and J. Kulmer, “Phase estimation in single-channel speech enhancement: Limits-potential,”
2015
Cited alongside, same era.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: an overview,”
2017
Later among the works it cites.
H.-G. Hirsch and M. Gref, “On the influence of modifying magnitude and phase spectrum to enhance noisy speech signals,” in
2017
Later among the works it cites.
2017
Later among the works it cites.
A. Ephrat, T. Halperin, and S. Peleg, “Improved speech reconstruction from silent video,” in
2017
Later among the works it cites.
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,”
2015
Cited alongside, same era.
2015
Cited alongside, same era.
P. Mowlaee, R. Saeidi, and Y. Stylianou, “Advances in phase-aware signal processing in speech communication,”
2016
Cited alongside, same era.
J. Fahringer, T. Schrank, J. Stahl, P. Mowlaee, and F. Pernkopf, “Phase-aware signal processing for automatic speech recognition,” in
2016
Cited alongside, same era.
D. S. Williamson, Y. Wang, and D. Wang, “Complex ratio masking for joint enhancement of magnitude and phase,” in
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in
2016
Cited alongside, same era.
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas, “Lipnet: Sentence-level lipreading,”
2016
Cited alongside, same era.
Later among the works it cites.
A. Gabbay, A. Shamir, and S. Peleg, “Visual Speech Enhancement using Noise-Invariant Training,”
2017
Later among the works it cites.
T. Stafylakis and G. Tzimiropoulos, “Combining Residual Networks with LSTMs for Lipreading,” in
2017
Later among the works it cites.
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” in
2017
Later among the works it cites.
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” in
2017
Later among the works it cites.
J. S. Chung and A. Zisserman, “Lip reading in profile,” in
2017
Later among the works it cites.
J.-C. Hou, S.-S. Wang, Y.-H. Lai, Y. Tsao, H.-W. Chang, and H.-M. Wang, “Audio-Visual Speech Enhancement Using Multimodal Deep Convolutional Neural Networks,”
2018
Closest in time.
2018
Closest in time.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,”
2018
Closest in time.
2018
Closest in time.