Fetching the paper…
Reading the bibliography…
Recent advances in deep learning have heightened interest among researchers in the field of visual speech recognition (VSR).
Visual contribution to speech intelligibility in noise
W. H. Sumby and I. Pollack · 1954
Earlier work this paper cites.
Which components of the face do humans and machines best speechread?
C. Benoit, T. Guiard-Marigny, B. Le Goff, and A. Adjoudani · 1996
Earlier work this paper cites.
Eye movement of perceivers during audiovisualspeech perception
E. Vatikiotis-Bateson, I.-M. Eigsti, S. Yano, and K. G. Munhall · 1998
Earlier work this paper cites.
Word identification and eye fixation locations in visual and visual-plus-auditory presentations of spoken sentences
C. R. Lansing and G. W. McConkie · 2003
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
M. Cooke, J. Barker, S. Cunningham, and X. Shao · 2006
Earlier work this paper cites.
Audio-visual speech perception off the top of the head
C. Davis and J. Kim · 2006
Earlier work this paper cites.
Seeing pitch: Visual information for lexical tones of mandarin-chinese
T. H. Chen and D. W. Massaro · 2008
Earlier work this paper cites.
Lipreading with local spatiotemporal descriptors
G. Zhao, M. Barnard, and M. Pietikäinen · 2009
Earlier work this paper cites.
Prosody off the top of the head: Prosodic contrasts can be discriminated by head motion
E. Cvejic, J. Kim, and C. Davis · 2010
Earlier work this paper cites.
Do the eyes really have it? dynamic allocation of attention when viewing moving faces
M. L.-H. Võ, T. J. Smith, P. K. Mital, and J. M. Henderson · 2012
Earlier work this paper cites.
Resolution limits on visual speech recognition
H. L. Bear, R. W. Harvey, B. Theobald, and Y. Lan · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
OuluVS2: A multi-view audiovisual database for non-rigid mouth motion analysis
I. Anina, Z. Zhou, G. Zhao, and M. Pietikäinen · 2015
Earlier work this paper cites.
Convolutional, long short-term memory, fully connected deep neural networks
T. N. Sainath, O. Vinyals, A. W. Senior, and H. Sak · 2015
Cited alongside, same era.
Optical flow based lip reading using non rectangular ROI and head motion reduction
J. Shiraishi and T. Saitoh · 2015
Cited alongside, same era.
Striving for simplicity: The all convolutional net
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. A. Riedmiller · 2015
Cited alongside, same era.
LipNet: End-to-end sentence-level lipreading
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas · 2016
Cited alongside, same era.
Lip reading in the wild
J. S. Chung and A. Zisserman · 2016
Cited alongside, same era.
You said that?
J. S. Chung, A. Jamaludin, and A. Zisserman · 2017
Cited alongside, same era.
The impact of reduced video quality on visual speech recognition
L. Dungan, A. Karaali, and N. Harte · 2018
Later among the works it cites.
Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein · 2018
Later among the works it cites.
Patch-gated CNN for occlusion-aware facial expression recognition
Y. Li, J. Zeng, S. Shan, and X. Chen · 2018
Later among the works it cites.
End-to-end audiovisual speech recognition
S. Petridis, T. Stafylakis, P. Ma, F. Cai, G. Tzimiropoulos, and M. Pantic · 2018
Later among the works it cites.
Pushing the boundaries of audiovisual word recognition using residual networks and LSTMs
T. Stafylakis, M. H. Khan, and G. Tzimiropoulos · 2018
Later among the works it cites.
Lip-Interact: Improving mobile device interaction with silent speech commands
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lip reading in profile
J. S. Chung and A. Zisserman · 2017
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
T. Devries and G. W. Taylor · 2017
Cited alongside, same era.
Exploring ROI size in deep learning based lipreading
A. Koumparoulis, G. Potamianos, Y. Mroueh, and S. J. Rennie · 2017
Cited alongside, same era.
Combining residual networks with LSTMs for lipreading
T. Stafylakis and G. Tzimiropoulos · 2017
Cited alongside, same era.
Random erasing data augmentation
Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang · 2017
Cited alongside, same era.
Deep audio-visual speech recognition
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman · 2018
Cited alongside, same era.
K. Sun, C. Yu, W. Shi, L. Liu, and Y. Shi · 2018
Later among the works it cites.
CBAM: convolutional block attention module
S. Woo, J. Park, J. Lee, and I. S. Kweon · 2018
Later among the works it cites.
LCANet: End-to-end lipreading with cascaded attention-ctc
K. Xu, D. Li, N. Cassimatis, and X. Wang · 2018
Later among the works it cites.
Robust remote heart rate estimation from face utilizing spatial-temporal attention
X. Niu, X. Zhao, H. Han, A. Das, A. Dantcheva, S. Shan, and X. Chen · 2019
Later among the works it cites.
SeetaFace2
SeetaTech · 2019
Later among the works it cites.
Words can shift: Dynamically adjusting word representations using nonverbal behaviors
Y. Wang, Y. Shen, Z. Liu, P. P. Liang, A. Zadeh, and L. Morency · 2019
Later among the works it cites.
LRW-1000: A naturally-distributed large-scale benchmark for lip reading in the wild
S. Yang, Y. Zhang, D. Feng, M. Yang, C. Wang, J. Xiao, K. Long, S. Shan, and X. Chen · 2019
Later among the works it cites.
A cascade sequence-to-sequence model for chinese mandarin lip reading
Y. Zhao, R. Xu, and M. Song · 2019
Later among the works it cites.