Fetching the paper…
Reading the bibliography…
Lipreading is the task of decoding text from the movement of a speaker's mouth.
The story of lip-reading, its genesis and development
F. DeLand · 1931
Earlier work this paper cites.
Phoneme perception in lipreading
M. F. Woodward and C. G. Barber · 1960
Earlier work this paper cites.
Confusions among visually perceived consonants
C. G. Fisher · 1968
Earlier work this paper cites.
Hearing lips and seeing voices
H. McGurk and J. MacDonald · 1976
Earlier work this paper cites.
Perceptual dominance during lipreading
R. D. Easton and M. Basala · 1982
Earlier work this paper cites.
Continuous automatic speech recognition by lipreading
A. J. Goldschen, O. N. Garcia, and E. D. Petajan · 1997
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Audio visual speech recognition
C. Neti, G. Potamianos, J. Luettin, I. Matthews, H. Glotin, D. Vergyri, J. Sison, and A. Mashari · 2000
Earlier work this paper cites.
Extraction of visual features for lipreading
I. Matthews, T. F. Cootes, J. A. Bangham, S. Cox, and R. Harvey · 2002
Earlier work this paper cites.
Recent advances in the automatic recognition of audiovisual speech
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A. W. Senior · 2003
Earlier work this paper cites.
Framewise phoneme classification with bidirectional LSTM and other neural network architectures
A. Graves and J. Schmidhuber · 2005
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
M. Cooke, J. Barker, S. Cunningham, and X. Shao · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Patch-based representation of visual speech
P. Lucey and S. Sridharan · 2006
Earlier work this paper cites.
Adaptive multimodal fusion by uncertainty compensation
V. Pitsikalis, A. Katsamanis, G. Papandreou, and P. Maragos · 2006
Earlier work this paper cites.
Multimodal fusion and learning with uncertain features applied to audiovisual speech recognition
G. Papandreou, A. Katsamanis, V. Pitsikalis, and P. Maragos · 2007
Earlier work this paper cites.
Classification and feature extraction by simplexization
Y. Fu, S. Yan, and T. S. Huang · 2008
Earlier work this paper cites.
Information theoretic feature extraction for audio-visual speech recognition
M. Gurban and J.-P. Thiran · 2009
Earlier work this paper cites.
Comparison of human and machine-based lip-reading
S. Hilder, R. Harvey, and B.-J. Theobald · 2009
Cited alongside, same era.
Dlib-ml: A machine learning toolkit
D. E. King · 2009
Cited alongside, same era.
Adaptive multimodal fusion by uncertainty compensation with application to audiovisual speech recognition
G. Papandreou, A. Katsamanis, V. Pitsikalis, and P. Maragos · 2009
Cited alongside, same era.
Lipreading with local spatiotemporal descriptors
G. Zhao, M. Barnard, and M. Pietikainen · 2009
Cited alongside, same era.
Formant frequencies of vowels in 13 accents of the british isles
E. Ferragne and F. Pellegrino · 2010
Cited alongside, same era.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Cited alongside, same era.
Lipreading using convolutional neural network
K. Noda, Y. Yamaguchi, K. Nakadai, H. G. Okuno, and T. Ogata · 2014
Later among the works it cites.
Striving for simplicity: The all convolutional net
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller · 2014
Later among the works it cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Later among the works it cites.
A review of recent advances in visual speech decoding
Z. Zhou, G. Zhao, X. Hong, and M. Pietikäinen · 2014
Later among the works it cites.
Deep Speech 2: End-to-end speech recognition in English and Mandarin
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, et al · 2015
Later among the works it cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition
G. E. Dahl, D. Yu, L. Deng, and A. Acero · 2012
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Understanding phonetics
P. Ashby · 2013
Cited alongside, same era.
3d convolutional neural networks for human action recognition
S. Ji, W. Xu, M. Yang, and K. Yu · 2013
Cited alongside, same era.
300 faces in-the-wild challenge: The first facial landmark localization challenge
C. Sagonas, G. Tzimiropoulos, S. Zafeiriou, and M. Pantic · 2013
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Later among the works it cites.
Deep learning of mouth shapes for sign language
O. Koller, H. Ney, and R. Bowden · 2015
Later among the works it cites.
Lexicon-free conversational speech recognition with neural networks
A. L. Maas, Z. Xie, D. Jurafsky, and A. Y. Ng · 2015
Later among the works it cites.
Integration of deep bottleneck features for audio-visual speech recognition
H. Ninomiya, N. Kitaoka, S. Tamura, Y. Iribe, and K. Takeda · 2015
Later among the works it cites.
Listening with your eyes: Towards a practical visual speech recognition system using deep boltzmann machines
C. Sui, M. Bennamoun, and R. Togneri · 2015
Later among the works it cites.
Improved speaker independent lip reading using speaker adaptive training and deep neural networks
I. Almajai, S. Cox, R. Harvey, and Y. Lan · 2016
Closest in time.
Lip reading using CNN and LSTM
A. Garg, J. Noyola, and S. Bagadia · 2016
Closest in time.
Dynamic stream weighting for turbo-decoding-based audiovisual ASR
S. Gergen, S. Zeiler, A. H. Abdelaziz, R. Nickel, and D. Kolossa · 2016
Closest in time.
Temporal multimodal learning in audiovisual speech recognition
D. Hu, X. Li, et al · 2016
Closest in time.
Deep complementary bottleneck features for visual speech recognition
S. Petridis and M. Pantic · 2016
Closest in time.
Audio-visual speech recognition using bimodal-trained bottleneck features for a person with severe hearing loss
Y. Takashima, R. Aihara, T. Takiguchi, Y. Ariki, N. Mitani, K. Omori, and K. Nakazono · 2016
Closest in time.
Lipreading with long short-term memory
M. Wand, J. Koutnik, and J. Schmidhuber · 2016
Closest in time.