Fetching the paper…
Reading the bibliography…
We propose an end-to-end deep learning architecture for word-level visual speech recognition.
C. G. Fisher, “Confusions among visually perceived consonants,”
1968
Earlier work this paper cites.
A. J. Goldschen, O. N. Garcia, and E. D. Petajan, “Continuous automatic speech recognition by lipreading,” in
1997
Earlier work this paper cites.
G. I. Chiou and J.-N. Hwang, “Lipreading from color video,”
1997
Earlier work this paper cites.
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A. W. Senior, “Recent advances in the automatic recognition of audiovisual speech,”
2003
Earlier work this paper cites.
G. Potamianos, C. Neti, J. Luettin, and I. Matthews, “Audio-visual automatic speech recognition: An overview,”
2004
Earlier work this paper cites.
A. Graves, S. Fernández, and J. Schmidhuber, “Bidirectional LSTM networks for improved phoneme classification and recognition,” in
2005
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition,”
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
C. Chandrasekaran, A. Trubanova, S. Stillittano, A. Caplier, and A. A. Ghazanfar, “The natural statistics of audiovisual speech,”
2009
Earlier work this paper cites.
G. Papandreou, A. Katsamanis, V. Pitsikalis, and P. Maragos, “Adaptive multimodal fusion by uncertainty compensation with application to audiovisual speech recognition,”
2009
Earlier work this paper cites.
A. A. Shaikh, D. K. Kumar, W. C. Yau, M. C. Azemin, and J. Gubbi, “Lip reading using optical flow and support vector machines,” in
2010
Earlier work this paper cites.
R. Collobert, K. Kavukcuoglu, and C. Farabet, “Torch7: A matlab-like environment for machine learning,” in
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath
2012
Cited alongside, same era.
H. L. Bear and R. Harvey, “Decoding visemes: improving machine lip-reading,” in
2013
Cited alongside, same era.
J. Huang and B. Kingsbury, “Audio-visual deep learning for noise robust speech recognition,” in
2013
Cited alongside, same era.
S. Bengio and G. Heigold, “Word embeddings for speech recognition.” in
2014
Cited alongside, same era.
Z. Zhou, G. Zhao, X. Hong, and M. Pietikäinen, “A review of recent advances in visual speech decoding,”
2014
Cited alongside, same era.
K. Noda, Y. Yamaguchi, K. Nakadai, H. G. Okuno, and T. Ogata, “Audio-visual speech recognition using deep learning,”
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in
2016
Later among the works it cites.
2016
Later among the works it cites.
I. Almajai, S. Cox, R. Harvey, and Y. Lan, “Improved speaker independent lip reading using speaker adaptive training and deep neural networks,” in
2016
Later among the works it cites.
S. Petridis and M. Pantic, “Deep complementary bottleneck features for visual speech recognition,” in
2016
Later among the works it cites.
A. Bulat and G. Tzimiropoulos, “Convolutional aggregation of local evidence for large pose face alignment,” in
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
K. Thangthai, R. W. Harvey, S. J. Cox, and B.-J. Theobald, “Improving lip-reading performance for robust audiovisual speech recognition using DNNs,” in
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Cited alongside, same era.
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention.” in
2015
Cited alongside, same era.
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas, “Lipnet: Sentence-level lipreading,”
2016
Cited alongside, same era.
M. Wand, J. Koutník, and J. Schmidhuber, “Lipreading with long short-term memory,” in
2016
Cited alongside, same era.
——, “Two-stage convolutional part heatmap regression for the 1st 3D face alignment in the wild (3DFAW) challenge,” in
2016
Later among the works it cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Later among the works it cites.
——, “Identity mappings in deep residual networks,” in
2016
Later among the works it cites.
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,”
2017
Closest in time.
2017
Closest in time.
E. K. Patterson, S. Gurbuz, Z. Tufekci, and J. N. Gowdy, “Cuave: A new audio-visual database for multimodal human-computer interface research,” in
2017
Closest in time.