Fetching the paper…
Reading the bibliography…
The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio.
Perceptual linear predictive (plp) analysis of speech
H. Hermansky · 1990
Earlier work this paper cites.
Selecting and tracking features for image sequence analysis
C. Tomasi and T. Kanade · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Reliable transition detection in videos: A survey and practitioner’s guide
R. Lienhart · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
M. Cooke, J. Barker, S. Cunningham, and X. Shao · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Pascal visual object classes challenge
D. Larlus, G. Dorko, D. Jurie, and B. Triggs · 2006
Earlier work this paper cites.
Speaker identification on the scotus corpus
J. Yuan and M. Liberman · 2008
Earlier work this paper cites.
Dlib-ml: A machine learning toolkit
D. E. King · 2009
Earlier work this paper cites.
Comparing visual features for lipreading
Y. Lan, R. Harvey, B. Theobald, E.-J. Ong, and R. Bowden · 2009
Earlier work this paper cites.
The Oxford handbook of deaf studies, language, and education
M. Marschark and P. E. Spencer · 2010
Earlier work this paper cites.
Audio-visual speech recognition incorporating facial depth information captured by the kinect
G. Galatas, G. Potamianos, and F. Makedon · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Hybrid speech recognition with deep bidirectional lstm
A. Graves, N. Jaitly, and A.-r. Mohamed · 2013
Cited alongside, same era.
Return of the devil in the details: Delving deep into convolutional nets
K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Cited alongside, same era.
End-to-end continuous speech recognition using attention-based recurrent nn: first results
J. Chorowski, D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Cited alongside, same era.
One millisecond face alignment with an ensemble of regression trees
V. Kazemi and J. Sullivan · 2014
Cited alongside, same era.
Lipreading using convolutional neural network
Audio-visual speech recognition using deep learning
K. Noda, Y. Yamaguchi, K. Nakadai, H. G. Okuno, and T. Ogata · 2015
Later among the works it cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, S. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. Berg, and F. Li · 2015
Later among the works it cites.
Improving neural machine translation models with monolingual data
R. Sennrich, B. Haddow, and A. Birch · 2015
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Later among the works it cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Later among the works it cites.
Audio-visual speech recognition using deep bottleneck features and high-performance lipreading
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Noda, Y. Yamaguchi, K. Nakadai, H. G. Okuno, and T. Ogata · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. Le · 2014
Cited alongside, same era.
A review of recent advances in visual speech decoding
Z. Zhou, G. Zhao, X. Hong, and M. Pietikäinen · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Cited alongside, same era.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals · 2015
Cited alongside, same era.
Attention-based models for speech recognition
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
S. Tamura, H. Ninomiya, N. Kitaoka, S. Osuga, Y. Iribe, K. Takeda, and S. Hayamizu · 2015
Later among the works it cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al · 2016
Closest in time.
Lipnet: Sentence-level lipreading
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas · 2016
Closest in time.
Lip reading in the wild
J. S. Chung and A. Zisserman · 2016
Closest in time.
Out of time: automated lip sync in the wild
J. S. Chung and A. Zisserman · 2016
Closest in time.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Closest in time.
Deep complementary bottleneck features for visual speech recognition
S. Petridis and M. Pantic · 2016
Closest in time.
Lipreading with long short-term memory
M. Wand, J. Koutn, et al · 2016
Closest in time.