Fetching the paper…
Reading the bibliography…
The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio.
An iterative image registration technique with an application to stereo vision
B. D. Lucas and T. Kanade · 1981
Earlier work this paper cites.
Reliable transition detection in videos: A survey and practitioner’s guide
R. Lienhart · 2001
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
M. Cooke, J. Barker, S. Cunningham, and X. Shao · 2006
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Speaker identification on the scotus corpus
J. Yuan and M. Liberman · 2008
Earlier work this paper cites.
Dlib-ml: A machine learning toolkit
D. E. King · 2009
Earlier work this paper cites.
Audio-visual speech recognition incorporating facial depth information captured by the kinect
G. Galatas, G. Potamianos, and F. Makedon · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
A. Graves · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition
G. Hinton, L. Deng, D. Yu, G. Dahl, A.-R. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, B. Kingsbury, and T. Sainath · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
End-to-end continuous speech recognition using attention-based recurrent NN: first results
J. Chorowski, D. Bahdanau, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Earlier work this paper cites.
Lipreading using convolutional neural network
K. Noda, Y. Yamaguchi, K. Nakadai, H. G. Okuno, and T. Ogata · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. Le · 2014
Earlier work this paper cites.
A review of recent advances in visual speech decoding
Z. Zhou, G. Zhao, X. Hong, and M. Pietikäinen · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
ADAM: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Deep learning of mouth shapes for sign language
O. Koller, H. Ney, and R. Bowden · 2015
Cited alongside, same era.
Lexicon-free conversational speech recognition with neural networks
A. L. Maas, Z. Xie, D. Jurafsky, and A. Y. Ng · 2015
Cited alongside, same era.
Deep multimodal learning for audio-visual speech recognition
Y. Mroueh, E. Marcheret, and V. Goel · 2015
Cited alongside, same era.
Audio-visual speech recognition using deep learning
K. Noda, Y. Yamaguchi, K. Nakadai, H. G. Okuno, and T. Ogata · 2015
Cited alongside, same era.
Lipreading with long short-term memory
M. Wand, J. Koutn, and J. Schmidhuber · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, J. Klingner, A. Shah, M. Johnson, X. Liu, L. Kaiser, S. Gouws, Y. Kato, T. Kudo, H. Kazawa, K. Stevens, G. Kurian, N. Patil, W. Wang, C. Young, J. Smith, J. Riesa, A. Rudnick, O. Vinyals, G. Corrado, M. Hughes, and J. Dean · 2016
Later among the works it cites.
State-of-the-art speech recognition with sequence-to-sequence models
C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, K. Gonina, N. Jaitly, B. Li, J. Chorowski, and M. Bacchiani · 2017
Later among the works it cites.
Lip reading sentences in the wild
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman · 2017
Later among the works it cites.
Lip reading in profile
J. S. Chung and A. Zisserman · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, S. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. Berg, and F. Li · 2015
Cited alongside, same era.
Fast and accurate recurrent neural network acoustic models for speech recognition
H. Sak, A. W. Senior, K. Rao, and F. Beaufays · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Audio-visual speech recognition using deep bottleneck features and high-performance lipreading
S. Tamura, H. Ninomiya, N. Kitaoka, S. Osuga, Y. Iribe, K. Takeda, and S. Hayamizu · 2015
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al · 2016
Cited alongside, same era.
Later among the works it cites.
An audio-visual corpus for multimodal automatic speech recognition
A. Czyzewski, B. Kostek, P. Bratoszewski, J. Kotus, and M. Szykulski · 2017
Later among the works it cites.
An analysis of incorporating an external language model into a sequence-to-sequence model
A. Kannan, Y. Wu, P. Nguyen, T. N. Sainath, Z. Chen, and R. Prabhavalkar · 2017
Later among the works it cites.
Letter-based speech recognition with gated convnets
V. Liptchinsky, G. Synnaeve, and R. Collobert · 2017
Later among the works it cites.
Combining residual networks with LSTMs for lipreading
T. Stafylakis and G. Tzimiropoulos · 2017
Later among the works it cites.
Attention Is All You Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Residual Convolutional CTC Networks for Automatic Speech Recognition
Y. Wang, X. Deng, S. Pu, and Z. Huang · 2017
Later among the works it cites.
Learning filterbanks from raw speech for phone recognition
N. Zeghidour, N. Usunier, I. Kokkinos, T. Schatz, G. Synnaeve, and E. Dupoux · 2017
Later among the works it cites.
Towards end-to-end speech recognition with deep convolutional neural networks
Y. Zhang, M. Pezeshki, P. Brakel, S. Zhang, C. Laurent, Y. Bengio, and A. C. Courville · 2017
Later among the works it cites.
Deep lip reading: A comparison of models and an online application
T. Afouras, J. S. Chung, and A. Zisserman · 2018
Closest in time.
LRS3-TED: a large-scale dataset for visual speech recognition
T. Afouras, J. S. Chung, and A. Zisserman · 2018
Closest in time.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
S. Bai, J. Z. Kolter, and V. Koltun · 2018
Closest in time.
End-to-end audiovisual speech recognition
S. Petridis, T. Stafylakis, P. Ma, F. Cai, G. Tzimiropoulos, and M. Pantic · 2018
Closest in time.
Large-Scale Visual Speech Recognition
B. Shillingford, Y. Assael, M. W. Hoffman, T. Paine, C. Hughes, U. Prabhu, H. Liao, H. Sak, K. Rao, L. Bennett, M. Mulville, B. Coppin, B. Laurie, A. Senior, and N. de Freitas · 2018
Closest in time.