Fetching the paper…
Reading the bibliography…
We focus on the word-level visual lipreading, which requires recognizing the word being spoken, given only the video but not the audio.
Hearing Lips and Seeing Voices
H. McGurk and J. MacDonald · 1976
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional Recurrent Neural Networks
Mike Schuster and Kuldip K Paliwal · 1997
Earlier work this paper cites.
A Cascade Visual Front End for Speaker Independent Automatic Speechreading
Gerasimos Potamianos, Chalapathy Neti, and Giridharan Iyengar · 2001
Earlier work this paper cites.
CUAVE: A New Audio-Visual Database for Multimodal Human-Computer Interface Research
E.K. Patterson, S. Gurbuz, Z. Tufekci, and J.N. Gowdy · 2002
Earlier work this paper cites.
Recent Advances in the Automatic Recognition of Audiovisual Speech
Gerasimos Potamianos, Chalapathy Neti, Guillaume Gravier, Ashutosh Garg, and Andrew W. Senior · 2003
Earlier work this paper cites.
A PCA Based Visual DCT Feature Extraction Method for Lipreading
Hong Xiaopeng, Yao Hongxun, Wan Yuqi, and Chen Rong · 2006
Earlier work this paper cites.
Stream Weight Estimation for Multistream Audio-Visual Speech Recognition in a Multispeaker Environment
Xu Shao and Jon Barker · 2008
Earlier work this paper cites.
An Investigation into Features for Multi-View Lipreading
Adrian Pass, Jianguo Zhang, and Darryl Stewart · 2010
Earlier work this paper cites.
Convolutional Learning of Spatio-Temporal Features
Graham W Taylor, Rob Fergus, Yann Lecun, and Christoph Bregler · 2010
Earlier work this paper cites.
On Dynamic Stream Weighting for Audio-Visual Speech Recognition
Virginia Estellers, Mihai Gurban, and Jean-Philippe Thiran · 2012
Earlier work this paper cites.
Imagenet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
3D Convolutional Neural Networks for Human Action Recognition
Shuiwang Ji, Ming Yang, and Kai Yu · 2013
Earlier work this paper cites.
Large-Scale Video Classification with Convolutional Neural Networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Fei Fei Li · 2014
Earlier work this paper cites.
Lipreading Using Convolutional Neural Network
Kuniaki Noda, Yuki Yamaguchi, Kazuhiro Nakadai, Hiroshi G Okuno, and Tetsuya Ogata · 2014
Earlier work this paper cites.
Two-Stream Convolutional Networks for Action Recognition in Videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Robust Audio-Visual Speech Recognition under Noisy Audio-Video Conditions
Darryl Stewart, Rowan Seymour, Adrian Pass, and Ji Ming · 2014
Cited alongside, same era.
Blind Video Temporal Consistency
Nicolas Bonneel, James Tompkin, Kalyan Sunkavalli, Deqing Sun, Sylvain Paris, and Hanspeter Pfister · 2015
Cited alongside, same era.
Long-Term Recurrent Convolutional Networks for Visual Recognition and Description
Jeff Donahue, Lisa Anne Hendricks, Marcus Rohrbach, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, and Trevor Darrell · 2015
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Beyond Short Snippets: Deep Networks for Video Classification
Joe Yue Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici · 2015
Cited alongside, same era.
Lipreading with Long Short-Term Memory
Michael Wand, Jan Koutník, and Jürgen Schmidhuber · 2016
Later among the works it cites.
Multi-Stream Multi-Class Fusion of Deep Networks for Video Classification
Zuxuan Wu, Yu-gang Jiang, Xi Wang, Hao Ye, and Xiangyang Xue · 2016
Later among the works it cites.
LipNet: End-to-End Sentence-Level Lipreading
Yannis M Assael, Brendan Shillingford, Shimon Whiteson, and Nando De Freitas · 2017
Later among the works it cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Joao Carreira and Andrew Zisserman · 2017
Later among the works it cites.
Lip Reading Sentences in the Wild
Joon Son Chung, Andrew W Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Later among the works it cites.
An Audio-Visual Corpus for Multimodal Automatic Speech Recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hiroshi Ninomiya, Norihide Kitaoka, Satoshi Tamura, Yurie Iribe, and Kazuya Takeda · 2015
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Listening With Your Eyes: Towards a Practical Visual Speech Recognition System Using Deep Boltzmann Machines
Chao Sui, Mohammed Bennamoun, and Roberto Togneri · 2015
Cited alongside, same era.
Going Deeper with Convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Learning Spatiotemporal Features with 3D Convolutional Networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Cited alongside, same era.
Improved Speaker Independent Lip Reading Using Speaker Adaptive Training and Deep Neural Networks
Ibrahim Almajai, Stephen Cox, Richard Harvey, and Yuxuan Lan · 2016
Cited alongside, same era.
Convolutional Two-Stream Network Fusion for Video Action Recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Cited alongside, same era.
Andrzej Czyzewski, Bozena Kostek, Piotr Bratoszewski, Jozef Kotus, and Marcin Szykulski · 2017
Later among the works it cites.
Computer Vision for Autonomous Vehicles: Problems, Datasets and State-of-the-Art
Joel Janai, Fatma Güney, Aseem Behl, and Andreas Geiger · 2017
Later among the works it cites.
Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Later among the works it cites.
Combining Residual Networks with LSTMs for Lipreading
Georgios Tzimiropoulos Themos Stafylakis · 2017
Later among the works it cites.
Deep Audio-Visual Speech Recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2018
Later among the works it cites.
Learning to Lip Read Words by Watching Videos
Joon Son Chung and Andrew Zisserman · 2018
Later among the works it cites.
T-C3D: Temporal Convolutional 3D Network for Real-time Action Recognition
Kun Liu, Wu Liu, Chuang Gan, Mingkui Tan, and Huadong Ma · 2018
Later among the works it cites.
End-to-End Audiovisual Speech Recognition
Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Feipeng Cai, Georgios Tzimiropoulos, and Maja Pantic · 2018
Later among the works it cites.
PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz · 2018
Later among the works it cites.
Hidden Two-Stream Convolutional Networks for Action Recognition
Yi Zhu, Zhenzhong Lan, Shawn Newsam, and Alexander G. Hauptmann · 2018
Later among the works it cites.
Future Near-Collision Prediction from Monocular Video: Feasibility, Dataset, and Challenges
Aashi Manglik, Xinshuo Weng, Eshed Ohn-bar, and Kris M Kitani · 2019
Closest in time.