Fetching the paper…
Reading the bibliography…
The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos.
Recent advances in the automatic recognition of audiovisual speech
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A.W. Senior · 2003
Earlier work this paper cites.
Dbn based multi-stream models for audio-visual speech recognition
John N Gowdy, Amarnag Subramanya, Chris Bartels, and Jeff Bilmes · 2004
Earlier work this paper cites.
Voice activity detection using visual information
Peng Liu and Zuoying Wang · 2004
Earlier work this paper cites.
Audio-visual automatic speech recognition: An overview
Gerasimos Potamianos, C. Neti, Juergen Luettin, and Iain Matthews · 2004
Earlier work this paper cites.
An analysis of visual speech information applied to voice activity detection
D. Sodoyer, B. Rivet, L. Girin, J.-L. Schwartz, and C. Jutten · 2006
Earlier work this paper cites.
Two novel visual voice activity detectors based on appearance models and retinal filtering
Andrew Aubrey, Bertrand Rivet, Yulia Hicks, Laurent Girin, Jonathon Chambers, and Christian Jutten · 2007
Earlier work this paper cites.
Articulatory feature-based methods for acoustic and audio-visual speech recognition: Summary from the 2006 jhu summer workshop
Karen Livescu, Ozgur Cetin, Mark Hasegawa-Johnson, Simon King, Chris Bartels, Nash Borges, Arthur Kantor, Partha Lal, Lisa Yung, Ari Bezman, et al · 2007
Earlier work this paper cites.
Voice activity detection. fundamentals and speech recognition system robustness
Javier Ramirez, Juan Manuel Górriz, and José Carlos Segura · 2007
Earlier work this paper cites.
Adaptive multimodal fusion by uncertainty compensation with application to audiovisual speech recognition
George Papandreou, Athanassios Katsamanis, Vassilis Pitsikalis, and Petros Maragos · 2009
Earlier work this paper cites.
Visual lip activity detection and speaker detection using mouth region intensities
Spyridon Siatras, Nikos Nikolaidis, Michail Krinidis, and Ioannis Pitas · 2009
Earlier work this paper cites.
A study of lip movements during spontaneous dialog and its application to voice activity detection
David Sodoyer, Bertrand Rivet, Laurent Girin, Christophe Savariaux, jean-luc Schwartz, and Christian Jutten · 2009
Earlier work this paper cites.
Visual voice activity detection with optical flow
Andrew J. Aubrey, Yulia A. Hicks, and Jonathon A. Chambers · 2010
Earlier work this paper cites.
A visual voice activity detection method with adaboosting
Qingju Liu, Wenwu Wang, and Philip Jackson · 2011
Earlier work this paper cites.
Learning temporal signatures for lip reading
Eng-Jon Ong and Richard Bowden · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Interference reduction in reverberant speech separation with visual voice activity detection
Qingju Liu, Andrew J. Aubrey, and Wenwu Wang · 2014
Earlier work this paper cites.
A review of recent advances in visual speech decoding
Ziheng Zhou, Guoying Zhao, Xiaopeng Hong, and Matti Pietikäinen · 2014
Earlier work this paper cites.
Deep learning of mouth shapes for sign language
Oscar Koller, Hermann Ney, and Richard Bowden · 2015
Earlier work this paper cites.
Lexicon-free conversational speech recognition with neural networks
Andrew L. Maas, Ziang Xie, Dan Jurafsky, and Andrew Y. Ng · 2015
Earlier work this paper cites.
Lipnet: Sentence-level lipreading
Yannis M. Assael, Brendan Shillingford, Shimon Whiteson, and Nando de Freitas · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals · 2016
Earlier work this paper cites.
Lip reading in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Visual voice activity detection in the wild
Foteini Patrona, Alexandros Iosifidis, Anastasios Tefas, Nikolaos Nikolaidis, and Ioannis Pitas · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Earlier work this paper cites.
Lip reading in profile
Joon Son Chung and Andrew Zisserman · 2017
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W. Taylor · 2017
Cited alongside, same era.
Exploring roi size in deep learning based lipreading
Alexandros Koumparoulis, Gerasimos Potamianos, Youssef Mroueh, and Steven J. Rennie · 2017
Cited alongside, same era.
Montreal forced aligner: Trainable text-speech alignment using kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, M. Wagner, and Morgan Sonderegger · 2017
Cited alongside, same era.
Combining residual networks with lstms for lipreading
Themos Stafylakis and Georgios Tzimiropoulos · 2017
Cited alongside, same era.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Hearing lips: Improving lip reading by distilling speech recognizers
Ya Zhao, Rui Xu, Xinchao Wang, Peng Hou, Haihong Tang, and Mingli Song · 2019
Later among the works it cites.
ASR is all you need: Cross-modal distillation for lip reading
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2020
Later among the works it cites.
Now you’re speaking my language: Visual language identification
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2020
Later among the works it cites.
Self-supervised learning of audio-visual objects from video
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman · 2020
Later among the works it cites.
Active speakers in context
Juan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi, Joon-Young Lee, Pablo Arbeláez, and Bernard Ghanem · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
LRS3-TED: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Cited alongside, same era.
VoxCeleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Cited alongside, same era.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Cited alongside, same era.
Learn to pay attention
Saumya Jetley, Nicholas A Lord, Namhoon Lee, and Philip HS Torr · 2018
Cited alongside, same era.
A review of image-based automatic facial landmark identification techniques
Benjamin Johnston and Philip Chazal · 2018
Cited alongside, same era.
An analysis of incorporating an external language model into a sequence-to-sequence model
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N. Sainath, Zhifeng Chen, and Rohit Prabhavalkar · 2018
Cited alongside, same era.
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
Lip reading sentences using deep learning with only visual cues
Souheil Fenghour, Daqing Chen, Kun Guo, and Perry Xiao · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Later among the works it cites.
Adversarial attacks against lipnet: End-to-end sentence level lipreading
Mahir Jethanandani and Derek Tang · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Later among the works it cites.
Seeing wake words: Audio-visual keyword spotting
Liliane Momeni, Triantafyllos Afouras, Themos Stafylakis, Samuel Albanie, and Andrew Zisserman · 2020
Later among the works it cites.
Ava-activespeaker: An audio-visual dataset for active speaker detection
Joseph Roth, Sourish Chaudhuri, Ondrej Klejch, Radhika Marvin, Andrew C. Gallagher, Liat Kaver, Sharadh Ramaswamy, Arkadiusz Stopczynski, Cordelia Schmid, Zhonghua Xi, and Caroline Pantofaru · 2020
Later among the works it cites.
Visual transformers: Token-based image representation and processing for computer vision
Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Masayoshi Tomizuka, Kurt Keutzer, and Peter Vajda · 2020
Later among the works it cites.
Audio-visual recognition of overlapped speech for the lrs2 dataset
Jianwei Yu, Shi-Xiong Zhang, Jian Wu, Shahram Ghorbani, Bo Wu, Shiyin Kang, Shansong Liu, Xunying Liu, Helen Meng, and Dong Yu · 2020
Later among the works it cites.
Can We Read Speech Beyond the Lips? Rethinking RoI Selection for Deep Visual Speech Recognition
Yuanhang Zhang, Shuang Yang, Jingyun Xiao, Shiguang Shan, and Xilin Chen · 2020
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Fatalread-fooling visual speech recognition models
Anup Kumar Gupta, Puneet Gupta, and Esa Rahtu · 2021
Closest in time.
Learning visual voice activity detection with an automatically annotated dataset
Sylvain Guy, Stéphane Lathuilière, Pablo Mesejo, and Radu Horaud · 2021
Closest in time.
Maas: Multi-modal assignation for active speaker detection
Juan León-Alcázar, Fabian Caba Heilbron, Ali Thabet, and Bernard Ghanem · 2021
Closest in time.
Detecting adversarial attacks on audiovisual speech recognition
Pingchuan Ma, Stavros Petridis, and Maja Pantic · 2021
Closest in time.
End-to-end audio-visual speech recognition with conformers
Pingchuan Ma, Stavros Petridis, and Maja Pantic · 2021
Closest in time.
Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection
Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li · 2021
Closest in time.
Efficient DETR: improving end-to-end object detector with dense prior
Zhuyu Yao, Jiangbo Ai, Boxun Li, and Chi Zhang · 2021
Closest in time.