Fetching the paper…
Reading the bibliography…
Visual speech recognition (VSR) aims to recognize the content of speech based on lip movements, without relying on the audio stream.
Visual contribution to speech intelligibility in noise
Sumby, W. H. & Pollack, I · 1954
Earlier work this paper cites.
Hearing lips and seeing voices
McGurk, H. & MacDonald, J · 1976
Earlier work this paper cites.
Audio-visual speech modeling for continuous speech recognition
Dupont, S. & Luettin, J · 2000
Earlier work this paper cites.
Recent advances in the automatic recognition of audiovisual speech
Potamianos, G., Neti, C., Gravier, G., Garg, A. & Senior, A. W · 2003
Earlier work this paper cites.
Silent speech interfaces
Denby, B. et al · 2010
Earlier work this paper cites.
Analysis of the visual lombard effect and automatic recognition experiments
Heracleous, P., Ishi, C. T., Sato, M., Ishiguro, H. & Hagita, N · 2013
Earlier work this paper cites.
Resolution limits on visual speech recognition
Bear, H. L., Harvey, R., Theobald, B.-J. & Lan, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. & Ba, J · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. & Zisserman, A · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D. & Khudanpur, S · 2015
Earlier work this paper cites.
Effects of mouthing and interlocutor presence on movements of visible vs. non-visible articulators
Bicevskis, K. et al · 2016
Earlier work this paper cites.
Hyperarticulation in lombard speech: Global coordination of the jaw, lips and the tongue
Šimko, J., Beňuš, Š. & Vainio, M · 2016
Earlier work this paper cites.
Lipnet: End-to-end sentence-level lipreading
Assael, Y., Shillingford, B., Whiteson, S. & De Freitas, N · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S. & Sun, J · 2016
Earlier work this paper cites.
Lip reading in the wild
Chung, J. S. & Zisserman, A · 2016
Earlier work this paper cites.
Lip reading sentences in the wild
Chung, J. S., Senior, A., Vinyals, O. & Zisserman, A · 2017
Earlier work this paper cites.
Improving speaker-independent lipreading with domain-adversarial training
Wand, M. & Schmidhuber, J · 2017
Earlier work this paper cites.
End-to-end multi-view lipreading
Petridis, S., Wang, Y., Li, Z. & Pantic, M · 2017
Earlier work this paper cites.
How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)
Bulat, A. & Tzimiropoulos, G · 2017
Earlier work this paper cites.
Multitask learning with low-level auxiliary tasks for encoder-decoder based speech recognition
Toshniwal, S., Tang, H., Lu, L. & Livescu, K · 2017
Earlier work this paper cites.
Combining residual networks with LSTMs for lipreading
Stafylakis, T. & Tzimiropoulos, G · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M. & Grangier, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A. et al · 2017
Earlier work this paper cites.
Hybrid ctc/attention architecture for end-to-end speech recognition
Watanabe, S., Hori, T., Kim, S., Hershey, J. R. & Hayashi, T · 2017
Earlier work this paper cites.
Deep audio-visual speech recognition
Afouras, T., Chung, J. S., Senior, A., Vinyals, O. & Zisserman, A · 2018
Earlier work this paper cites.
Audio-visual speech recognition with a hybrid CTC/attention architecture
Petridis, S., Stafylakis, T., Ma, P., Tzimiropoulos, G. & Pantic, M · 2018
Earlier work this paper cites.
The conversation: Deep audio-visual speech enhancement
Afouras, T., Chung, J. S. & Zisserman, A · 2018
Earlier work this paper cites.
Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation
Ephrat, A. et al · 2018
Earlier work this paper cites.
The impact of reduced video quality on visual speech recognition
Dungan, L., Karaali, A. & Harte, N · 2018
Cited alongside, same era.
Visual-only recognition of normal, whispered and silent speech
Petridis, S., Shen, J., Cetin, D. & Pantic, M · 2018
Cited alongside, same era.
LRS3-TED: a large-scale dataset for visual speech recognition
Afouras, T., Chung, J. S. & Zisserman, A · 2018
Cited alongside, same era.
ESPnet: End-to-end speech processing toolkit
Watanabe, S. et al · 2018
Cited alongside, same era.
TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation
Hernandez, F., Nguyen, V., Ghannay, S., Tomashenko, N. A. & Estève, Y · 2018
Cited alongside, same era.
ShufflenetV2: practical guidelines for efficient CNN architecture design
Retinaface: Single-stage dense face localisation in the wild
Deng, J. et al · 2020
Later among the works it cites.
Learning speech representations from raw audio by joint audiovisual self-supervision
Shukla, A., Petridis, S. & Pantic, M · 2020
Later among the works it cites.
Common voice: A massively-multilingual speech corpus
Ardila, R. et al · 2020
Later among the works it cites.
MLS: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G. & Collobert, R · 2020
Later among the works it cites.
Audio-visual speech recognition is worth 32×32×8 voxels
Serdyuk, D., Braga, O. & Siohan, O · 2021
Later among the works it cites.
End-to-end audio-visual speech recognition with conformers
Ma, P., Petridis, S. & Pantic, M · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ma, N., Zhang, X., Zheng, H. & Sun, J · 2018
Cited alongside, same era.
Large-scale visual speech recognition
Shillingford, B. et al · 2019
Cited alongside, same era.
Understanding pictograph with facial features: End-to-end sentence-level lip reading of chinese
Zhang, X. et al · 2019
Cited alongside, same era.
A cascade sequence-to-sequence model for chinese mandarin lip reading
Zhao, Y., Xu, R. & Song, M · 2019
Cited alongside, same era.
Recurrent neural network transducer for audio-visual speech recognition
Makino, T. et al · 2019
Cited alongside, same era.
Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness
Geirhos, R. et al · 2019
Cited alongside, same era.
Investigating the lombard effect influence on end-to-end audio-visual speech recognition
Ma, P., Petridis, S. & Pantic, M · 2019
Cited alongside, same era.
Learning from the master: Distilling cross-modal advanced knowledge for lip reading
Ren, S., Du, Y., Lv, J., Han, G. & He, S · 2021
Later among the works it cites.
Fusing information streams in end-to-end audio-visual speech recognition
Yu, W., Zeiler, S. & Kolossa, D · 2021
Later among the works it cites.
Look who’s talking: Active speaker detection in the wild
Kim, Y. J. et al · 2021
Later among the works it cites.
Lips don’t lie: A generalisable and robust approach to face forgery detection
Haliassos, A., Vougioukas, K., Petridis, S. & Pantic, M · 2021
Later among the works it cites.
https://www.nytimes.com/2018/10/22/business/efforts-to-acknowledge-the-risks-of-new-ai-technology.html (2018)
Efforts to Acknowledge the Risks of New A.I. Technology · 2021
Later among the works it cites.
https://www.vice.com/en/article/bvzvdw/tech-companies-are-training-ai-to-read-your-lips (2021)
Tech Companies Are Training AI to Read Your Lips · 2021
Later among the works it cites.
https://liopa.ai
Liopa - the world’s only startup focused on automated lipreading via visual speech recognition · 2021
Later among the works it cites.
https://www.wired.com/story/facial-recognition-laws-are-literally-all-over-the-map/ (2019)
Facial Recognition Laws Are (Literally) All Over the Map · 2021
Later among the works it cites.
https://innotechtoday.com/13-cities-where-police-are-banned-from-using-facial-recognition-tech/ (2020)
13 Cities Where Police Are Banned From Using Facial Recognition Tech · 2021
Later among the works it cites.
https://about.fb.com/news/2021/11/update-on-use-of-face-recognition/ (2021)
An Update On Our Use of Face Recognition · 2021
Later among the works it cites.
https://edition.cnn.com/2021/05/18/tech/amazon-police-facial-recognition-ban/index.html (2021)
Amazon will block police indefinitely from using its facial-recognition software · 2021
Later among the works it cites.
https://www.washingtonpost.com/technology/2020/06/11/microsoft-facial-recognition (2020)
Microsoft won’t sell police its facial-recognition technology, following similar moves by Amazon and IBM · 2021
Later among the works it cites.
The Multilingual TEDx Corpus for Speech Recognition and Translation
Salesky, E. et al · 2021
Later among the works it cites.
VoxLingua107: a dataset for spoken language recognition
Valk, J. & Alumäe, T · 2021
Later among the works it cites.
Towards practical lipreading with distilled and efficient models
Ma, P., Martinez, B., Petridis, S. & Pantic, M · 2021
Later among the works it cites.
Improving rnn transducer based asr with auxiliary tasks
Liu, C. et al · 2021
Later among the works it cites.
Intermediate loss regularization for ctc-based speech recognition
Lee, J. & Watanabe, S · 2021
Later among the works it cites.
LiRA: Learning Visual Speech Representations from Audio Through Self-Supervision
Ma, P., Mira, R., Petridis, S., Schuller, B. W. & Pantic, M · 2021
Later among the works it cites.
Transformer-based video front-ends for audio-visual speech recognition for single and multi-person video
Serdyuk, D., Braga, O. & Siohan, O · 2022
Closest in time.
End-to-end video-to-speech synthesis using generative adversarial networks
Mira, R. et al · 2022
Closest in time.
mpc001/visual_speech_recognition_for_multiple_languages: Visual speech recognition for multiple languages
Ma, P., Petridis, S. & Pantic, M · 2022
Closest in time.