Fetching the paper…
Reading the bibliography…
Speechreading is a notoriously difficult task for humans to perform.
Acoustic theory of speech production: with calculations based on X-ray studies of Russian articulations
Gunnar Fant, · 1971
Earlier work this paper cites.
“Line spectrum representation of linear predictor coefficients of speech signals,”
Fumitada Itakura, · 1975
Earlier work this paper cites.
Automatic lipreading to enhance speech recognition (speech reading)
Eric David Petajan, · 1984
Earlier work this paper cites.
“Analysis, synthesis, and perception of voice quality variations among female and male talkers,”
Dennis H Klatt and Laura C Klatt, · 1990
Earlier work this paper cites.
“Opencv,”
G. Bradski, · 2000
Earlier work this paper cites.
“Extraction of visual features for lipreading,”
Iain Matthews, Timothy F Cootes, J Andrew Bangham, Stephen Cox, and Richard Harvey, · 2002
Earlier work this paper cites.
“A neural network model of the articulatory-acoustic forward mapping trained on recordings of articulatory parameters,”
Christopher T Kello and David C Plaut, · 2004
Earlier work this paper cites.
“An audio-visual corpus for speech perception and automatic speech recognition,”
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao, · 2006
Earlier work this paper cites.
“Decoding visemes: Improving machine lip-reading,”
Helen L Bear and Richard Harvey, · 2013
Cited alongside, same era.
“Rectifier nonlinearities improve neural network acoustic models,”
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng, · 2013
Cited alongside, same era.
“A review of recent advances in visual speech decoding,”
Ziheng Zhou, Guoying Zhao, Xiaopeng Hong, and Matti Pietikäinen, · 2014
Cited alongside, same era.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2014
Cited alongside, same era.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Keras,” https://github.com/fchollet/keras
François Chollet, · 2015
Later among the works it cites.
“Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2015
Later among the works it cites.
“Lipreading with long short-term memory,”
Michael Wand, Jan Koutn, et al., · 2016
Later among the works it cites.
“Lipnet: End-to-end sentence-level lipreading,”
Yannis M Assael, Brendan Shillingford, Shimon Whiteson, and Nando de Freitas, · 2016
Later among the works it cites.
“Lip reading sentences in the wild,”
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman, · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Cited alongside, same era.
“Reconstructing intelligible audio speech from visual speech features,”
Thomas Le Cornu and Ben Milner, · 2015
Cited alongside, same era.
“Tensorflow,” Software available from http://tensorflow.org/
Cited in the paper.
“Speech signal processing toolkit,” Available from http://sp-tk.sourceforge.net/readme.php
Cited in the paper.
“Statistical conversion of silent articulation into audible speech using full-covariance hmm,”
Thomas Hueber and Gérard Bailly, · 2016
Later among the works it cites.
“Visually indicated sounds,”
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman, · 2016
Later among the works it cites.