Fetching the paper…
Reading the bibliography…
Machine lipreading is a special type of automatic speech recognition (ASR) which transcribes human speech by visually interpreting the movement of related face regions including lips, face, and tongue.
J. Luettin et al. , “Visual speech recognition using active shape models and hidden markov models,” in ICASSP , 1996
1996
Earlier work this paper cites.
G. Potamianos et al. , “An image transform approach for hmm based automatic lipreading,” in ICIP , 1998
1998
Earlier work this paper cites.
P. Neti et al. , “Audio-visual speech recognition,” in Technical report. Center for Language and Speech Processing, Johns Hopkins University, Baltimore , 2000
2000
Earlier work this paper cites.
T. F. Cootes et al. , “Active appearance models,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 23, no. 6, pp. 681–685, 2001
2001
Earlier work this paper cites.
K. Papineni et al. , “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting on association for computational linguistics , 2002
2002
Earlier work this paper cites.
A. Graves et al. , “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in ICML , 2006
2006
Earlier work this paper cites.
G. Zhao et al. , “Local spatiotemporal descriptors for visual recognition of spoken phrases,” in Proceedings of the International Workshop on Human-centered Multimedia , 2007
2007
Earlier work this paper cites.
S. Hilder et al. , “Comparison of human and machine-based lip-reading,” in INTERSPEECH , 2009
2009
Earlier work this paper cites.
C. Chandrasekaran et al. , “The natural statistics of audiovisual speech,” PLoS computational biology , vol. 5, no. 7, p. e1000436, 2009
2009
Earlier work this paper cites.
D. E. King, “Dlib-ml: A machine learning toolkit,” Journal of Machine Learning Research , vol. 10, pp. 1755–1758, 2009
2009
Earlier work this paper cites.
A. A. Shaikh et al. , “Lip reading using optical flow and support vector machines,” in International Congress on Image and Signal Processing , 2010
2010
Earlier work this paper cites.
G. E. Dahl et al. , “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,” IEEE Transactions on audio, speech, and language processing , vol. 20, no. 1, pp. 30–42, 2012
2012
Earlier work this paper cites.
G. Hinton et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
X. Wang, L. Pollock, and K. Vijay-Shanker, “Automatic segmentation of method code into meaningful blocks: Design and evaluation,” Journal of Software: Evolution and Process , 2013
2013
Earlier work this paper cites.
A. Graves et al. , “Speech recognition with deep recurrent neural networks,” in ICASSP , 2013
2013
Earlier work this paper cites.
K. Noda et al. , “Lipreading using convolutional neural network,” in INTERSPEECH , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
D. Wu and Q. Ruan, “Lip reading based on cascade feature extraction and hmm,” in International Conference on Signal Processing (ICSP) , 2014
2014
Earlier work this paper cites.
S. S. Morade and S. Patnaik, “Lip reading using dwt and lsda,” in IEEE International Advance Computing Conference (IACC) , 2014
2014
Earlier work this paper cites.
X. Liu and Y. m. Cheung, “Learning multi-boosted hmms for lip-password based speaker verification,” in IEEE Transactions on Information Forensics and Security , 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization.” in ICLR , 2014
2014
Earlier work this paper cites.
J. Wu, E. E. Abdel-Fatah, and M. R. Mahfouz, “Fully automatic initialization of two-dimensional–three-dimensional medical image registration using hybrid classifier,” Journal of Medical Imaging , vol. 2, no. 2, p. 024007, 2015
2015
Cited alongside, same era.
Y. Miao et al. , “Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding.” in In IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Later among the works it cites.
M. Abadi et al. , “Tensorflow: A system for large-scale machine learning,” in USENIX Symposium on Operating Systems Design and Implementation (OSDI) , 2016
2016
Later among the works it cites.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in ACCV , 2016
2016
Later among the works it cites.
K. He et al. , “Deep residual learning for image recognition,” in CVPR , 2016
2016
Later among the works it cites.
S. Petridis et al. , “End-to-end visual speech recognition with lstms,” in ICASSP , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
K. Palecek, “Comparison of depth-based features for lipreading,” in International Conference on Telecommunications and Signal Processing (TSP) , 2015
2015
Cited alongside, same era.
D. Tran et al. , “Learning spatiotemporal features with 3d convolutional networks,” in ICCV , 2015
2015
Cited alongside, same era.
R. K. Srivastava et al. , “Training very deep networks,” in NIPS , 2015
2015
Cited alongside, same era.
Y. M. Assael et al. , “Lipnet: Sentence-level lipreading,” CoRR , vol. abs/1611.01599, 2016
2016
Cited alongside, same era.
M. Wand et al. , “Lipreading with long short-term memory,” in ICASSP , 2016
2016
Cited alongside, same era.
S. Miao, Z. J. Wang, and R. Liao, “A cnn regression approach for real-time 2d/3d registration,” IEEE transactions on medical imaging , vol. 35, no. 5, pp. 1352–1363, 2016
2016
Cited alongside, same era.
J. Wu and M. R. Mahfouz, “Robust x-ray image segmentation by spectral clustering and active shape model,” Journal of Medical Imaging , vol. 3, no. 3, p. 034005, 2016
2016
Cited alongside, same era.
2017
Later among the works it cites.
Y. Geng, G. Zhang, W. Li, Y. Gu, R.-Z. Liang, G. Liang, J. Wang, Y. Wu, N. Patil, and J.-Y. Wang, “A novel image tag completion method based on convolutional neural transformation,” in International Conference on Artificial Neural Networks . Springer, 2017, pp. 539–546
2017
Later among the works it cites.
G. Zhang, G. Liang, W. Li, J. Fang, J. Wang, Y. Geng, and J.-Y. Wang, “Learning convolutional ranking-score function by query preference regularization,” in International Conference on Intelligent Data Engineering and Automated Learning . Springer, 2017, pp. 1–8
2017
Later among the works it cites.
D. Tome, C. Russell, and L. Agapito, “Lifting from the deep: Convolutional 3d pose estimation from a single image,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , July 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
J. Liang, C. Chen, Y. Yi, X. Xu, and M. Ding, “Bilateral two-dimensional neighborhood preserving discriminant embedding for face recognition,” IEEE Access , vol. 5, 2017
2017
Later among the works it cites.
——, “Automatically generating natural language descriptions for object-related statement sequences,” in Proceedings of the 24th International Conference on Software Analysis, Evolution and Reengineering (SANER) , Feb 2017, pp. 205–216
2017
Later among the works it cites.
2017
Later among the works it cites.
J. Gehring et al. , “Convolutional sequence to sequence learning,” aarXiv:1705.03122 , 2017
2017
Later among the works it cites.
A. Vaswani et al. , “Attention is all you need,” arXiv:1706.03762 , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
S. Han, J. Kang, H. Mao, Y. Hu, X. Li, Y. Li, D. Xie, H. Luo, S. Yao, Y. Wang, H. Yang, and W. B. J. Dally, “Ese: Efficient speech recognition engine with sparse lstm on fpga,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , ser. FPGA ’17. New York, NY, USA: ACM, 2017, pp. 75–84. [Online]. Available: http://doi.acm.org/10.1145/3020078.3021745
2017
Later among the works it cites.
Y. Mao, J. Oak, A. Pompili, D. Beer, T. Han, and P. Hu, “Draps: Dynamic and resource-aware placement scheme for docker containers in a heterogeneous cluster,” in 2017 IEEE 36th International Performance Computing and Communications Conference (IPCCC) , Dec 2017, pp. 1–8
2017
Later among the works it cites.
T. Hori et al. , “Joint ctc/attention decoding for end-to-end speech recognition,” in Annual Meeting of the Association for Computational Linguistics , 2017
2017
Later among the works it cites.
T. Stafylakis and G. Tzimiropoulos, “Combining residual networks with lstms for lipreading,” in Interspeech , 2017
2017
Later among the works it cites.