Fetching the paper…
Reading the bibliography…
Several end-to-end deep learning approaches have been recently presented which simultaneously extract visual features from the input images and perform visual speech classification.
A. Varga and H. Steeneken, “Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems,”
1993
Earlier work this paper cites.
S. Dupont and J. Luettin, “Audio-visual speech modeling for continuous speech recognition,”
2000
Earlier work this paper cites.
S. Young, G. Evermann, M. Gales, T. Hain, D. Kershaw, X. Liu, G. Moore, J. Odell, D. Ollason, D. Povey
2002
Earlier work this paper cites.
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A. W. Senior, “Recent advances in the automatic recognition of audiovisual speech,”
2003
Earlier work this paper cites.
P. S. Aleksic and A. K. Katsaggelos, “Audio-visual biometrics,”
2006
Earlier work this paper cites.
G. Hinton and R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,”
2006
Earlier work this paper cites.
B. Schuller, R. Mueller, F. Eyben, J. Gast, B. Hoernler, M. Woellmer, G. Rigoll, A. Hoethker, and H. Konosu, “Being bored? Recognising natural interest by extensive audiovisual integration for real-life application,”
2009
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks.” in
2010
Earlier work this paper cites.
S. Petridis and M. Pantic, “Audiovisual discrimination between speech and laughter: Why and when visual information might help,”
2011
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in
2011
Earlier work this paper cites.
F. Eyben, S. Petridis, B. Schuller, G. Tzimiropoulos, S. Zafeiriou, and M. Pantic, “Audiovisual classification of vocal outbursts in human conversation using long-short-term memory networks,” in
2011
Cited alongside, same era.
Y. Lan, B. J. Theobald, and R. Harvey, “View independent computer lip-reading,” in
2012
Cited alongside, same era.
G. Hinton, “A practical guide to training restricted boltzmann machines,” in
2012
Cited alongside, same era.
H. Gunes and B. Schuller, “Categorical and dimensional affect analysis in continuous input: Current trends and future directions,”
2013
Cited alongside, same era.
V. Kazemi and J. Sullivan, “One millisecond face alignment with an ensemble of regression trees,” in
2014
Cited alongside, same era.
D. Hu, X. Li
2016
Later among the works it cites.
Y. Takashima, R. Aihara, T. Takiguchi, Y. Ariki, N. Mitani, K. Omori, and K. Nakazono, “Audio-visual speech recognition using bimodal-trained bottleneck features for a person with severe hearing loss,”
2016
Later among the works it cites.
M. Wand, J. Koutn, and J. Schmidhuber, “Lipreading with long short-term memory,” in
2016
Later among the works it cites.
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas, “Lipnet: Sentence-level lipreading,”
2016
Later among the works it cites.
S. Petridis and M. Pantic, “Prediction-based audiovisual fusion for classification of non-linguistic vocalisations,”
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
H. Ninomiya, N. Kitaoka, S. Tamura, Y. Iribe, and K. Takeda, “Integration of deep bottleneck features for audio-visual speech recognition,” in
2015
Cited alongside, same era.
Y. Mroueh, E. Marcheret, and V. Goel, “Deep multimodal learning for audio-visual speech recognition,” in
2015
Cited alongside, same era.
I. Anina, Z. Zhou, G. Zhao, and M. Pietikäinen, “Ouluvs2: A multi-view audiovisual database for non-rigid mouth motion analysis,” in
2015
Cited alongside, same era.
N. Harte and E. Gillen, “Tcd-timit: An audio-visual corpus of continuous speech,”
2015
Cited alongside, same era.
Z. Zeng, M. Pantic, G. Roisman, and T. Huang, “A survey of affect recognition methods: Audio, visual and spontaneous expressions,”
Cited in the paper.
T. Saitoh, Z. Zhou, G. Zhao, and M. Pietikäinen, “Concatenated frame image based cnn for visual speech recognition,” in
2016
Later among the works it cites.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in
2016
Later among the works it cites.
S. Petridis, Z. Li, and M. Pantic, “End-to-end visual speech recognition with lstms,” in
2017
Closest in time.
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,”
2017
Closest in time.