Fetching the paper…
Reading the bibliography…
Several end-to-end deep learning approaches have been recently presented which extract either audio or visual features from the input images or audio signals and perform speech recognition.
A. Varga and H. Steeneken, · 1993
Earlier work this paper cites.
“Audio-visual speech modeling for continuous speech recognition,”
S. Dupont and J. Luettin, · 2000
Earlier work this paper cites.
“Moving-talker, speaker-independent feature study, and baseline results using the CUAVE multimodal speech corpus,”
E. Patterson, S. Gurbuz, Z. Tufekci, and J. Gowdy, · 2002
Earlier work this paper cites.
“Recent advances in the automatic recognition of audiovisual speech,”
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A. W. Senior, · 2003
Earlier work this paper cites.
“An audio-visual corpus for speech perception and automatic speech recognition,”
M. Cooke, J. Barker, S. Cunningham, and X. Shao, · 2006
Earlier work this paper cites.
“Multimodal deep learning,”
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y Ng, · 2011
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Integration of deep bottleneck features for audio-visual speech recognition,”
H. Ninomiya, N. Kitaoka, S. Tamura, Y. Iribe, and K. Takeda, · 2015
Earlier work this paper cites.
“Deep multimodal learning for audio-visual speech recognition,”
Y. Mroueh, E. Marcheret, and V. Goel, · 2015
Earlier work this paper cites.
“Ouluvs2: A multi-view audiovisual database for non-rigid mouth motion analysis,”
I. Anina, Z. Zhou, G. Zhao, and M. Pietikäinen, · 2015
Cited alongside, same era.
“Prediction-based audiovisual fusion for classification of non-linguistic vocalisations,”
S. Petridis and M. Pantic, · 2016
Cited alongside, same era.
“Temporal multimodal learning in audiovisual speech recognition,”
D. Hu, X. Li, and X. Lu, · 2016
Cited alongside, same era.
“Audio-visual speech recognition using bimodal-trained bottleneck features for a person with severe hearing loss,”
Y. Takashima, R. Aihara, T. Takiguchi, Y. Ariki, N. Mitani, K. Omori, and K. Nakazono, · 2016
Cited alongside, same era.
“Deep complementary bottleneck features for visual speech recognition,”
S. Petridis and M. Pantic, · 2016
Cited alongside, same era.
“Lipreading with long short-term memory,”
“Lip reading in the wild,”
J. S. Chung and A. Zisserman, · 2016
Later among the works it cites.
“Deep residual learning for image recognition,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Later among the works it cites.
“Identity mappings in deep residual networks,”
K. He, X. Zhang, S. Ren, and J. Sun, · 2016
Later among the works it cites.
“End-to-end visual speech recognition with LSTMs,”
S. Petridis, Z. Li, and M. Pantic, · 2017
Later among the works it cites.
“Combining residual networks with LSTMs for lipreading,”
T. Stafylakis and G. Tzimiropoulos, · 2017
Later among the works it cites.
“Lip reading sentences in the wild,”
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Wand, J. Koutnik, and J. Schmidhuber, · 2016
Cited alongside, same era.
“Lipnet: Sentence-level lipreading,”
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas, · 2016
Cited alongside, same era.
“Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network,”
G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. Nicolaou, B. Schuller, and S. Zafeiriou, · 2016
Cited alongside, same era.
S. Petridis, Y. Wang, Z. Li, and M. Pantic, · 2017
Later among the works it cites.
“End-to-end multimodal emotion recognition using deep neural networks,”
P. Tzirakis, G. Trigeorgis, M. A. Nicolaou, B. Schuller, and S. Zafeiriou, · 2017
Later among the works it cites.