Fetching the paper…
Reading the bibliography…
Speech is one of the most effective ways of communication among humans.
J Jeffers and M Barley, · 1971
Earlier work this paper cites.
“Hearing lips and seeing voices,”
Harry McGurk and John MacDonald, · 1976
Earlier work this paper cites.
“Coarticulation in recent speech production models,”
Raymond D Kent and Fred D Minifie, · 1977
Earlier work this paper cites.
“Temporal patterns of coarticulation: Lip rounding,”
Fredericka Bell-Berti and Katherine S Harris, · 1982
Earlier work this paper cites.
“The dynamics of audiovisual behavior in speech,”
Eric Vatikiotis-Bateson, Kevin G Munhall, Makoto Hirayama, Y Victor Lee, and Demetri Terzopoulos, · 1996
Earlier work this paper cites.
“Characterizing audiovisual information during speech.,”
Eric Vatikiotis-Bateson, Kevin G Munhall, Y Kasahara, Frederique Garcia, and Hani Yehia, · 1996
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
Features for audio-visual speech recognition
Iain Matthews, · 1998
Earlier work this paper cites.
“The moving face during speech communication,”
Eric Vatikiotis-Bateson, · 1998
Cited alongside, same era.
“Audio visual speech recognition,”
Chalapathy Neti, Gerasimos Potamianos, Juergen Luettin, Iain Matthews, Herve Glotin, Dimitra Vergyri, June Sison, and Azad Mashari, · 2000
Cited alongside, same era.
“Large-vocabulary audio-visual speech recognition by machines and humans.,”
Gerasimos Potamianos, Chalapathy Neti, Giridharan Iyengar, and Eric Helmuth, · 2001
Cited alongside, same era.
“Asynchronous stream modeling for large vocabulary audio-visual speech recognition,”
Juergen Luettin, Gerasimos Potamianos, and Chalapathy Neti, · 2001
Cited alongside, same era.
“Large vocabulary audio-visual speech recognition using the janus speech recognition toolkit,”
Jan Kratt, Florian Metze, Rainer Stiefelhagen, and Alex Waibel, · 2004
Cited alongside, same era.
“Distinctive image features from scale-invariant keypoints,”
“Phoneme-to-viseme mapping for visual speech recognition.,”
Luca Cappelletta and Naomi Harte, · 2012
Later among the works it cites.
“The effect of speaking rate on audio and visual speech,”
Sarah Taylor, Barry-John Theobald, and Iain Matthews, · 2014
Later among the works it cites.
“Deep multimodal learning for audio-visual speech recognition,”
Youssef Mroueh, Etienne Marcheret, and Vaibhava Goel, · 2015
Later among the works it cites.
“Audio-visual speech recognition using deep learning,”
Kuniaki Noda, Yuki Yamaguchi, Kazuhiro Nakadai, Hiroshi G Okuno, and Tetsuya Ogata, · 2015
Later among the works it cites.
“Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Later among the works it cites.
“Intraface,”
Fernando De la Torre, Wen-Sheng Chu, Xuehan Xiong, Francisco Vicente, Xiaoyu Ding, and Jeffrey Cohn, · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David G Lowe, · 2004
Cited alongside, same era.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Cited alongside, same era.
“Multimodal deep learning,”
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng, · 2011
Cited alongside, same era.
Later among the works it cites.
“Scaling recurrent neural network language models,”
Will Williams, Niranjani Prasad, David Mrva, Tom Ash, and Tony Robinson, · 2015
Later among the works it cites.
“Temporal multimodal learning in audiovisual speech recognition,”
Di Hu, Xuelong Li, et al., · 2016
Closest in time.