Fetching the paper…
Reading the bibliography…
This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition.
R. Lienhart, “Reliable transition detection in videos: A survey and practitioner’s guide,”
2001
Earlier work this paper cites.
M. Cooke, J. Barker, S. Cunningham, and X. Shao, “An audio-visual corpus for speech perception and automatic speech recognition,”
2006
Earlier work this paper cites.
J. Yuan and M. Liberman, “Speaker identification on the scotus corpus,”
2008
Earlier work this paper cites.
Y. M. Assael, B. Shillingford, S. Whiteson, and N. de Freitas, “Lipnet: Sentence-level lipreading,”
2016
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in
2016
Earlier work this paper cites.
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “Ssd: Single shot multibox detector,” in
2016
Cited alongside, same era.
J. S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in
2016
Cited alongside, same era.
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” in
2017
Cited alongside, same era.
T. Stafylakis and G. Tzimiropoulos, “Combining residual networks with LSTMs for lipreading,” in
2017
Cited alongside, same era.
A. Czyzewski, B. Kostek, P. Bratoszewski, J. Kotus, and M. Szykulski, “An audio-visual corpus for multimodal automatic speech recognition,”
2017
Later among the works it cites.
2018
Closest in time.
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman, “Deep audio-visual speech recognition,” in
2018
Closest in time.
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…