Fetching the paper…
Reading the bibliography…
In this study, we propose a deep neural network for reconstructing intelligible speech from silent lip movement videos.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Perceptual evaluation of speech quality (pesq), an objective method for end-to-end speech quality assessment of narrowband telephone networks and speech codecs
Rix, A., Beerends, J., Hollier, M., and Hekstra, A. (2001) · 2001
Earlier work this paper cites.
A spectro-temporal modulation index (stmi) for assessment of speech intelligibility
Elhilali, M., Chi, T., and Shamma, S. A. (2003) · 2003
Earlier work this paper cites.
Multiresolution spectrotemporal analysis of complex sounds
Chi, T., Ru, P., and Shamma, S. A. (2005) · 2005
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Cooke, M., Barker, J., Cunningham, S., and Shao, X. (2006) · 2006
Earlier work this paper cites.
Multimodal deep learning
Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Maas, A. L., Hannun, A. Y., and Ng, A. Y. (2013) · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J. (2014) · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X. (2015) · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S. (2015) · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Cited alongside, same era.
Lipnet: Sentence-level lipreading
Assael, Y. M., Shillingford, B., Whiteson, S., and de Freitas, N. (2016) · 2016
Later among the works it cites.
Lip reading in the wild
Chung, J. S. and Zisserman, A. (2016) · 2016
Later among the works it cites.
Visually indicated sounds
Owens, A., Isola, P., McDermott, J., Torralba, A., Adelson, E. H., and Freeman, W. T. (2016) · 2016
Later among the works it cites.
Deep complementary bottleneck features for visual speech recognition
Petridis, S. and Pantic, M. (2016) · 2016
Later among the works it cites.
Lipreading with long short-term memory
Wand, M., Koutník, J., and Schmidhuber, J. (2016) · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ioffe, S. and Szegedy, C. (2015) · 2015
Cited alongside, same era.
Reconstructing intelligible audio speech from visual speech features
Milner, B. and Le Cornu, T. (2015) · 2015
Cited alongside, same era.
Audio-visual speech recognition using deep learning
Noda, K., Yamaguchi, Y., Nakadai, K., Okuno, H. G., and Ogata, T. (2015) · 2015
Cited alongside, same era.
Ephrat, A. and Peleg, S. (2017) · 2017
Closest in time.