Fetching the paper…
Reading the bibliography…
Multimodal learning allows us to leverage information from multiple sources (visual, acoustic and text), similar to our experience of the real world.
Probing the need for visual context in multimodal machine translation
Caglayan, O., Madhyastha, P., Specia, L., and Barrault, L · 1903
Earlier work this paper cites.
Hearing lips and seeing voices
McGurk, H. and MacDonald, J · 1976
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
Toutanova, K., Klein, D., Manning, C. D., and Singer, Y · 2003
Earlier work this paper cites.
Understanding the exploding gradient problem
Pascanu, R., Mikolov, T., and Bengio, Y · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Deep multimodal semantic embeddings for speech and images
Harwath, D. and Glass, J · 2015
Earlier work this paper cites.
Deep multimodal learning for audio-visual speech recognition
Mroueh, Y., Marcheret, E., and Goel, V · 2015
Earlier work this paper cites.
Lipnet: End-to-end sentence-level lipreading
Assael, Y. M., Shillingford, B., Whiteson, S., and De Freitas, N · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, W., Jaitly, N., Le, Q., and Vinyals, O · 2016
Cited alongside, same era.
Unsupervised learning of spoken language with visual context
Harwath, D., Torralba, A., and Glass, J · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Open-domain audio-visual speech recognition: A deep learning approach
Miao, Y. and Metze, F · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2016
Cited alongside, same era.
Look, listen, and decode: Multimodal speech recognition with images
Attention strategies for multi-source sequence-to-sequence learning
Libovickỳ, J. and Helcl, J · 2017
Later among the works it cites.
Findings of the third shared task on multimodal machine translation
Barrault, L., Bougares, F., Specia, L., Lala, C., Elliott, D., and Frank, S · 2018
Later among the works it cites.
Adversarial evaluation of multimodal machine translation
Elliott, D · 2018
Later among the works it cites.
Lstm language model adaptation with images and titles for multimedia automatic speech recognition
Moriya, Y. and Jones, G. J · 2018
Later among the works it cites.
End-to-end multimodal speech recognition
Palaskar, S., Sanabria, R., and Metze, F · 2018
Later among the works it cites.
How2: A large-scale dataset for multimodal language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sun, F., Harwath, D., and Glass, J · 2016
Cited alongside, same era.
Nmtpy: A flexible toolkit for advanced neural machine translation systems
Caglayan, O., García-Martínez, M., Bardet, A., Aransa, W., Bougares, F., and Barrault, L · 2017
Cited alongside, same era.
Lip reading sentences in the wild
Chung, J. S., Senior, A., Vinyals, O., and Zisserman, A · 2017
Cited alongside, same era.
Visual features for context-aware speech recognition
Gupta, A., Miao, Y., Neves, L., and Metze, F · 2017
Cited alongside, same era.
Sanabria, R., Caglayan, O., Palaskar, S., Elliott, D., Barrault, L., Specia, L., and Metze, F · 2018
Later among the works it cites.
Multimodal grounding for sequence-to-sequence speech recognition
Caglayan, O., Sanabria, R., Palaskar, S., Barraul, L., and Metze, F · 2019
Closest in time.
Multimodal speaker adaptation of acoustic model and language model for asr using speaker face embedding
Moriya, Y. and Jones, G. J · 2019
Closest in time.
Modality attention for end-to-end audio-visual speech recognition
Zhou, P., Yang, W., Chen, W., Wang, Y., and Jia, J · 2019
Closest in time.