Fetching the paper…
Reading the bibliography…
Speech is understood better by using visual context; for this reason, there have been many attempts to use images to adapt automatic speech recognition (ASR) systems.
“Xsede: Accelerating scientific discovery,”
J. Towns, T. Cockerill, M. Dahan, I. Foster, K. Gaither, A. Grimshaw, V. Hazlewood, S. Lathrop, D. Lifka, G. D. Peterson, R. Roskies, J. Scott, and N. Wilkins-Diehr, · 2002
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Petr Qian, Yanmand Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Earlier work this paper cites.
“VQA: Visual Question Answering,”
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh, · 2015
Earlier work this paper cites.
“Deep multimodal semantic embeddings for speech and images,”
David Harwath and James Glass, · 2015
Earlier work this paper cites.
“Framing image description as a ranking task: Data, models and evaluation metrics (extended abstract),”
Micah Hodosh, Peter Young, and Julia Hockenmaier, · 2015
Earlier work this paper cites.
“A shared task on multimodal machine translation and crosslingual image description (WMT),”
Lucia Specia, Stella Frank, Khalil Sima’an, and Desmond Elliott, · 2016
Earlier work this paper cites.
“Open-domain audio-visual speech recognition: A deep learning approach.,”
Yajie Miao and Florian Metze, · 2016
Earlier work this paper cites.
“Look, listen, and decode: Multimodal speech recognition with images,”
Felix Sun, David Harwath, and James Glass, · 2016
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Cited alongside, same era.
“Visual Dialog,”
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M.F. Moura, Devi Parikh, and Dhruv Batra, · 2017
Cited alongside, same era.
“Visual features for context-aware speech recognition,”
Abhinav Gupta, Yajie Miao, Leonardo Neves, and Florian Metze, · 2017
Cited alongside, same era.
“Places: A 10 million image database for scene recognition,”
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba, · 2017
“End-to-end multimodal speech recognition,”
Shruti Palaskar, Ramon Sanabria, and Florian Metze, · 2018
Later among the works it cites.
“LSTM language model adaptation with images and titles for multimedia automatic speech recognition,”
Yasufumi Moriya and Gareth J. F. Jones, · 2018
Later among the works it cites.
“Adversarial evaluation of multimodal machine translation,”
Desmond Elliott, · 2018
Later among the works it cites.
“How2: a large-scale dataset for multimodal language understanding,”
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze, · 2018
Later among the works it cites.
“Multimodal Grounding for Sequence-to-Sequence Speech Recognition,”
Ozan Caglayan, Ramon Sanabria, Shruti Palaskar, Loïc Barrault, and Florian Metze, · 2019
Later among the works it cites.
“Analyzing utility of visual context multimodal speech recognition under noisy conditions,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Nmtpy: A flexible toolkit for advanced neural machine translation systems,”
Ozan Caglayan, Mercedes García-Martínez, Adrien Bardet, Walid Aransa, Fethi Bougares, and Loïc Barrault, · 2017
Cited alongside, same era.
Tejas Srinivasan, Ramon Sanabria, and Florian Metze, · 2019
Later among the works it cites.
“Probing the need for visual context multimodal machine translation,”
Ozan Caglayan, Pranava Madhyastha, Lucia Specia, and Loïc Barrault, · 2019
Later among the works it cites.