Fetching the paper…
Reading the bibliography…
In this study, we try to address the problem of leveraging visual signals to improve Automatic Speech Recognition (ASR), also known as visual context-aware ASR (VC-ASR).
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Lipnet: End-to-end sentence-level lipreading,”
Yannis M Assael, Brendan Shillingford, Shimon Whiteson, and Nando De Freitas, · 2016
Earlier work this paper cites.
“Open-domain audio-visual speech recognition: A deep learning approach.,”
Yajie Miao and Florian Metze, · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Rethinking the inception architecture for computer vision,”
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna, · 2016
Earlier work this paper cites.
“Visual features for context-aware speech recognition,”
Abhinav Gupta, Yajie Miao, Leonardo Neves, and Florian Metze, · 2017
Earlier work this paper cites.
“Deliberation networks: Sequence generation beyond one-pass decoding,”
Yingce Xia, Fei Tian, Lijun Wu, Jianxin Lin, Tao Qin, Nenghai Yu, and Tie-Yan Liu, · 2017
Earlier work this paper cites.
“End-to-end multimodal speech recognition,”
Shruti Palaskar, Ramon Sanabria, and Florian Metze, · 2018
Earlier work this paper cites.
“Lstm language model adaptation with images and titles for multimedia automatic speech recognition,”
Yasufumi Moriya and Gareth JF Jones, · 2018
Cited alongside, same era.
“Multimodal abstractive summarization of open-domain videos,”
Jindrich Libovickỳ, Shruti Palaskar, Spandana Gella, and Florian Metze, · 2018
Cited alongside, same era.
“Deep audio-visual speech recognition,”
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman, · 2018
Cited alongside, same era.
“Adversarial evaluation of multimodal machine translation,”
Desmond Elliott, · 2018
Cited alongside, same era.
“Jointly discovering visual objects and spoken words from raw sensory input,”
David Harwath, Adria Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass, · 2018
Cited alongside, same era.
“Semantic speech retrieval with a visually grounded model of untranscribed speech,”
“Multimodal grounding for sequence-to-sequence speech recognition,”
Ozan Caglayan, Ramon Sanabria, Shruti Palaskar, Loic Barraul, and Florian Metze, · 2019
Later among the works it cites.
“Probing the need for visual context in multimodal machine translation,”
Ozan Caglayan, Pranava Swaroop Madhyastha, Lucia Specia, and Loïc Barrault, · 2019
Later among the works it cites.
“End-to-end audio visual scene-aware dialog using multimodal attention-based video features,”
Chiori Hori, Huda Alamri, Jue Wang, Gordon Wichern, Takaaki Hori, Anoop Cherian, Tim K Marks, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das, et al., · 2019
Later among the works it cites.
“Two-pass end-to-end speech recognition,”
Tara N Sainath, Ruoming Pang, David Rybach, Yanzhang He, Rohit Prabhavalkar, Wei Li, Mirkó Visontai, Qiao Liang, Trevor Strohman, Yonghui Wu, et al., · 2019
Later among the works it cites.
“Towards end-to-end speech-to-text translation with two-pass decoding,”
Tzu-Wei Sung, Jun-You Liu, Hung-yi Lee, and Lin-shan Lee, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Herman Kamper, Gregory Shakhnarovich, and Karen Livescu, · 2018
Cited alongside, same era.
“How2: A large-scale dataset for multimodal language understanding,”
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze, · 2018
Cited alongside, same era.
“Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?,”
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh, · 2018
Cited alongside, same era.
“Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
“Analyzing utility of visual context in multimodal speech recognition under noisy conditions,”
Tejas Srinivasan, Ramon Sanabria, and Florian Metze, · 2019
Cited alongside, same era.
“Looking enhances listening: Recovering missing speech using images,”
Tejas Srinivasan, Ramon Sanabria, and Florian Metze, · 2020
Closest in time.
“Investigating topics, audio representations and attention for multimodal scene-aware dialog,”
Shachi H Kumar, Eda Okur, Saurav Sahay, Jonathan Huang, and Lama Nachman, · 2020
Closest in time.
“End-to-end learning of visual representations from uncurated instructional videos,”
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman, · 2020
Closest in time.
“Hero: Hierarchical encoder for video+ language omni-representation pre-training,”
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu, · 2020
Closest in time.
“Deliberation model based two-pass end-to-end speech recognition,”
Ke Hu, Tara N Sainath, Ruoming Pang, and Rohit Prabhavalkar, · 2020
Closest in time.