Fetching the paper…
Reading the bibliography…
Previous work on multimodal machine translation (MMT) has focused on the way of incorporating vision features into translation but little attention is on the quality of vision models.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Multimodal attention for neural machine translation
Ozan Caglayan, Loïc Barrault, and Fethi Bougares. 2016 · 2016
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Earlier work this paper cites.
A shared task on multimodal machine translation and crosslingual image description
Lucia Specia, Stella Frank, Khalil Sima’an, and Desmond Elliott. 2016 · 2016
Earlier work this paper cites.
Incorporating global visual features into attention-based neural machine translation
Iacer Calixto and Qun Liu. 2017 · 2017
Earlier work this paper cites.
An empirical study on the effectiveness of images in multimodal neural machine translation
Jean-Benoit Delbrouck and Stéphane Dupont. 2017 · 2017
Earlier work this paper cites.
Findings of the second shared task on multimodal machine translation and multilingual image description
Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Imagination improves multimodal translation
Desmond Elliott and Ákos Kádár. 2017 · 2017
Earlier work this paper cites.
Attention strategies for multi-source sequence-to-sequence learning
Jindřich Libovický and Jindřich Helcl. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Adversarial evaluation of multimodal machine translation
Desmond Elliott. 2018 · 2018
Cited alongside, same era.
The MeMAD submission to the WMT18 multimodal translation task
Stig-Arne Grönroos, Benoit Huet, Mikko Kurimo, Jorma Laaksonen, Bernard Merialdo, Phu Pham, Mats Sjöberg, Umut Sulubacak, Jörg Tiedemann, Raphael Troncy, and Raúl Vázquez. 2018 · 2018
Cited alongside, same era.
Sheffield submissions for WMT18 multimodal translation shared task
Chiraag Lala, Pranava Swaroop Madhyastha, Carolina Scarton, and Lucia Specia. 2018 · 2018
Cited alongside, same era.
Probing the need for visual context in multimodal machine translation
Ozan Caglayan, Pranava Madhyastha, Lucia Specia, and Loïc Barrault. 2019 · 2019
Cited alongside, same era.
fairseq: A fast, extensible toolkit for sequence modeling
A novel graph-based multi-modal fusion encoder for neural machine translation
Yongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou, Zhengyuan Yang, Jie Zhou, and Jiebo Luo. 2020 · 2020
Later among the works it cites.
Neural machine translation with universal visual representation
Zhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita, Zuchao Li, and Hai Zhao. 2020 · 2020
Later among the works it cites.
Double attention-based multimodal neural machine translation with semantic image regions
Yuting Zhao, Mamoru Komachi, Tomoyuki Kajiwara, and Chenhui Chu. 2020 · 2020
Later among the works it cites.
Cross-lingual visual pre-training for multimodal machine translation
Ozan Caglayan, Menekse Kuyu, Mustafa Sercan Amac, Pranava Madhyastha, Erkut Erdem, Aykut Erdem, and Lucia Specia. 2021 · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Cited alongside, same era.
Does multi-encoder help? a case study on context-aware neural machine translation
Bei Li, Hui Liu, Ziyang Wang, Yufan Jiang, Tong Xiao, Jingbo Zhu, Tongran Liu, and Changliang Li. 2020 · 2020
Cited alongside, same era.
Dynamic context-guided capsule network for multimodal machine translation
Huan Lin, Fandong Meng, Jinsong Su, Yongjing Yin, Zhengyuan Yang, Yubin Ge, Jie Zhou, and Jiebo Luo. 2020 · 2020
Cited alongside, same era.
Multimodal transformer for multimodal machine translation
Shaowei Yao and Xiaojun Wan. 2020 · 2020
Cited alongside, same era.
Gumbel-attention for multi-modal machine translation
Pengbo Liu, Hailong Cao, and Tiejun Zhao. 2021a
Cited in the paper.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021b
Cited in the paper.
Yuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Efficient object-level visual context modeling for multimodal machine translation: Masking irrelevant objects helps grounding
Dexin Wang and Deyi Xiong. 2021 · 2021
Later among the works it cites.
Good for misconceived reasons: An empirical revisiting on the need for visual context in multimodal machine translation
Zhiyong Wu, Lingpeng Kong, Wei Bi, Xiang Li, and Ben Kao. 2021 · 2021
Later among the works it cites.