Fetching the paper…
Reading the bibliography…
Despite the achievements of large-scale multimodal pre-training approaches, cross-modal retrieval, e.g., image-text retrieval, remains a challenging task.
What does BERT look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 1906
Earlier work this paper cites.
Do attention heads in BERT track syntactic dependencies?
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R. Bowman. 2019 · 1911
Earlier work this paper cites.
Learning the semantics of words and pictures
Kobus Barnard and David A. Forsyth. 2001 · 2001
Earlier work this paper cites.
Matching words and pictures
Kobus Barnard, Pinar Duygulu, David A. Forsyth, Nando de Freitas, David M. Blei, and Michael I. Jordan. 2003 · 2003
Earlier work this paper cites.
Interbert: Vision-and-language interaction for multi-modal pretraining
Junyang Lin, An Yang, Yichang Zhang, Jie Liu, Jingren Zhou, and Hongxia Yang. 2020 · 2003
Earlier work this paper cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020 · 2006
Earlier work this paper cites.
Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders
Nicola Messina, Giuseppe Amato, Andrea Esuli, Fabrizio Falchi, Claudio Gennaro, and Stéphane Marchand-Maillet. 2020 · 2008
Earlier work this paper cites.
Multimodal pretraining unmasked: Unifying the vision and language berts
Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and Desmond Elliott. 2020 · 2011
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li. 2015 · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Cited alongside, same era.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
CAMP: cross-modal adaptive message passing for text-image retrieval
Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao. 2019 · 2019
Later among the works it cites.
UNITER: universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Large-scale adversarial training for vision-and-language representation learning
Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Graph structured network for image-text matching
Chunxiao Liu, Zhendong Mao, Tianzhu Zhang, Hongtao Xie, Bin Wang, and Yongdong Zhang. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
VSE++: improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018 · 2018
Cited alongside, same era.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018 · 2018
Cited alongside, same era.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
LXMERT: learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
VL-BERT: pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong, and Quoc Le. 2020 · 2020
Later among the works it cites.
Visual agreement regularized training for multi-modal machine translation
Pengcheng Yang, Boxing Chen, Pei Zhang, and Xu Sun. 2020 · 2020
Later among the works it cites.
Learning to represent image and text with denotation graph
Bowen Zhang, Hexiang Hu, Vihan Jain, Eugene Ie, and Fei Sha. 2020a · 2020
Later among the works it cites.
DCA: diversified co-attention towards informative live video commenting
Zhihan Zhang, Zhiyi Yin, Shuhuai Ren, Xinhang Li, and Shicheng Li. 2020b · 2020
Later among the works it cites.
M6: A chinese multimodal pretrainer
Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, Jie Zhang, Jianwei Zhang, Xu Zou, Zhikang Li, Xiaodong Deng, Jie Liu, Jinbao Xue, Huiling Zhou, Jianxin Ma, Jin Yu, Yong Li, Wei Lin, Jingren Zhou, Jie Tang, and Hongxia Yang. 2021 · 2021
Closest in time.