Fetching the paper…
Reading the bibliography…
This article aims to provide the information retrieval community with some reflections on recent advances in retrieval learning by analyzing the reproducibility of image-text retrieval models.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K. Paliwal. 1997 · 1997
Earlier work this paper cites.
Cumulated gain-based evaluation of IR techniques
Kalervo Järvelin and Jaana Kekäläinen. 2002 · 2002
Earlier work this paper cites.
Object retrieval with large vocabularies and fast spatial matching. In CVPR
James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman. 2007 · 2007
Earlier work this paper cites.
The Probabilistic Relevance Framework: BM25 and Beyond
Stephen E. Robertson and Hugo Zaragoza. 2009 · 2009
Earlier work this paper cites.
Im2Text: Describing Images Using 1 Million Captioned Photographs. In NeurIPS
Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. 2011 · 2011
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In ECCV
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Neural Machine Translation by Jointly Learning to Align and Translate. In ICLR
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In CVPR
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Andrej Karpathy and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Feature Pyramid Networks for Object Detection. In CVPR . 936–944
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. 2017 · 2017
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In NeurIPS
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Deep Metric Learning with Angular Loss. In ICCV
Jian Wang, Feng Zhou, Shilei Wen, Xiao Liu, and Yuanqing Lin. 2017 · 2017
Earlier work this paper cites.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In CVPR
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Earlier work this paper cites.
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives. In BMVC
Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018 · 2018
Earlier work this paper cites.
Stacked cross attention for image-text matching. In ECCV
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018 · 2018
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In ACL
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
Representation Learning with Contrastive Predictive Coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
Graph Attention Networks. In ICLR
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Comprehensive distance-preserving autoencoders for cross-modal retrieval. In ACM Multimedia
Yibing Zhan, Jun Yu, Zhou Yu, Rong Zhang, Dacheng Tao, and Qi Tian. 2018 · 2018
Cited alongside, same era.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT . Association for Computational Linguistics
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
VL-BERT: Pre-training of Generic Visual-Linguistic Representations. In ICLR
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
On the Gap between Adoption and Understanding in NLP. In Findings of ACL
Federico Bianchi and Dirk Hovy. 2021 · 2021
Later among the works it cites.
Similarity Reasoning and Filtration for Image-Text Matching. In AAAI
Haiwen Diao, Ying Zhang, Lin Ma, and Huchuan Lu. 2021 · 2021
Later among the works it cites.
Progressive Multi-Granularity Training for Non-Autoregressive Translation. In Fingdings of ACL
Liang Ding, Longyue Wang, Xuebo Liu, Derek F Wong, Dacheng Tao, and Zhaopeng Tu. 2021 · 2021
Later among the works it cites.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision. In ICML
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019a · 2019
Cited alongside, same era.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In NeurIPS
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval. In ICCV
Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao. 2019 · 2019
Cited alongside, same era.
Learning Fragment Self-Attention Embeddings for Image-Text Matching. In ACM Multimedia
Yiling Wu, Shuhui Wang, Guoli Song, and Qingming Huang. 2019 · 2019
Cited alongside, same era.
UNITER: UNiversal Image-TExt Representation Learning. In ECCV
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Cited alongside, same era.
Self-attention with cross-lingual position representation. In ACL
Liang Ding, Longyue Wang, and Dacheng Tao. 2020 · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. In ICML
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021 · 2021
Later among the works it cites.
UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning. In ACL/IJCNLP
Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. 2021 · 2021
Later among the works it cites.
VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words. In ACL
Xiaopeng Lu, Tiancheng Zhao, and Kyusong Lee. 2021 · 2021
Later among the works it cites.
Do Transformer Modifications Transfer Across Implementations and Applications?. In EMNLP
Sharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel. 2021 · 2021
Later among the works it cites.
Student Can Also be a Good Teacher: Extracting Knowledge from Vision-and-Language Model for Cross-Modal Retrieval. In CIKM
Jun Rao, Tao Qian, Shuhan Qi, Yulin Wu, Qing Liao, and Xuan Wang. 2021 · 2021
Later among the works it cites.
Contrastive Learning with Hard Negative Samples. In ICLR
Joshua David Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka. 2021 · 2021
Later among the works it cites.
Slua: A super lightweight unsupervised word alignment model via cross-lingual contrastive learning
Di Wu, Liang Ding, Shuo Yang, and Dacheng Tao. 2021 · 2021
Later among the works it cites.
Vitae: Vision transformer advanced by exploring intrinsic inductive bias
Yufei Xu, Qiming Zhang, Jing Zhang, and Dacheng Tao. 2021 · 2021
Later among the works it cites.
Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal Hashing. In AAAI
Jun Yu, Hao Zhou, Yibing Zhan, and Dacheng Tao. 2021 · 2021
Later among the works it cites.
VinVL: Revisiting Visual Representations in Vision-Language Models. In CVPR
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
Bridging Cross-Lingual Gaps During Leveraging the Multilingual Sequence-to-Sequence Pretraining for Text Generation. In arXiv preprint
Changtong Zan, Liang Ding, Li Shen, Yu Cao, Weifeng Liu, and Dacheng Tao. 2022 · 2022
Closest in time.
ViTAEv2: Vision Transformer Advanced by Exploring Inductive Bias for Image Recognition and Beyond
Qiming Zhang, Yufei Xu, Jing Zhang, and Dacheng Tao. 2022 · 2022
Closest in time.