Fetching the paper…
Reading the bibliography…
In the past few years, cross-modal image-text retrieval (ITR) has experienced increased interest in the research community due to its excellent research value and broad real-world application.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Vse++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2017
Earlier work this paper cites.
Pairwise relationship guided deep hashing for cross-modal retrieval
Erkun Yang, Cheng Deng, Wei Liu, Xianglong Liu, Dacheng Tao, and Xinbo Gao · 2017
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He · 2018
Earlier work this paper cites.
Self-supervised adversarial hashing networks for cross-modal retrieval
Chao Li, Cheng Deng, Ning Li, Wei Liu, Xinbo Gao, and Dacheng Tao · 2018
Earlier work this paper cites.
Learning two-branch neural networks for image-text matching tasks
Liwei Wang, Yin Li, Jing Huang, and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Attention-aware deep adversarial hashing for cross-modal retrieval
Xi Zhang, Hanjiang Lai, and Jiashi Feng · 2018
Earlier work this paper cites.
Saliency-guided attention network for image-sentence matching
Zhong Ji, Haoran Wang, Jungong Han, and Yanwei Pang · 2019
Earlier work this paper cites.
Visual semantic reasoning for image-text matching
Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Focus your attention: A bidirectional focal attention network for image-text matching
Chunxiao Liu, Zhendong Mao, An-An Liu, Tianzhu Zhang, Bin Wang, and Yongdong Zhang · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Adversarial representation learning for text-to-image matching
Nikolaos Sarafianos, Xiang Xu, and Ioannis A Kakadiaris · 2019
Earlier work this paper cites.
Polysemous visual-semantic embedding for cross-modal retrieval
Yale Song and Mohammad Soleymani · 2019
Earlier work this paper cites.
Camp: Cross-modal adaptive message passing for text-image retrieval
Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao · 2019
Earlier work this paper cites.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma · 2019
Earlier work this paper cites.
Learning fragment self-attention embeddings for image-text matching
Yiling Wu, Shuhui Wang, Guoli Song, and Qingming Huang · 2019
Earlier work this paper cites.
Imram: Iterative matching with recurrent attention memory for cross-modal image-text retrieval
Hui Chen, Guiguang Ding, Xudong Liu, Zijia Lin, Ji Liu, and Jungong Han · 2020
Earlier work this paper cites.
Review of recent deep learning based methods for image-text retrieval
Jianan Chen, Lu Zhang, Cong Bai, and Kidiyo Kpalma · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Fashionbert: Text and image matching with adaptive loss for cross-modal retrieval
Dehong Gao, Linbo Jin, Ben Chen, Minghui Qiu, Peng Li, Yi Wei, Yi Hu, and Hao Wang · 2020
Cited alongside, same era.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu · 2020
Cited alongside, same era.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang · 2020
Cited alongside, same era.
Step-wise hierarchical alignment network for image-text matching
Zhong Ji, Kexin Chen, and Haoran Wang · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Later among the works it cites.
Kd-vlp: Improving end-to-end vision-and-language pretraining with object knowledge distillation
Yongfei Liu, Chenfei Wu, Shao-yen Tseng, Vasudev Lal, Xuming He, and Nan Duan · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Cited alongside, same era.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti · 2020
Cited alongside, same era.
Consensus-aware visual-semantic embedding for image-text matching
Haoran Wang, Ying Zhang, Zhong Ji, Yanwei Pang, and Lin Ma · 2020
Cited alongside, same era.
Multi-modality cross attention network for image and sentence matching
Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, and Feng Wu · 2020
Cited alongside, same era.
Context-aware attention network for image-text retrieval
Qi Zhang, Zhen Lei, Zhaoxiang Zhang, and Stan Z Li · 2020
Cited alongside, same era.
Dual-path convolutional image-text embeddings with instance loss
Zhedong Zheng, Liang Zheng, Michael Garrett, Yi Yang, Mingliang Xu, and Yi-Dong Shen · 2020
Cited alongside, same era.
Stacmr: scene-text aware cross-modal retrieval
Andrés Mafla, Rafael S Rezende, Lluis Gomez, Diane Larlus, and Dimosthenis Karatzas · 2021
Later among the works it cites.
Thinking fast and slow: Efficient text-to-visual retrieval with transformers
Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2021
Later among the works it cites.
Dynamic modality interaction modeling for image-text retrieval
Leigang Qu, Meng Liu, Jianlong Wu, Zan Gao, and Liqiang Nie · 2021
Later among the works it cites.
Learning relation alignment for calibrated cross-modal retrieval
Shuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men, An Yang, Jingren Zhou, Xu Sun, and Hongxia Yang · 2021
Later among the works it cites.
E2e-vlp: End-to-end vision-language pre-training enhanced by visual learning
Haiyang Xu, Ming Yan, Chenliang Li, Bin Bi, Songfang Huang, Wenming Xiao, and Fei Huang · 2021
Later among the works it cites.
Probing inter-modality: Visual parsing with self-attention for vision-and-language pre-training
Hongwei Xue, Yupan Huang, Bei Liu, Houwen Peng, Jianlong Fu, Houqiang Li, and Jiebo Luo · 2021
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2021
Later among the works it cites.
Deep graph-neighbor coherence preserving network for unsupervised cross-modal hashing
Jun Yu, Hao Zhou, Yibing Zhan, and Dacheng Tao · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Later among the works it cites.
Kaleido-bert: Vision-language pre-training on fashion domain
Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Ben Chen, Haoming Zhou, Minghui Qiu, and Ling Shao · 2021
Later among the works it cites.
Playing lottery tickets with vision and language
Zhe Gan, Yen-Chun Chen, Linjie Li, Tianlong Chen, Yu Cheng, Shuohang Wang, and Jingjing Liu · 2022
Closest in time.
Multimodal research in vision and language: A review of current and emerging trends
Shagun Uppal, Sarthak Bhagat, Devamanyu Hazarika, Navonil Majumder, Soujanya Poria, Roger Zimmermann, and Amir Zadeh · 2022
Closest in time.
Filip: Fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2022
Closest in time.