Fetching the paper…
Reading the bibliography…
Text-to-image retrieval is an essential task in cross-modal information retrieval, i.e., retrieving relevant images from a large and unlabelled dataset given textual queries.
Introduction to information retrieval
Christopher D Manning, Hinrich Schütze, and Prabhakar Raghavan · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Elasticsearch in action
Radu Gheorghe, Matthew Lee Hinman, and Roy Russo · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2016
Earlier work this paper cites.
Vse++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2017
Earlier work this paper cites.
Instance-aware image and sentence matching with selective multimodal lstm
Yan Huang, Wei Wang, and Liang Wang · 2017
Earlier work this paper cites.
Dual attention networks for multimodal reasoning and matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Finding beans in burgers: Deep semantic-visual embedding with localization
Martin Engilberge, Louis Chevallier, Patrick Pérez, and Matthieu Cord · 2018
Cited alongside, same era.
Learning semantic concepts and order for image and sentence matching
Yan Huang, Qi Wu, Chunfeng Song, and Liang Wang · 2018
Cited alongside, same era.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He · 2018
Cited alongside, same era.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston · 2018
Cited alongside, same era.
Context-aware sentence/passage term importance estimation for first stage retrieval
Zhuyun Dai and Jamie Callan · 2019
Cited alongside, same era.
Language-agnostic visual-semantic embeddings
Jonatas Wehrmann, Douglas M Souza, Mauricio A Lopes, and Rodrigo C Barros · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
Array programming with NumPy
Charles R. Harris, K. Jarrod Millman, St’efan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fern’andez del R’ıo, Mark Wiebe, Pearu Peterson, Pierre G’erard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant · 2020
Later among the works it cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, Daxin Jiang, and Ming Zhou · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Convert: Efficient and accurate conversational representations from transformers
Matthew Henderson, Iñigo Casanueva, Nikola Mrkšić, Pei-Hao Su, Tsung-Hsien Wen, and Ivan Vulić · 2019
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Cited alongside, same era.
Position focused attention network for image-text matching
Yaxiong Wang, Hao Yang, Xueming Qian, Lin Ma, Jing Lu, Biao Li, and Xin Fan · 2019
Cited alongside, same era.
Camp: Cross-modal adaptive message passing for text-image retrieval
Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao · 2019
Cited alongside, same era.
Later among the works it cites.
12-in-1: Multi-task vision and language representation learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee · 2020
Later among the works it cites.
Expansion via prediction of importance with contextualization
Sean MacAvaney, Franco Maria Nardini, Raffaele Perego, Nicola Tonellotto, Nazli Goharian, and Ophir Frieder · 2020
Later among the works it cites.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti · 2020
Later among the works it cites.
Consensus-aware visual-semantic embedding for image-text matching
Haoran Wang, Ying Zhang, Zhong Ji, Yanwei Pang, and Lin Ma · 2020
Later among the works it cites.
Sparta: Efficient open-domain question answering via sparse transformer matching retrieval
Tiancheng Zhao, Xiaopeng Lu, and Kyusong Lee · 2020
Later among the works it cites.