Fetching the paper…
Reading the bibliography…
Image-Text Matching (ITM) task, a fundamental vision-language (VL) task, suffers from the inherent ambiguity arising from multiplicity and imperfect annotations.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, and Kevin Murphy · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
VSE++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2018
Earlier work this paper cites.
Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models
Jiuxiang Gu, Jianfei Cai, Shafiq R Joty, Li Niu, and Gang Wang · 2018
Earlier work this paper cites.
Learning semantic concepts and order for image and sentence matching
Yan Huang, Qi Wu, Chunfeng Song, and Liang Wang · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
Nsml: Meet the mlaas platform with a real-world case study
Hanjoo Kim, Minkyu Kim, Dongjoo Seo, Jinwoong Kim, Heungseok Park, Soeun Park, Hyunwoo Jo, KyungHyun Kim, Youngil Yang, Youngkwan Kim, et al · 2018
Earlier work this paper cites.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He · 2018
Earlier work this paper cites.
Exploring the limits of weakly supervised pretraining
Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens Van Der Maaten · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou · 2019
Cited alongside, same era.
Visual semantic reasoning for image-text matching
Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu · 2019
Cited alongside, same era.
Modeling uncertainty with hedged instance embedding
Seong Joon Oh, Kevin Murphy, Jiyan Pan, Joseph Roth, Florian Schroff, and Andrew Gallagher · 2019
Cited alongside, same era.
Probabilistic face embeddings
Yichun Shi and Anil K Jain · 2019
Cited alongside, same era.
Polysemous visual-semantic embedding for cross-modal retrieval
Yale Song and Mohammad Soleymani · 2019
Cited alongside, same era.
RedCaps: Web-curated image-text data created by the people, for the people
Karan Desai, Gaurav Kaul, Zubin Aysola, and Justin Johnson · 2021
Later among the works it cites.
Similarity reasoning and filtration for image-text matching
Haiwen Diao, Ying Zhang, Lin Ma, and Huchuan Lu · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, and Jung-Woo Ha · 2021
Later among the works it cites.
Learning with noisy correspondence for cross-modal matching
Zhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding, Xinyan Xiao, hua wu, and Xi Peng · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language-agnostic visual-semantic embeddings
Jonatas Wehrmann, Douglas M Souza, Mauricio A Lopes, and Rodrigo C Barros · 2019
Cited alongside, same era.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo · 2019
Cited alongside, same era.
Data uncertainty learning in face recognition
Jie Chang, Zhonghao Lan, Changmao Cheng, and Yichen Wei · 2020
Cited alongside, same era.
Dividemix: Learning with noisy labels as semi-supervised learning
Junnan Li, Richard Socher, and Steven C.H. Hoi · 2020
Cited alongside, same era.
A metric learning reality check
Kevin Musgrave, Serge Belongie, and Ser-Nam Lim · 2020
Cited alongside, same era.
Openclip, July 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for ms-coco
Zarana Parekh, Jason Baldridge, Daniel Cer, Austin Waters, and Yinfei Yang · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
Ross Wightman, Hugo Touvron, and Hervé Jégou · 2021
Later among the works it cites.
Is an image worth five sentences? a new look into semantics for image-text matching
Ali Furkan Biten, Andres Mafla, Lluís Gómez, and Dimosthenis Karatzas · 2022
Later among the works it cites.
Eccv caption: Correcting false negatives by collecting machine-and-human-verified image-caption associations for ms-coco
Sanghyuk Chun, Wonjae Kim, Song Park, Minsuk Chang Chang, and Seong Joon Oh · 2022
Later among the works it cites.
Probabilistic compositional embeddings for multimodal image retrieval
Andrei Neculai, Yanbei Chen, and Zeynep Akata · 2022
Later among the works it cites.
Deep evidential learning with noisy correspondence for cross-modal retrieval
Yang Qin, Dezhong Peng, Xi Peng, Xu Wang, and Peng Hu · 2022
Later among the works it cites.
Point to rectangle matching for image text retrieval
Zheng Wang, Zhenwei Gao, Xing Xu, Yadan Luo, Yang Yang, and Heng Tao Shen · 2022
Later among the works it cites.
Map: Multimodal uncertainty-aware vision-language pre-training model
Yatai Ji, Junjie Wang, Yuan Gong, Lin Zhang, Yanru Zhu, Hongfa Wang, Jiaxing Zhang, Tetsuya Sakai, and Yujiu Yang · 2023
Closest in time.
Improving cross-modal retrieval with set of diverse embeddings
Dongwon Kim, Namyup Kim, and Suha Kwak · 2023
Closest in time.
Probabilistic contrastive learning recovers the correct aleatoric uncertainty of ambiguous inputs
Michael Kirchhof, Enkelejda Kasneci, and Seong Joon Oh · 2023
Closest in time.
Bayesian metric learning for uncertainty quantification in image retrieval
Frederik Warburg, Marco Miani, Silas Brack, and Soren Hauberg · 2023
Closest in time.