Fetching the paper…
Reading the bibliography…
Advances in the field of vision-language contrastive learning have made it possible for many downstream applications to be carried out efficiently and accurately by simply taking the dot product between image and text representations.
Y. Tian, D. Krishnan, and P. Isola · 1906
Earlier work this paper cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
G. Li, N. Duan, Y. Fang, D. Jiang, and M. Zhou · 1908
Earlier work this paper cites.
UNITER: learning universal image-text representations
Y. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 1909
Earlier work this paper cites.
Tiger: Text-to-image grounding for image caption evaluation
M. Jiang, Q. Huang, L. Zhang, X. Wang, P. Zhang, Z. Gan, J. Diesner, and J. Gao · 1909
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. B. Girshick · 1911
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. E. Hinton · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
X. Chen, H. Fan, R. B. Girshick, and K. He · 2003
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
X. Chen, H. Fan, R. B. Girshick, and K. He · 2003
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, Y. Choi, and J. Gao · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
T. Wang and P. Isola · 2005
Earlier work this paper cites.
Big self-supervised models are strong semi-supervised learners
T. Chen, S. Kornblith, K. Swersky, M. Norouzi, and G. E. Hinton · 2006
Earlier work this paper cites.
Large-scale adversarial training for vision-and-language representation learning
Z. Gan, Y. Chen, L. Li, C. Zhu, Y. Cheng, and J. Liu · 2006
Earlier work this paper cites.
Bootstrap your own latent: A new approach to self-supervised learning
J. Grill, F. Strub, F. Altché, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. Á. Pires, Z. D. Guo, M. G. Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. Valko · 2006
Earlier work this paper cites.
Consensus-aware visual-semantic embedding for image-text matching
H. Wang, Y. Zhang, Z. Ji, Y. Pang, and L. Ma · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Collecting image annotations using amazon’s mechanical turk
C. Rashtchian, P. Young, M. Hodosh, and J. Hockenmaier · 2010
Earlier work this paper cites.
Contrastive learning with hard negative samples
J. Robinson, C. Chuang, S. Sra, and S. Jegelka · 2010
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba · 2010
Earlier work this paper cites.
Contrastive learning of medical visual representations from paired images and text
Y. Zhang, H. Jiang, Y. Miura, C. D. Manning, and C. P. Langlotz · 2010
Earlier work this paper cites.
A survey on contrastive self-supervised learning
A. Jaiswal, A. R. Babu, M. Z. Zadeh, D. Banerjee, and F. Makedon · 2011
Earlier work this paper cites.
UNIMO: towards unified-modal understanding and generation via cross-modal contrastive learning
W. Li, C. Gao, G. Niu, X. Xiao, H. Liu, J. Liu, H. Wu, and H. Wang · 2012
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier · 2013
Cited alongside, same era.
Some improvements on deep convolutional neural network based image classification
A. G. Howard · 2013
Cited alongside, same era.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Cited alongside, same era.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Improving image captioning evaluation by considering inter references variance
Y. Yi, H. Deng, and J. Hu · 2020
Later among the works it cites.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Later among the works it cites.
Clipscore: A reference-free evaluation metric for image captioning
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig · 2021
Later among the works it cites.
Transparent human evaluation for image captioning
J. Kasai, K. Sakaguchi, L. Dunagan, J. Morrison, R. L. Bras, Y. Choi, and N. A. Smith · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2014
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Cited alongside, same era.
SPICE: semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Revisiting batch normalization for practical domain adaptation
Y. Li, N. Wang, J. Shi, J. Liu, and X. Hou · 2016
Cited alongside, same era.
Improved deep metric learning with multi-class n-pair loss objective
K. Sohn · 2016
Cited alongside, same era.
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision, 2021
W. Kim, B. Son, and I. Kim · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
J. Li, R. R. Selvaraju, A. D. Gotmare, S. R. Joty, C. Xiong, and S. C. H. Hoi · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation, 2021
X. L. Li and P. Liang · 2021
Later among the works it cites.
Visualsparta: Sparse transformer fragment-level matching for large-scale text-to-image search
X. Lu, T. Zhao, and K. Lee · 2021
Later among the works it cites.
Styleclip: Text-driven manipulation of stylegan imagery
O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischinski · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Later among the works it cites.
Better aggregation in test-time augmentation
D. Shanmugam, D. Blalock, G. Balakrishnan, and J. Guttag · 2021
Later among the works it cites.
C. You, R. Zhao, L. H. Staib, and J. S. Duncan · 2021
Later among the works it cites.
Multimodal contrastive training for visual representation learning
X. Yuan, Z. Lin, J. Kuen, J. Zhang, Y. Wang, M. Maire, A. Kale, and B. Faieta · 2021
Later among the works it cites.
Lit: Zero-shot transfer with locked-image text tuning
X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer · 2021
Later among the works it cites.
Calip: Zero-shot enhancement of clip with parameter-free attention, 2022
Z. Guo, R. Zhang, L. Qiu, X. Ma, X. Miao, X. He, and B. Cui · 2022
Later among the works it cites.
Caltech 101, Apr 2022
F.-F. Li, M. Andreeto, M. Ranzato, and P. Perona · 2022
Later among the works it cites.
Crafting better contrastive views for siamese representation learning
X. Peng, K. Wang, Z. Zhu, and Y. You · 2022
Later among the works it cites.
Vt-clip: Enhancing vision-language models with visual-guided texts, 2022
L. Qiu, R. Zhang, Z. Guo, Z. Zeng, Y. Li, and G. Zhang · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models, 2022
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev · 2022
Later among the works it cites.
Test-time prompt tuning for zero-shot generalization in vision-language models, 2022
M. Shu, W. Nie, D.-A. Huang, Z. Yu, T. Goldstein, A. Anandkumar, and C. Xiao · 2022
Later among the works it cites.
Vision-language pre-training with triple contrastive learning, 2022
J. Yang, J. Duan, S. Tran, Y. Xu, S. Chanda, L. Chen, B. Zeng, T. Chilimbi, and J. Huang · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models, 2022
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Later among the works it cites.
Sus-x: Training-free name-only transfer of vision-language models, 2023
V. Udandarao, A. Gupta, and S. Albanie · 2023
Closest in time.