Fetching the paper…
Reading the bibliography…
Vision-language models (VLMs) are impactful in part because they can be applied to a variety of visual understanding tasks in a zero-shot fashion, without any fine-tuning.
Approche mixte pour l’extraction automatique de terminologie: statistiques lexicales et filtres linguistiques
Daille, B · 1994
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C · 2003
Earlier work this paper cites.
Monte carlo sampling methods
Shapiro, A · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Co-reranking by mutual reinforcement for image search
Yao, T., Mei, T., and Ngo, C.-W · 2010
Earlier work this paper cites.
Handling the impact of low frequency events on co-occurrence based measures of word similarity
Role, F. and Nadif, M · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Mutual information and diverse decoding improve neural machine translation, 2016
Li, J. and Jurafsky, D · 2016
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
Li, J., Galley, M., Brockett, C., Gao, J., and Dolan, B · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Estimating the information gap between textual and visual representations
Henning, C. A. and Ewerth, R · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Agrawal, H., Desai, K., Wang, Y., Chen, X., Jain, R., Johnson, M., Batra, D., Parikh, D., Lee, S., and Anderson, P · 2019
Earlier work this paper cites.
Approximating cnns with bag-of-local-features models works surprisingly well on imagenet
Brendel, W. and Bethge, M · 2019
Earlier work this paper cites.
Decoupling representation and classifier for long-tailed recognition
Kang, B., Xie, S., Rohrbach, M., Yan, Z., Gordo, A., Feng, J., and Kalantidis, Y · 2019
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Towards unique and informative captioning of images
Wang, Z., Feng, B., Narasimhan, K., and Russakovsky, O · 2020
Earlier work this paper cites.
How effective is bert without word ordering? implications for language understanding and data privacy
Hessel, J. and Schofield, A · 2021
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J., Holtzman, A., Forbes, M., Bras, R. L., and Choi, Y · 2021
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., and Hoi, S. C. H · 2021
Cited alongside, same era.
A survey on bias and fairness in machine learning
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A · 2021
Cited alongside, same era.
Thinking fast and slow: Efficient text-to-visual retrieval with transformers
Miech, A., Alayrac, J.-B., Laptev, I., Sivic, J., and Zisserman, A · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Vl-checklist: Evaluating pre-trained vision-language models with objects, attributes and relations
Zhao, T., Zhang, T., Zhu, M., Shen, H., Lee, K., Lu, X., and Yin, J · 2022
Later among the works it cites.
Improving image generation with better captions
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al · 2023
Closest in time.
Going beyond nouns with vision & language models using synthetic data
Cascante-Bonilla, P., Shehada, K., Smith, J. S., Doveh, S., Kim, D., Panda, R., Varol, G., Oliva, A., Ordonez, V., Feris, R., et al · 2023
Closest in time.
Dense and aligned captions (dac) promote compositional reasoning in vl models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Clip-lite: information efficient visual representation learning from textual annotations
Shrivastava, A., Selvaraju, R. R., Naik, N., and Ordonez, V · 2021
Cited alongside, same era.
Sinha, K., Jia, R., Hupkes, D., Pineau, J., Williams, A., and Kiela, D · 2021
Cited alongside, same era.
A fistful of words: Learning transferable visual models from bag-of-words supervision
Tejankar, A., Sanjabi, M., Wu, B., Xie, S., Khabsa, M., Pirsiavash, H., and Firooz, H · 2021
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
Yuan, W., Neubig, G., and Liu, P · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zhao, T., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Testing large language models on compositionality and inference with phrase-level adjective-noun entailment
Bertolini, L., Weeds, J., and Weir, D · 2022
Cited alongside, same era.
Doveh, S., Arbelle, A., Harary, S., Alfassy, A., Herzig, R., Kim, D., Giryes, R., Feris, R., Panda, R., Ullman, S., et al · 2023
Closest in time.
Gptscore: Evaluate as you desire
Fu, J., Ng, S.-K., Jiang, Z., and Liu, P · 2023
Closest in time.
Incorporating structured representations into pretrained vision & language models using scene graphs
Herzig, R., Mendelson, A., Karlinsky, L., Arbelle, A., Feris, R., Darrell, T., and Globerson, A · 2023
Closest in time.
Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Hsieh, C.-Y., Zhang, J., Ma, Z., Kembhavi, A., and Krishna, R · 2023
Closest in time.
Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Hu, Y., Liu, B., Kasai, J., Wang, Y., Ostendorf, M., Krishna, R., and Smith, N. A · 2023
Closest in time.
Structure-clip: Enhance multi-modal language representations with structure knowledge
Huang, Y., Tang, J., Chen, Z., Zhang, R., Zhang, X., Chen, W., Zhao, Z., Lv, T., Hu, Z., and Zhang, W · 2023
Closest in time.
Text encoders are performance bottlenecks in contrastive vision-language models
Kamath, A., Hessel, J., and Chang, K.-W · 2023
Closest in time.
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Closest in time.
Multimodality helps unimodality: Cross-modal few-shot learning with multimodal models
Lin, Z., Yu, S., Kuang, Z., Pathak, D., and Ramana, D · 2023
Closest in time.
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Closest in time.
Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation
Lu, Y., Yang, X., Li, X., Wang, X. E., and Wang, W. Y · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Singh, H., Zhang, P., Wang, Q., Wang, M., Xiong, W., Du, J., and Chen, Y · 2023
Closest in time.
Image captioners are scalable vision learners too
Tschannen, M., Kumar, M., Steiner, A., Zhai, X., Houlsby, N., and Beyer, L · 2023
Closest in time.
Equivariant similarity for vision-language foundation models
Wang, T., Lin, K., Li, L., Lin, C.-C., Yang, Z., Zhang, H., Liu, Z., and Wang, L · 2023
Closest in time.
Multimodal dataset distillation for image-text retrieval
Wu, X., Deng, Z., and Russakovsky, O · 2023
Closest in time.
What you see is what you read? improving text-image alignment evaluation
Yarom, M., Bitton, Y., Changpinyo, S., Aharoni, R., Herzig, J., Lang, O., Ofek, E., and Szpektor, I · 2023
Closest in time.
Evaluating and improving compositional text-to-visual generation
Li, B., Lin, Z., Pathak, D., Li, J., Fei, Y., Wu, K., Xia, X., Zhang, P., Neubig, G., and Ramanan, D · 2024
Closest in time.
Evaluating text-to-visual generation with image-to-text generation
Lin, Z., Pathak, D., Li, B., Li, J., Xia, X., Neubig, G., Zhang, P., and Ramanan, D · 2024
Closest in time.
The neglected tails of vision-language models
Parashar, S., Lin, Z., Liu, T., Dong, X., Li, Y., Ramanan, D., Caverlee, J., and Kong, S · 2024
Closest in time.