Fetching the paper…
Reading the bibliography…
Benefiting from the powerful capabilities of Large Language Models (LLMs), pre-trained visual encoder models connected to an LLMs can realize Vision Language Models (VLMs).
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild, 2015
Liu, Z., Luo, P., Wang, X., and Tang, X · 2015
Earlier work this paper cites.
A hierarchical approach for generating descriptive image paragraphs, 2017
Krause, J., Johnson, J., Krishna, R., and Fei-Fei, L · 2017
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Nsfw data scraper
Kim, A · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Towards understanding and detecting cyberbullying in real-world images
Vishwamitra, N., Hu, H., Luo, F., and Cheng, L · 2021
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Earlier work this paper cites.
Knowledge mining with scene text for fine-grained recognition, 2022
Wang, H., Liao, J., Cheng, T., Gao, Z., Liu, H., Ren, B., Bai, X., and Liu, W · 2022
Earlier work this paper cites.
Privacyalert: A dataset for image privacy prediction
Zhao, C., Mangat, J., Koujalgi, S., Squicciarini, A., and Caragea, C · 2022
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J · 2023
Cited alongside, same era.
Image hijacks: Adversarial images can control generative models at runtime, 2023
Bailey, L., Ong, E., Russell, S., and Emmons, S · 2023
Cited alongside, same era.
Introducing our multimodal models, 2023
Bavishi, R., Elsen, E., Hawthorne, C., Nye, M., Odena, A., Somani, A., and Taşırlar, S · 2023
Cited alongside, same era.
Antifakeprompt: Prompt-tuned vision-language models are fake image detectors, 2023
Chang, Y.-M., Yeh, C., Chiu, W.-C., and Yu, N · 2023
Cited alongside, same era.
Sharegpt4v: Improving large multi-modal models with better captions
Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D · 2023
Cited alongside, same era.
Image safeguarding: Reasoning with conditional vision language model and obfuscating unsafe content counterfactually, 2024
Bethany, M., Wherry, B., Vishwamitra, N., and Najafirad, P · 2024
Closest in time.
Honeybee: Locality-enhanced projector for multimodal llm
Cha, J., Kang, W., Mun, J., and Roh, B · 2024
Closest in time.
Dong, X., Zhang, P., Zang, Y., Cao, Y., Wang, B., Ouyang, L., Wei, X., Zhang, S., Duan, H., Cao, M., Zhang, W., Li, Y., Yan, H., Gao, Y., Zhang, X., Li, W., Li, J., Chen, K., He, C., Zhang, X., Qiao, Y., Lin, D., and Wang, J · 2024
Closest in time.
Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., Wu, Y., and Ji, R · 2024
Closest in time.
Inducing high energy-latency of large vision-language models with verbose images, 2024
Gao, K., Bai, Y., Gu, J., Xia, S.-T., Torr, P., Li, Z., and Liu, W · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Figstep: Jailbreaking large vision-language models via typographic visual prompts, 2023
Gong, Y., Ran, D., Liu, J., Wang, C., Cong, T., Wang, A., Duan, S., and Wang, X · 2023
Cited alongside, same era.
Onellm: One framework to align all modalities with language
Han, J., Gong, K., Zhang, Y., Wang, J., Zhang, K., Lin, D., Qiao, Y., Gao, P., and Yue, X · 2023
Cited alongside, same era.
Stable bias: Evaluating societal representations in diffusion models
Luccioni, S., Akiki, C., Mitchell, M., and Jernite, Y · 2023
Cited alongside, same era.
How many unicorns are in this image? a safety evaluation benchmark for vision llms, 2023
Tu, H., Cui, C., Wang, Z., Zhou, Y., Zhao, B., Han, J., Zhou, W., Yao, H., and Xie, C · 2023
Cited alongside, same era.
Jailbroken: How does llm safety training fail?, 2023
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Cited alongside, same era.
On evaluating adversarial robustness of large vision-language models, 2023
Zhao, Y., Pang, T., Du, C., Yang, X., Li, C., Cheung, N.-M., and Lin, M · 2023
Cited alongside, same era.
Mquake: Assessing knowledge editing in language models via multi-hop questions, 2023
Zhong, Z., Wu, Z., Manning, C. D., Potts, C., and Chen, D · 2023
Cited alongside, same era.
Closest in time.
Red teaming visual language models, 2024
Li, M., Li, L., Yin, Y., Ahmed, M., Liu, Z., and Liu, Q · 2024
Closest in time.
Vl-trojan: Multimodal instruction backdoor attacks against autoregressive visual language models, 2024
Liang, J., Liang, S., Luo, M., Liu, A., Han, D., Chang, E.-C., and Cao, X · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI · 2024
Closest in time.
llama3-vision-alpha
QResearch · 2024
Closest in time.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models, 2024
Röttger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D · 2024
Closest in time.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models, 2024
Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Hospedales, T · 2024
Closest in time.