Fetching the paper…
Reading the bibliography…
Multimodal large language models (MLLMs) have revolutionized vision-language understanding but remain vulnerable to multimodal jailbreak attacks, where adversarial inputs are meticulously crafted to elicit harmful or inappropriate responses.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Earlier work this paper cites.
Sparse and imperceivable adversarial attacks
Croce, F. and Hein, M · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Diffusion models for adversarial purification
Nie, W., Guo, B., Huang, Y., Xiao, C., Vahdat, A., and Anandkumar, A · 2022
Earlier work this paper cites.
A-okvqa: A benchmark for visual question answering using world knowledge
Schwenk, D., Khandelwal, A., Clark, C., Marino, K., and Mottaghi, R · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Awadalla, A., Gao, I., Gardner, J., Hessel, J., Hanafy, Y., Zhu, W., Marathe, K., Bitton, Y., Gadre, S., Sagawa, S., et al · 2023
Earlier work this paper cites.
Image hijacks: Adversarial images can control generative models at runtime, 2023
Bailey, L., Ong, E., Russell, S., and Emmons, S · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Earlier work this paper cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning. arxiv 2023
Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Earlier work this paper cites.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y · 2023
Cited alongside, same era.
Online advertisements with llms: Opportunities and challenges
Feizi, S., Hajiaghayi, M., Rezaei, K., and Shin, S · 2023
Cited alongside, same era.
Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries
Jin, Y., Chandra, M., Verma, G., Hu, Y., De Choudhury, M., and Kumar, S · 2023
Cited alongside, same era.
Adversarial robustness of prompt-based few-shot learning for natural language understanding
Nookala, V. P. S., Verma, G., Mukherjee, S., and Kumar, S · 2023
Cited alongside, same era.
Pointing out human answer mistakes in a goal-oriented visual dialogue
Oshima, R., Shinagawa, S., Tsunashima, H., Feng, Q., and Morishima, S · 2023
Cited alongside, same era.
Deng, C., Duan, Y., Jin, X., Chang, H., Tian, Y., Liu, H., Zou, H. P., Jin, Y., Xiao, Y., Wang, Y., et al · 2024
Closest in time.
Musechat: A conversational music recommendation system for videos
Dong, Z., Liu, X., Chen, B., Polak, P., and Zhang, P · 2024
Closest in time.
Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation
Gou, Y., Chen, K., Liu, Z., Hong, L., Xu, H., Li, Z., Yeung, D.-Y., Kwok, J. T., and Zhang, Y · 2024
Closest in time.
Mm-soc: Benchmarking multimodal large language models in social media platforms
Jin, Y., Choi, M., Verma, G., Wang, J., and Kumar, S · 2024
Closest in time.
Gs2p: a generative pre-trained learning to rank model with over-parameterization for web-scale search
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visual adversarial examples jailbreak aligned large language models, 2023
Qi, X., Huang, K., Panda, A., Henderson, P., Wang, M., and Mittal, P · 2023
Cited alongside, same era.
Smoothllm: Defending large language models against jailbreaking attacks
Robey, A., Wong, E., Hassani, H., and Pappas, G · 2023
Cited alongside, same era.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, E., Dong, Y., and Abu-Ghazaleh, N · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Foundation model-oriented robustness: Robust image model evaluation with pretrained models
Zhang, P., Liu, H., Li, C., Xie, X., Kim, S., and Wang, H · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Cited alongside, same era.
Li, Y., Xiong, H., Kong, L., Bian, J., Wang, S., Chen, G., and Yin, D · 2024
Closest in time.
Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering
Liu, X., Dong, Z., and Zhang, P · 2024
Closest in time.
MUFFIN: Curating multi-faceted instructions for improving instruction following
Lou, R., Zhang, K., Xie, J., Sun, Y., Ahn, J., Xu, H., su, Y., and Yin, W · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Niu, Z., Ren, H., Gao, X., Hua, G., and Jin, R · 2024
Closest in time.
Mllm-protector: Ensuring mllm’s safety without hurting performance
Pi, R., Han, T., Xie, Y., Pan, R., Lian, Q., Dong, H., Zhang, J., and Zhang, T · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Closest in time.
Large language models can be good privacy protection learners
Xiao, Y., Jin, Y., Bai, Y., Wu, Y., Yang, X., Luo, X., Yu, W., Zhao, X., Liu, Y., Chen, H., et al · 2024
Closest in time.
When search engine services meet large language models: Visions and challenges
Xiong, H., Bian, J., Li, Y., Li, X., Du, M., Wang, S., Yin, D., and Helal, S · 2024
Closest in time.
Competeai: Understanding the competition behaviors in large language model-based agents
Zhao, Q., Wang, J., Zhang, Y., Jin, Y., Zhu, K., Chen, H., and Xie, X · 2024
Closest in time.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models
Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Hospedales, T · 2024
Closest in time.
Liu, S., Jin, Y., Li, C., Wong, D. F., Wen, Q., Sun, L., Chen, H., Xie, X., and Wang, J · 2025
Closest in time.