Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have showcased impressive performance in a variety of multimodal tasks.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Gray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Earlier work this paper cites.
Ignore Previous Prompt: Attack Techniques For Language Models
Perez, F.; and Ribeiro, I. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Earlier work this paper cites.
Gaining wisdom from setbacks: Aligning large language models via mistake analysis
Chen, K.; Wang, C.; Yang, K.; Han, J.; Hong, L.; Mi, F.; Xu, H.; Liu, Z.; Huang, W.; Li, Z.; et al. 2023 · 2023
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J.; Pan, X.; Sun, R.; Ji, J.; Xu, X.; Liu, M.; Wang, Y.; and Yang, Y. 2023 · 2023
Earlier work this paper cites.
Figstep: Jailbreaking large vision-language models via typographic visual prompts
Gong, Y.; Ran, D.; Liu, J.; Wang, C.; Cong, T.; Wang, A.; Duan, S.; and Wang, X. 2023 · 2023
Earlier work this paper cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Cited alongside, same era.
An Image Is Worth 1000 Lies: Transferability of Adversarial Images across Prompts on Vision-Language Models
Luo, H.; Gu, J.; Liu, F.; and Torr, P. 2023 · 2023
Cited alongside, same era.
Visual adversarial examples jailbreak large language models
Qi, X.; Huang, K.; Panda, A.; Wang, M.; and Mittal, P. 2023 · 2023
Cited alongside, same era.
Universal jailbreak backdoors from poisoned human feedback
Rando, J.; and Tramèr, F. 2023 · 2023
Cited alongside, same era.
On the adversarial robustness of multi-modal foundation models
Schlarmann, C.; and Hein, M. 2023 · 2023
Cited alongside, same era.
Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation
Gou, Y.; Chen, K.; Liu, Z.; Hong, L.; Xu, H.; Li, Z.; Yeung, D.-Y.; Kwok, J. T.; and Zhang, Y. 2024 · 2024
Closest in time.
Li, Y.; Guo, H.; Zhou, K.; Zhao, W. X.; and Wen, J.-R. 2024 · 2024
Closest in time.
Luo, W.; Ma, S.; Liu, X.; Guo, X.; and Xiao, C. 2024 · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Niu, Z.; Ren, H.; Gao, X.; Hua, G.; and Jin, R. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, E.; Dong, Y.; and Abu-Ghazaleh, N. 2023 · 2023
Cited alongside, same era.
How many unicorns are in this image? a safety evaluation benchmark for vision llms
Tu, H.; Cui, C.; Wang, Z.; Zhou, Y.; Zhao, B.; Han, J.; Zhou, W.; Yao, H.; and Xie, C. 2023 · 2023
Cited alongside, same era.
A mutation-based method for multi-modal jailbreaking attack detection
Zhang, X.; Zhang, C.; Li, T.; Huang, Y.; Jia, X.; Xie, X.; Liu, Y.; and Shen, C. 2023 · 2023
Cited alongside, same era.
On evaluating adversarial robustness of large vision-language models
Zhao, Y.; Pang, T.; Du, C.; Yang, X.; Li, C.; Cheung, N.-M. M.; and Lin, M. 2023 · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A.; Wang, Z.; Kolter, J. Z.; and Fredrikson, M. 2023 · 2023
Cited alongside, same era.
He, Y.; Li, Y.; Wu, J.; Sui, Y.; Chen, Y.; and Hooi, B. 2025a
Cited in the paper.
UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs
He, Y.; Sui, Y.; He, X.; Liu, Y.; Sun, Y.; and Hooi, B. 2025b
Cited in the paper.
Pi, R.; Han, T.; Xie, Y.; Pan, R.; Lian, Q.; Dong, H.; Zhang, J.; and Zhang, T. 2024 · 2024
Closest in time.
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
Tao, X.; Zhong, S.; Li, L.; Liu, Q.; and Kong, L. 2024 · 2024
Closest in time.
Backdooring instruction-tuned large language models with virtual prompt injection
Yan, J.; Yadav, V.; Li, S.; Chen, L.; Tang, Z.; Wang, H.; Srinivasan, V.; Ren, X.; and Jin, H. 2024 · 2024
Closest in time.
Extracting Prompts by Inverting LLM Outputs
Zhang, C.; Morris, J. X.; and Shmatikov, V. 2024 · 2024
Closest in time.
Spa-vl: A comprehensive safety preference alignment dataset for vision language model
Zhang, Y.; Chen, L.; Zheng, G.; Gao, Y.; Zheng, R.; Fu, J.; Yin, Z.; Jin, S.; Qiao, Y.; Huang, X.; et al. 2024 · 2024
Closest in time.
GuardReasoner: Towards Reasoning-based LLM Safeguards
Liu, Y.; Gao, H.; Zhai, S.; Jun, X.; Wu, T.; Xue, Z.; Chen, Y.; Kawaguchi, K.; Zhang, J.; and Hooi, B. 2025 · 2025
Closest in time.