Fetching the paper…
Reading the bibliography…
Vision-Language Models (VLMs) have achieved remarkable performance on a variety of tasks, yet they remain vulnerable to jailbreak attacks that compromise safety and reliability.
Image segmentation techniques
Haralick, R. M. and Shapiro, L. G · 1985
Earlier work this paper cites.
Fundamentals of digital image processing
Jain, A. K · 1989
Earlier work this paper cites.
Concepts and applications of voronoi diagrams, 1992
Tessellations, S · 1992
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S. and Lavie, A · 2005
Earlier work this paper cites.
Elements of information theory
Cover, T. and Thomas, J · 2006
Earlier work this paper cites.
Digital image processing
Gonzalez, R. C · 2009
Earlier work this paper cites.
Automatic keyword extraction from individual documents
Rose, S., Engel, D., Cramer, N., and Cowley, W · 2010
Earlier work this paper cites.
Adversarial examples are not easily detected: Bypassing ten detection methods
Carlini, N. and Wagner, D · 2017
Earlier work this paper cites.
Detecting adversarial samples from artifacts
Feinman, R., Curtin, R. R., Shintre, S., and Gardner, A. B · 2017
Earlier work this paper cites.
Early methods for detecting adversarial images
Hendrycks, D. and Gimpel, K · 2017
Earlier work this paper cites.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Lee, K., Lee, K., Lee, H., and Shin, J · 2018
Earlier work this paper cites.
Feature squeezing: Detecting adversarial exa mples in deep neural networks
Xu, W · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Earlier work this paper cites.
Nic: Detecting adversarial samples with neural network invariant checking
Ma, S. and Liu, Y · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
Uniter: Universal image-text representation learning
Chen, Y.-C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J · 2020
Earlier work this paper cites.
Detoxify
Hanu, L. and Unitary team · 2020
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X., Yin, X., Li, C., Zhang, P., Hu, X., Zhang, L., Wang, L., Hu, H., Dong, L., Wei, F., Choi, Y., and Gao, J · 2020
Earlier work this paper cites.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Mathur, N., Baldwin, T., and Cohn, T · 2020
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Su, W., Zhu, X., Cao, Y., Li, B., Lu, L., Wei, F., and Dai, J · 2020
Earlier work this paper cites.
On adaptive attacks to adversarial example defenses
Tramer, F., Carlini, N., Brendel, W., and Madry, A · 2020
Earlier work this paper cites.
Gat: Generative adversarial training for adversarial example detection and robust classification
Yin, X., Kolouri, S., and Rohde, G. K · 2020
Cited alongside, same era.
Adversarial attacks on deep-learning models in natural language processing: A survey
Zhang, W. E., Sheng, Q. Z., Alhazmi, A., and Li, C · 2020
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Provably robust classification of adversarial examples with detection
Sheikholeslami, F., Lotfi, A., and Kolter, J. Z · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, J · 2021
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Jailbroken: How Does LLM Safety Training Fail?, July 2023
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Defending ChatGPT against jailbreak attack via self-reminders
Xie, Y., Yi, J., Shao, J., Curl, J., Lyu, L., Chen, Q., Xie, X., and Wu, F · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, October 2023
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models, December 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Bao, H., Wang, W., Dong, L., Liu, Q., Mohammed, O. K., Aggarwal, K., Som, S., Piao, S., and Wei, F · 2022
Cited alongside, same era.
Be your own neighborhood: Detecting adversarial examples by the neighborhood relations built on self-supervised learning
He, Z., Yang, Y., Chen, P.-Y., Xu, Q., and Ho, T.-Y · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Detecting adversarial examples is (nearly) as hard as classifying them
Tramer, F · 2022
Cited alongside, same era.
Towards adversarial attack on vision-language pre-training models
Zhang, J., Yi, Q., and Sang, J · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Unveiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model
Cheng, H., Xiao, E., Gu, J., Yang, L., Duan, J., Zhang, J., Cao, J., Xu, K., and Xu, R · 2024
Closest in time.
Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., Li, W., Shen, Y., Ma, S., Liu, H., et al · 2024
Closest in time.
Cold-attack: Jailbreaking llms with stealthiness and controllability
Guo, X., Yu, F., Zhang, H., Qin, L., and Hu, B · 2024
Closest in time.
Catastrophic jailbreak of open-source llms via exploiting generation
Huang, Y., Gupta, S., Xia, M., Li, K., and Chen, D · 2024
Closest in time.
Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models
Li, Y., Guo, H., Zhou, K., Zhao, W. X., and Wen, J · 2024
Closest in time.
Luo, W., Ma, S., Liu, X., Guo, X., and Xiao, C · 2024
Closest in time.
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically, February 2024
Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A · 2024
Closest in time.
Sca: Highly efficient semantic-consistent unrestricted adversarial attack
Pan, Z., Wu, W., Cao, Y., and Zheng, Z · 2024
Closest in time.
Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework, August 2024
Pisano, M., Ly, P., Sanders, A., Yao, B., Wang, D., Strzalkowski, T., and Si, M · 2024
Closest in time.
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks, June 2024
Robey, A., Wong, E., Hassani, H., and Pappas, G. J · 2024
Closest in time.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Shayegani, E., Dong, Y., and Abu-Ghazaleh, N · 2024
Closest in time.
Highly transferable diffusion-based unrestricted adversarial attack on pre-trained vision-language models
Xu, W., Chen, K., Gao, Z., Wei, Z., Chen, J., and Jiang, Y.-G · 2024
Closest in time.
Jailbreak vision language models via bi-modal adversarial prompt, 2024
Ying, Z., Liu, A., Zhang, T., Yu, Z., Liang, S., Liu, X., and Tao, D · 2024
Closest in time.
Low-Resource Languages Jailbreak GPT-4, January 2024
Yong, Z.-X., Menghini, C., and Bach, S. H · 2024
Closest in time.
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts, June 2024
Yu, J., Lin, X., Yu, Z., and Xing, X · 2024
Closest in time.
Zhang, J., Ye, J., Ma, X., Li, Y., Yang, Y., Sang, J., and Yeung, D.-Y · 2024
Closest in time.
Retention score: Quantifying jailbreak risks for vision language models
Li, Z., Chen, P.-Y., and Ho, T.-Y · 2025
Closest in time.