Fetching the paper…
Reading the bibliography…
The integration of new modalities into frontier AI systems offers exciting capabilities, but also increases the possibility such systems can be adversarially manipulated in undesirable ways.
A technique for the measurement of attitudes
R. Likert · 1932
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples, 2015
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
Delving into transferable adversarial examples and black-box attacks
Y. Liu, X. Chen, C. Liu, and D. Song · 2016
Earlier work this paper cites.
Transferability in machine learning: from phenomena to black-box attacks using adversarial samples
N. Papernot, P. McDaniel, and I. Goodfellow · 2016
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
D. P. Kingma and J. Ba · 2017
Earlier work this paper cites.
Universal adversarial perturbations
S.-M. Moosavi-Dezfooli, A. Fawzi, O. Fawzi, and P. Frossard · 2017
Earlier work this paper cites.
Boosting adversarial attacks with momentum
Y. Dong, F. Liao, T. Pang, H. Su, J. Zhu, X. Hu, and J. Li · 2018
Earlier work this paper cites.
Understanding and enhancing the transferability of adversarial examples
L. Wu, Z. Zhu, C. Tai, et al · 2018
Earlier work this paper cites.
Feature space perturbations yield more transferable adversarial examples
N. Inkawhich, W. Wen, H. H. Li, and Y. Chen · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. D. M.-W. C. Kenton and L. K. Toutanova · 2019
Earlier work this paper cites.
Universal adversarial perturbations: A survey, 2020
A. Chaubey, N. Agrawal, K. Barnwal, K. K. Guliani, and P. Mehta · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Adversarial evaluation of multimodal models under realistic gray box assumption, 2021
I. Evtimov, R. Howes, B. Dolhansky, H. Firooz, and C. C. Ferrer · 2021
Earlier work this paper cites.
Multimodal neurons in artificial neural networks
G. Goh, N. Cammarata, C. Voss, S. Carter, M. Petrov, L. Schubert, A. Radford, and C. Olah · 2021
Earlier work this paper cites.
Reading isn’t believing: Adversarial attacks on multi-modal neurons
D. A. Noever and S. E. M. Noever · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Learning transferable adversarial perturbations
M. Salzmann et al · 2021
Earlier work this paper cites.
Reproducible scaling laws for contrastive language-image learning, 2022
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models, 2022
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V. Zhao, Y. Huang, A. Dai, H. Yu, S. Petrov, E. H. Chi, J. Dean, J. Devlin, A. Roberts, D. Zhou, Q. V. Le, and J. Wei · 2022
Earlier work this paper cites.
Eva: Exploring the limits of masked visual representation learning at scale, 2022
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, A. Jones, S. Bowman, A. Chen, T. Conerly, N. DasSarma, D. Drain, N. Elhage, S. El-Showk, S. Fort, Z. Hatfield-Dodds, T. Henighan, D. Hernandez, T. Hume, J. Jacobson, S. Johnston, S. Kravec, C. Olsson, S. Ringer, E. Tran-Johnson, D. Amodei, T. Brown, N. Joseph, S. McCandlish, C. Olah, J. Kaplan, and J. Clark · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Earlier work this paper cites.
Dual-key multimodal backdoors for visual question answering
M. Walmer, K. Sikka, I. Sur, A. Shrivastava, and S. Jha · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models, 2022
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models, 2023
A. Awadalla, I. Gao, J. Gardner, J. Hessel, Y. Hanafy, W. Zhu, K. Marathe, Y. Bitton, S. Gadre, S. Sagawa, J. Jitsev, S. Kornblith, P. W. Koh, G. Ilharco, M. Wortsman, and L. Schmidt · 2023
Earlier work this paper cites.
Abusing images and sounds for indirect instruction injection in multi-modal llms, 2023
E. Bagdasaryan, T.-Y. Hsieh, B. Nassi, and V. Shmatikov · 2023
Earlier work this paper cites.
Image hijacks: Adversarial images can control generative models at runtime
L. Bailey, E. Ong, S. Russell, and S. Emmons · 2023
Cited alongside, same era.
Pali-3 vision language models: Smaller, faster, stronger, 2023
X. Chen, X. Wang, L. Beyer, A. Kolesnikov, J. Wu, P. Voigtlaender, B. Mustafa, S. Goodman, I. Alabdulmohsin, P. Padlewski, D. Salz, X. Xiong, D. Vlasic, F. Pavetic, K. Rong, T. Yu, D. Keysers, X. Zhai, and R. Soricut · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2023
Cited alongside, same era.
Defending against unforeseen failure modes with latent adversarial training, 2024
S. Casper, L. Schulze, O. Patel, and D. Hadfield-Menell · 2024
Closest in time.
Unbridled icarus: A survey of the potential perils of image inputs in multimodal large language model security, 2024
Y. Fan, Y. Cao, Z. Zhao, Z. Liu, and S. Li · 2024
Closest in time.
Adversarial robustness for visual grounding of multimodal large language models, 2024
K. Gao, Y. Bai, J. Bai, Y. Yang, and S.-T. Xia · 2024
Closest in time.
Try bard and share your feedback
Google · 2024
Closest in time.
Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast, 2024
X. Gu, X. Zheng, T. Pang, C. Du, Q. Liu, Y. Wang, J. Jiang, and M. Lin · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Dong, H. Chen, J. Chen, Z. Fang, X. Yang, Y. Zhang, Y. Tian, H. Su, and J. Zhu · 2023
Cited alongside, same era.
Multi-attacks: Many images + + the same adversarial attack → \to many target labels, 2023
S. Fort · 2023
Cited alongside, same era.
Misusing tools in large language models with visual adversarial examples
X. Fu, Z. Wang, S. Li, R. K. Gupta, N. Mireshghallah, T. Berg-Kirkpatrick, and E. Fernandes · 2023
Cited alongside, same era.
Llama-adapter v2: Parameter-efficient visual instruction model, 2023
P. Gao, J. Han, R. Zhang, Z. Lin, S. Geng, A. Zhou, W. Zhang, P. Lu, C. He, X. Yue, H. Li, and Y. Qiao · 2023
Cited alongside, same era.
Figstep: Jailbreaking large vision-language models via typographic visual prompts, 2023
Y. Gong, D. Ran, J. Liu, C. Wang, T. Cong, A. Wang, S. Duan, and X. Wang · 2023
Cited alongside, same era.
Mistral 7b, 2023
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2023
Cited alongside, same era.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson · 2023
Cited alongside, same era.
Llava-gemma: Accelerating multimodal foundation models with a compact language model, 2024
M. Hinck, M. L. Olson, D. Cobbley, S.-Y. Tseng, and V. Lal · 2024
Closest in time.
What makes and breaks safety fine-tuning? mechanistic study, 2024
S. Jain, E. S. Lubana, K. Oksuz, T. Joy, P. H. S. Torr, A. Sanyal, and P. K. Dokania · 2024
Closest in time.
Brave: Broadening the visual encoding of vision-language models, 2024
O. F. Kar, A. Tonioni, P. Poklukar, A. Kulshrestha, A. Zamir, and F. Tombari · 2024
Closest in time.
Prismatic vlms: Investigating the design space of visually-conditioned language models, 2024
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Closest in time.
Analyzing and editing inner mechanisms of backdoored language models
M. Lamparth and A. Reuel · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity, 2024
A. Lee, X. Bai, I. Pres, M. Wattenberg, J. K. Kummerfeld, and R. Mihalcea · 2024
Closest in time.
An image is worth 1000 lies: Adversarial transferability across prompts on vision-language models, 2024
H. Luo, J. Gu, F. Liu, and P. Torr · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks · 2024
Closest in time.
Mm1: Methods, analysis & insights from multimodal llm pre-training, 2024
B. McKinzie, Z. Gan, J.-P. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, F. Weers, A. Belyi, H. Zhang, K. Singh, D. Kang, A. Jain, H. Hè, M. Schwarzer, T. Gunter, X. Kong, A. Zhang, J. Wang, C. Wang, N. Du, T. Lei, S. Wiseman, G. Yin, M. Lee, Z. Wang, R. Pang, P. Grasch, A. Toshev, and Y. Yang · 2024
Closest in time.
Universal adversarial triggers are not universal, 2024
N. Meade, A. Patel, and S. Reddy · 2024
Closest in time.
Jailbreaking attack against multimodal large language model, 2024
Z. Niu, H. Ren, X. Gao, G. Hua, and R. Jin · 2024
Closest in time.
Gpt-4v(ision) system card
OpenAI · 2024
Closest in time.
Dinov2: Learning robust visual features without supervision, 2024
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2024
Closest in time.
Tricking llms into disobedience: Formalizing, analyzing, and detecting jailbreaks, 2024
A. Rao, S. Vashistha, A. Naik, S. Aditya, and M. Choudhury · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Open problems in technical ai governance, 2024
A. Reuel, B. Bucknall, S. Casper, T. Fist, L. Soder, O. Aarne, L. Hammond, L. Ibrahim, A. Chan, P. Wills, M. Anderljung, B. Garfinkel, L. Heim, A. Trask, G. Mukobi, R. Schaeffer, M. Baker, S. Hooker, I. Solaiman, A. S. Luccioni, N. Rajkumar, N. Moës, N. Guha, J. Newman, Y. Bengio, T. South, A. Pentland, J. Ladish, S. Kyoejo, M. J. Kochenderfer, and R. Trager · 2024
Closest in time.
"do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models, 2024
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang · 2024
Closest in time.
A strongreject for empty jailbreaks, 2024
A. Souly, Q. Lu, D. Bowen, T. Trinh, E. Hsieh, S. Pandey, P. Abbeel, J. Svegliato, S. Emmons, O. Watkins, and S. Toyer · 2024
Closest in time.
Imgtrojan: Jailbreaking vision-language models with one image, 2024
X. Tao, S. Zhong, L. Li, Q. Liu, and L. Kong · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models, 2024
Y. Zong, O. Bohdal, T. Yu, Y. Yang, and T. Hospedales · 2024
Closest in time.
Improving alignment and robustness with circuit breakers, 2024
A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks · 2024
Closest in time.