Fetching the paper…
Reading the bibliography…
With the advent of Large Vision-Language Models (LVLMs), new attack vectors, such as cognitive bias, prompt injection, and jailbreaking, have emerged.
Quantifying perceptual distortion of adversarial examples
Matt Jordan, Naren Manoj, Surbhi Goel, et al · 1902
Earlier work this paper cites.
On the effectiveness of low frequency perturbations
Yash Sharma, Gavin Weiguang Ding, and Marcus Brubaker. 2019 · 1903
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu. 2019 · 1907
Earlier work this paper cites.
Visualizing and understanding the effectiveness of BERT
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2019 · 1908
Earlier work this paper cites.
An alternative surrogate loss for pgd-based adversarial testing
Sven Gowal, Jonathan Uesato, et al · 1910
Earlier work this paper cites.
The gradient projection method for nonlinear programming. Part I. Linear constraints
Jo Bo Rosen. 1960 · 1960
Earlier work this paper cites.
Discrete Cosine Transfom
Nasir Ahmed, T Natarajan, and Kamisetty R Rao. 1974 · 1974
Earlier work this paper cites.
Nonlinear total variation based noise removal algorithms
Leonid I Rudin, Stanley Osher, et al · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998 · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In ACL
Kishore Papineni et al · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, et al · 2004
Earlier work this paper cites.
Histograms of oriented gradients for human detection. In CVPR
Navneet Dalal and Bill Triggs. 2005 · 2005
Earlier work this paper cites.
Towards frequency-based explanation for robust cnn
Zifan Wang, Yilin Yang, Ankit Shrivastava, et al · 2005
Earlier work this paper cites.
Perceptual adversarial robustness: Defense against unseen threat models
Cassidy Laidlaw et al · 2006
Earlier work this paper cites.
Labeled faces in the wild: A database forstudying face recognition in unconstrained environments. In ECCV Workshop
Gary B Huang et al · 2008
Earlier work this paper cites.
Nus-wide: a real-world web image database from national university of singapore. In ACM CIVR
Tat-Seng Chua, Jinhui Tang, et al · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. In CVPR
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman et al · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images. Toronto, ON, Canada
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Attribute and simile classifiers for face verification. In ICCV
Neeraj Kumar, Alexander C Berg, Peter N Belhumeur, and Shree K Nayar. 2009 · 2009
Earlier work this paper cites.
Hatebert: Retraining bert for abusive language detection in english
Tommaso Caselli et al · 2010
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy. 2020 · 2010
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, et al · 2010
Earlier work this paper cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, et al · 2010
Earlier work this paper cites.
Collecting image annotations using amazon’s mechanical turk. In NAACL-HLT Workshop
Cyrus Rashtchian, Peter Young, et al · 2010
Earlier work this paper cites.
A new approach to cross-modal multimedia retrieval. In ACM MM
Nikhil Rasiwasia, Jose Costa Pereira, Emanuele Coviello, et al · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning. In AISTATS
Adam Coates, Andrew Ng, and Honglak Lee. 2011 · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning. In NIPS Workshop
Yuval Netzer, Tao Wang, Adam Coates, et al · 2011
Earlier work this paper cites.
Patch-wise++ perturbation for adversarial targeted attacks
Lianli Gao, Qilong Zhang, Jingkuan Song, et al · 2012
Earlier work this paper cites.
Vision-based traffic sign detection and analysis for intelligent driver assistance systems: Perspectives and survey
Andreas Mogelmose, Mohan Manubhai Trivedi, and Thomas B Moeslund. 2012 · 2012
Earlier work this paper cites.
Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition
Johannes Stallkamp et al · 2012
Earlier work this paper cites.
Evasion attacks against machine learning at test time. In ECML PKDD
Battista Biggio, Igino Corona, Davide Maiorca, et al · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma. 2013 · 2013
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning. In ICASSP
Quoc V Le. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan. 2013 · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR
Ross Girshick, Jeff Donahue, et al · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes. In EMNLP
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, et al · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, et al · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, et al · 2015
Earlier work this paper cites.
Model inversion attacks that exploit confidence information and basic countermeasures. In ACM SIGSAC
Matt Fredrikson et al · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild. In ICCV
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015 · 2015
Earlier work this paper cites.
Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In CVPR
Anh Nguyen et al · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In ICCV
Bryan A Plummer et al · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, et al · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation. In CVPR
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop
Fisher Yu et al · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding. In CVPR
Marius Cordts, Mohamed Omran, Sebastian Ramos, et al · 2016
Earlier work this paper cites.
Robustness of classifiers: from adversarial to random noise. In NIPS
Alhussein Fawzi, Seyed-Mohsen Moosavi-Dezfooli, et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In CVPR
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Deep networks with stochastic depth. In ECCV
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. 2016 · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution. In ECCV
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Adversarial machine learning at scale
Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2016 · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric. In ICML
Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, et al · 2016
Earlier work this paper cites.
Delving into transferable adversarial examples and black-box attacks
Yanpei Liu, Xinyun Chen, et al · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions. In CVPR
Junhua Mao, Jonathan Huang, Alexander Toshev, et al · 2016
Earlier work this paper cites.
Deepfool: a simple and accurate method to fool deep neural networks. In CVPR
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, et al · 2016
Earlier work this paper cites.
Simple black-box adversarial perturbations for deep networks
Nina Narodytska and Shiva Prasad Kasiviswanathan. 2016 · 2016
Earlier work this paper cites.
Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In CCS
Mahmood Sharif, Sruti Bhagavatula, et al · 2016
Earlier work this paper cites.
Adversarial images for variational autoencoders
Pedro Tabacof, Julia Tavares, and Eduardo Valle. 2016 · 2016
Earlier work this paper cites.
Exploring the space of adversarial images. In IJCNN
Pedro Tabacof and Eduardo Valle. 2016 · 2016
Earlier work this paper cites.
A boundary tilting persepective on the phenomenon of adversarial examples
Thomas Tanay and Lewis Griffin. 2016 · 2016
Earlier work this paper cites.
Stealing machine learning models via prediction { \{ APIs } \} . In USENIX Security
Florian Tramèr, Fan Zhang, Ari Juels, Michael K Reiter, et al · 2016
Earlier work this paper cites.
Modeling context in referring expressions. In ECCV
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. 2016 · 2016
Earlier work this paper cites.
Sergey Zagoruyko. 2016 · 2016
Earlier work this paper cites.
ImageNet-Compatible
NIPS 2017 Adversarial Attack and Defense Competition. 2017 · 2017
Earlier work this paper cites.
Adversarial transformation networks: Learning to generate adversarial examples
Shumeet Baluja and Ian Fischer. 2017 · 2017
Earlier work this paper cites.
Decision-based adversarial attacks: Reliable attacks against black-box machine learning models
W Brendel et al · 2017
Earlier work this paper cites.
Tom B Brown, Dandelion Mané, Aurko Roy, Martín Abadi, and Justin Gilmer. 2017 · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks. In S&P
Nicholas Carlini and David Wagner. 2017 · 2017
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. In AISec
P Chen et al · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences. In NIPS
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, et al · 2017
Earlier work this paper cites.
Houdini: Fooling deep structured prediction models
Moustapha Cisse, Yossi Adi, Natalia Neverova, et al · 2017
Earlier work this paper cites.
Towards interpretable deep neural networks by leveraging adversarial examples
Yinpeng Dong, Hang Su, et al · 2017
Earlier work this paper cites.
The robustness of deep networks: A geometrical perspective
Alhussein Fawzi et al · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In CVPR
Yash Goyal, Tejas Khot, et al · 2017
Earlier work this paper cites.
Countering adversarial images using input transformations
Chuan Guo, Mayank Rana, et al · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NIPS
Martin Heusel, Hubert Ramsauer, et al · 2017
Earlier work this paper cites.
Adversarial attacks on neural network policies
Sandy Huang, Nicolas Papernot, Ian Goodfellow, et al · 2017
Earlier work this paper cites.
Perspective API
Google Jigsaw. 2017 · 2017
Earlier work this paper cites.
Tactics of adversarial attack on deep reinforcement learning agents
Yen-Chen Lin et al · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry. 2017 · 2017
Earlier work this paper cites.
Universal adversarial perturbations. In CVPR
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. 2017 · 2017
Earlier work this paper cites.
NIPS 2017 Competition Track
NIPS. 2017 · 2017
Earlier work this paper cites.
Practical black-box attacks against machine learning. In ACM ASIACCS
Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, et al · 2017
Earlier work this paper cites.
Deepxplore: Automated whitebox testing of deep learning systems. In SOSP
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana. 2017 · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization. In CVPR
Ramprasaath R Selvaraju et al · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models. In SP
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Earlier work this paper cites.
Feature squeezing: Detecting adversarial exa mples in deep neural networks
W Xu. 2017 · 2017
Earlier work this paper cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
Jieyu Zhao et al · 2017
Earlier work this paper cites.
Threat of adversarial attacks on deep learning in computer vision: A survey
Naveed Akhtar and Ajmal Mian. 2018 · 2018
Earlier work this paper cites.
Synthesizing robust adversarial examples. In ICML
Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok. 2018 · 2018
Earlier work this paper cites.
Unrestricted adversarial examples
Tom B Brown, Nicholas Carlini, Chiyuan Zhang, Catherine Olsson, et al · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin. 2018 · 2018
Cited alongside, same era.
Measuring and mitigating unintended bias in text classification. In AIES
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, et al · 2018
Cited alongside, same era.
Boosting adversarial attacks with momentum. In CVPR
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, et al · 2018
Cited alongside, same era.
Robust physical-world attacks on deep learning visual classification. In CVPR
Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, et al · 2018
Cited alongside, same era.
Diffusion models for adversarial purification
Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, et al · 2022
Later among the works it cites.
Moderation API
OpenAI. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback. In NIPS
Long Ouyang, Jeffrey Wu, Xu Jiang, et al · 2022
Later among the works it cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Later among the works it cites.
Boosting the transferability of adversarial attacks with reverse adversarial perturbation. In NIPS
Zeyu Qin, Yanbo Fan, Yi Liu, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GeekPwn CAAD 2018
GeekPwn. 2018 · 2018
Cited alongside, same era.
Adversarial attacks on variational autoencoders
George Gondim-Ribeiro, Pedro Tabacof, and Eduardo Valle. 2018 · 2018
Cited alongside, same era.
Low frequency adversarial perturbation
Chuan Guo, Jared S Frank, and Kilian Q Weinberger. 2018 · 2018
Cited alongside, same era.
Black-box adversarial attacks with limited queries and information. In ICML
Andrew Ilyas, Logan Engstrom, Anish Athalye, and Jessy Lin. 2018 · 2018
Cited alongside, same era.
Harini Kannan, Alexey Kurakin, and Ian Goodfellow. 2018 · 2018
Cited alongside, same era.
Adversarial examples for generative models. In SPW
Jernej Kos, Ian Fischer, and Dawn Song. 2018 · 2018
Cited alongside, same era.
Adversarial examples in the physical world
Alexey Kurakin, Ian J Goodfellow, and Samy Bengio. 2018 · 2018
Cited alongside, same era.
Robin Rombach, Andreas Blattmann, Dominik Lorenz, et al · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models. In NIPS
Jason Wei, Xuezhi Wang, Dale Schuurmans, et al · 2022
Later among the works it cites.
Stochastic variance reduced ensemble adversarial attack for boosting the adversarial transferability. In CVPR
Yifeng Xiong, Jiadong Lin, et al · 2022
Later among the works it cites.
Ila-da: Improving transferability of intermediate level attack with data augmentation. In ICLR
Chiu Wai Yan, Tsz-Him Cheung, et al · 2022
Later among the works it cites.
Vision-language pre-training with triple contrastive learning. In CVPR
Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, et al · 2022
Later among the works it cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Anas Awadalla et al · 2023
Later among the works it cites.
(Ab) using Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs
Eugene Bagdasaryan et al · 2023
Later among the works it cites.
Ernie bot Introduction Report
Baidu. 2023 · 2023
Later among the works it cites.
Image hijacks: Adversarial images can control generative models at runtime
Luke Bailey, Euan Ong, et al · 2023
Later among the works it cites.
One transformer fits all distributions in multi-modal diffusion at scale. In ICML
Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, et al · 2023
Later among the works it cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, et al · 2023
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning. In CVPR
Mehdi Cherti, Romain Beaumont, Ross Wightman, et al · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, et al · 2023
Later among the works it cites.
Image-aware Decoder Enhanced à la Flamingo with Interleaved Cross-attentionS
Hugging Face and LAION. 2023 · 2023
Later among the works it cites.
Eva: Exploring the limits of masked visual representation learning at scale. In CVPR
Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, et al · 2023
Later among the works it cites.
Misusing tools in large language models with visual adversarial examples
Xiaohan Fu, Zihan Wang, et al · 2023
Later among the works it cites.
Imagebind: One embedding space to bind them all. In CVPR
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, et al · 2023
Later among the works it cites.
Bard (Gemini)
Google. 2023 · 2023
Later among the works it cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In AISec
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Later among the works it cites.
From images to textual prompts: Zero-shot visual question answering with frozen large language models. In CVPR
Jiaxian Guo et al · 2023
Later among the works it cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, et al · 2023
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, et al · 2023
Later among the works it cites.
Reaching 80 zero-shot accuracy with openclip: Vit-g/14 trained on laion-2b
LAION. 2023 · 2023
Later among the works it cites.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Zi Lin et al · 2023
Later among the works it cites.
Vicuna-7b-v1.5
LMSYS. 2023 · 2023
Later among the works it cites.
Tdc 2023 (llm edition): The trojan detection challenge. In NIPS Competition Track
Zou A. Mu N. Phan L. Mazeika, M. et al · 2023
Later among the works it cites.
Introducing MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs
Mosaic. 2023 · 2023
Later among the works it cites.
Scalable extraction of training data from (production) language models
Milad Nasr, Nicholas Carlini, et al · 2023
Later among the works it cites.
Nous Hermes 2 - Yi-34B
NousResearch. 2023 · 2023
Later among the works it cites.
Llm self defense: By self examination, llms know they are being tricked
Mansi Phute, Alec Helbling, et al · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi et al · 2023
Later among the works it cites.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger et al · 2023
Later among the works it cites.
On the adversarial robustness of multi-modal foundation models. In ICCV
Christian Schlarmann and Matthias Hein. 2023 · 2023
Later among the works it cites.
New jailbreak! Proudly unveiling the tried and tested DAN 5.0
SessionGloomy. 2023 · 2023
Later among the works it cites.
Glaze: Protecting artists from style mimicry by { \{ Text-to-Image } \} models. In USENIX Security
Shawn Shan, Jenna Cryan, et al · 2023
Later among the works it cites.
Tiny lvlm-ehub: Early multimodal experiments with bard
Wenqi Shao, Yutao Hu, Peng Gao, Meng Lei, et al · 2023
Later among the works it cites.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models. In ICLR
Erfan Shayegani, Yue Dong, et al · 2023
Later among the works it cites.
ChatGPT gives you free Windows 10 Pro keys!
Sid. 2023 · 2023
Later among the works it cites.
Pandagpt: One model to instruction-follow them all
Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, et al · 2023
Later among the works it cites.
Releasing 3B and 7B RedPajama-INCITE family of models including base, instruction-tuned & chat models
Together. 2023 · 2023
Later among the works it cites.
VisualGLM-6B
Tsinghua. 2023 · 2023
Later among the works it cites.
How many unicorns are in this image? a safety evaluation benchmark for vision llms
Haoqin Tu et al · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, et al · 2023
Later among the works it cites.
DAN is my new friend
Walkerspider. 2023 · 2023
Later among the works it cites.
Shared adversarial unlearning: Backdoor mitigation by unlearning shared adversarial examples. In NIPS
Shaokui Wei, Mingda Zhang, et al · 2023
Later among the works it cites.
Jailbreaking gpt-4v via self-adversarial attacks with system prompts
Dong Yin, Raphael Gontijo Lopes, et al · 2023
Later among the works it cites.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, et al · 2023
Later among the works it cites.
Robust physical-world attacks on face recognition
Xin Zheng, Yanbo Fan, Baoyuan Wu, Yong Zhang, Jue Wang, and Shirui Pan. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, et al · 2023
Later among the works it cites.
CosPGD: an efficient white-box adversarial attack for pixel-wise prediction tasks. In ICML
Shashank Agnihotri, Steffen Jung, et al · 2024
Closest in time.
Llama 2 - acceptable use policy
Meta AI. 2024 · 2024
Closest in time.
Adversarial attacks and countermeasures on image classification-based deep learning models in autonomous driving systems: A systematic review
Bakary Badjie, José Cecílio, and Antonio Casimiro. 2024 · 2024
Closest in time.
Are aligned neural networks adversarially aligned?. In NIPS
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, et al · 2024
Closest in time.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Patrick Chao et al · 2024
Closest in time.
How deep learning sees the world: A survey on adversarial attacks & defenses
Joana C Costa, Tiago Roxo, et al · 2024
Closest in time.
On the robustness of large multimodal models against image adversarial attacks. In CVPR
Xuanming Cui, Alejandro Aparcedo, et al · 2024
Closest in time.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Chaoyou Fu et al · 2024
Closest in time.
Gemini usage policies
Google. 2024 · 2024
Closest in time.
Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation
Yunhao Gou, Kai Chen, et al · 2024
Closest in time.
Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast
Xiangming Gu et al · 2024
Closest in time.
Qi Guo, Shanmin Pang, Xiaojun Jia, and Qing Guo. 2024 · 2024
Closest in time.
Pruning for protection: Increasing jailbreak resistance in aligned llms without fine-tuning
Adib Hasan et al · 2024
Closest in time.
Adversarial perturbations cannot reliably protect artists from generative ai
Robert Hönig, Javier Rando, et al · 2024
Closest in time.
Towards Transferable Targeted 3D Adversarial Attack in the Physical World. In CVPR
Yao Huang, Yinpeng Dong, Shouwei Ruan, et al · 2024
Closest in time.
Haibo Jin et al · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, et al · 2024
Closest in time.
Test-time backdoor attacks on multimodal large language models
Dong Lu, Tianyu Pang, Chao Du, et al · 2024
Closest in time.
Siyuan Ma et al · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, et al · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Zhenxing Niu, Haodong Ren, et al · 2024
Closest in time.
OpenAI Usage policies
OpenAI. 2024 · 2024
Closest in time.
Visual adversarial examples jailbreak aligned large language models. In AAAI
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, et al · 2024
Closest in time.
Vision-llms can fool themselves with self-generated typographic attacks
Maan Qraitem, Nazia Tasnim, et al · 2024
Closest in time.
" do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Xinyue Shen et al · 2024
Closest in time.
Eva-clip-18b: Scaling clip to 18 billion parameters
Quan Sun, Jinsheng Wang, Qiying Yu, et al · 2024
Closest in time.
The Wolf Within: Covert Injection of Malice into MLLM Societies via an MLLM Operative
Zhen Tan et al · 2024
Closest in time.
AI Threats: Adversarial Examples with a Quantum-Inspired Algorithm
Kuo-Chun Tseng et al · 2024
Closest in time.
Adversarial Attacks on Multimodal Agents
Chen Henry Wu, Jing Yu Koh, Ruslan Salakhutdinov, et al · 2024
Closest in time.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In CVPR
Qinghao Ye, Haiyang Xu, et al · 2024
Closest in time.
Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark. In NIPS
Zhenfei Yin, Jiong Wang, et al · 2024
Closest in time.
Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. In CVPR
Tianyu Yu et al · 2024
Closest in time.
On evaluating adversarial robustness of large vision-language models. In NIPS
Yunqing Zhao, Tianyu Pang, Chao Du, et al · 2024
Closest in time.
Sample-agnostic Adversarial Perturbation for Vision-Language Pre-training Models
Haonan Zheng et al · 2024
Closest in time.
Safety fine-tuning at (almost) no cost: A baseline for vision large language models
Yongshuo Zong et al · 2024
Closest in time.
Martin Kuo, Jianyi Zhang, Aolin Ding, et al · 2025
Closest in time.
Zhaoyi Li, Xiaohan Zhao, Dong-Dong Wu, Jiacheng Cui, and Zhiqiang Shen. 2025 · 2025
Closest in time.
Adversarial attacks in explainable machine learning: A survey of threats against models and humans
Jon Vadillo, Roberto Santana, and Jose A Lozano. 2025 · 2025
Closest in time.
Adversarial-Inspired Backdoor Defense via Bridging Backdoor and Adversarial Attacks. In AAAI
Jia-Li Yin, Weijian Wang, Wei Lin, et al · 2025
Closest in time.