Fetching the paper…
Reading the bibliography…
The rapid advancement of large models, driven by their exceptional abilities in learning and generalization through large-scale pre-training, has reshaped the landscape of Artificial Intelligence (AI).
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton, “Adaptive mixtures of local experts,”
1991
Earlier work this paper cites.
G. W. Stewart, “On the early history of the singular value decomposition,”
1993
Earlier work this paper cites.
L. Fei-Fei, R. Fergus, and P. Perona, “Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,” in
2004
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,”
2004
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
A. Krizhevsky, G. Hinton
2009
Earlier work this paper cites.
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba, “Sun database: Large-scale scene recognition from abbey to zoo,” in
2010
Earlier work this paper cites.
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar, “Cats and dogs,” in
2012
Earlier work this paper cites.
K. Soomro, “Ucf101: A dataset of 101 human actions classes from videos in the wild,”
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in
2013
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in
2014
Earlier work this paper cites.
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, “Describing textures in the wild,” in
2014
Earlier work this paper cites.
L. Bossard, M. Guillaumin, and L. V. Gool, “Food-101–mining discriminative components with random forests,” in
2014
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein
2015
Earlier work this paper cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in
2015
Earlier work this paper cites.
Z. Liu, P. Luo, X. Wang, and X. Tang, “Deep learning face attributes in the wild,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in
2016
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,”
2017
Earlier work this paper cites.
S.-M. Moosavi-Dezfooli, A. Fawzi, O. Fawzi, and P. Frossard, “Universal adversarial perturbations,” in
2017
Earlier work this paper cites.
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba, “Scene parsing through ade20k dataset,” in
2017
Earlier work this paper cites.
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” in
2017
Earlier work this paper cites.
J. Zhu, R. Kaplan, J. Johnson, and L. Fei-Fei, “Hidden: Hiding data with deep networks,” in
2018
Earlier work this paper cites.
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in
2018
Earlier work this paper cites.
X. Xu, X. Chen, C. Liu, A. Rohrbach, T. Darrell, and D. Song, “Fooling vision and language models despite localization and attention mechanism,” in
2018
Earlier work this paper cites.
Q. Cao, L. Shen, W. Xie, O. M. Parkhi, and A. Zisserman, “Vggface2: A dataset for recognising faces across pose and age,” in
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Carlini, C. Liu, Ú. Erlingsson, J. Kos, and D. Song, “The secret sharer: Evaluating and testing unintended memorization in neural networks,” in
2019
Earlier work this paper cites.
T.-N. Le, T. V. Nguyen, Z. Nie, M.-T. Tran, and A. Sugimoto, “Anabranch network for camouflaged object segmentation,”
2019
Earlier work this paper cites.
M. Shah, X. Chen, M. Rohrbach, and D. Parikh, “Cycle-consistency for robust visual question answering,” in
2019
Earlier work this paper cites.
P. Helber, B. Bischke, A. Dengel, and D. Borth, “Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,”
2019
Earlier work this paper cites.
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar, “Do imagenet classifiers generalize to imagenet?” in
2019
Earlier work this paper cites.
H. Wang, S. Ge, Z. Lipton, and E. P. Xing, “Learning robust global representations by penalizing local predictive power,”
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. Jin, Z. Jin, J. T. Zhou, and P. Szolovits, “Is bert really robust? a strong baseline for natural language attack on text classification and entailment,” in
2020
Earlier work this paper cites.
L. Li, R. Ma, Q. Guo, X. Xue, and X. Qiu, “Bert-attack: Adversarial attack against bert using bert,” in
2020
Earlier work this paper cites.
Z. Gan, Y.-C. Chen, L. Li, C. Zhu, Y. Cheng, and J. Liu, “Large-scale adversarial training for vision-and-language representation learning,” in
2020
Earlier work this paper cites.
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “Realtoxicityprompts: Evaluating neural toxic degeneration in language models,” in
2020
Earlier work this paper cites.
F. Croce and M. Hein, “Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks,” in
2020
Earlier work this paper cites.
C. Zhu, Y. Cheng, Z. Gan, S. Sun, T. Goldstein, and J. Liu, “Freelb: Enhanced adversarial training for natural language understanding,” in
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”
2020
Earlier work this paper cites.
D. Chen, N. Yu, Y. Zhang, and M. Fritz, “Gan-leaks: A taxonomy of membership inference attacks against generative models,” in
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Joshi, G. Jagatap, and C. Hegde, “Adversarial token attacks on vision transformers,”
2021
Earlier work this paper cites.
C. Guo, A. Sablayrolles, H. Jégou, and D. Kiela, “Gradient-based adversarial attacks against text transformers,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in
2021
Earlier work this paper cites.
H. Huang, X. Ma, S. M. Erfani, J. Bailey, and Y. Wang, “Unlearnable examples: Making personal data unexploitable,” in
2021
Earlier work this paper cites.
B. Wang, C. Xu, S. Wang, Z. Gan, Y. Cheng, J. Gao, A. H. Awadallah, and B. Li, “Adversarial glue: A multi-task benchmark for robustness evaluation of language models,” in
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Earlier work this paper cites.
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in
2021
Earlier work this paper cites.
K. Yang, W.-Y. Lin, M. Barman, F. Condessa, and Z. Kolter, “Defending multimodal fusion models against single-source adversaries,” in
2021
Earlier work this paper cites.
D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song, “Natural adversarial examples,” in
2021
Earlier work this paper cites.
D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in
2021
Earlier work this paper cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Y. Fu, S. Zhang, S. Wu, C. Wan, and Y. Lin, “Patch-fool: Are vision transformers always robust against adversarial perturbations?” in
2022
Earlier work this paper cites.
G. Lovisotto, N. Finnie, M. Munoz, C. K. Mummadi, and J. H. Metzen, “Give me your attention: Dot-product attention considered harmful for adversarial patch robustness,” in
2022
Earlier work this paper cites.
Y. Wang, J. Wang, Z. Yin, R. Gong, J. Wang, A. Liu, and X. Liu, “Generating transferable adversarial examples against vision transformers,” in
2022
Earlier work this paper cites.
Z. Wei, J. Chen, M. Goldblum, Z. Wu, T. Goldstein, and Y.-G. Jiang, “Towards transferable adversarial attacks on vision transformers,” in
2022
Earlier work this paper cites.
Y. Shi, Y. Han, Y.-a. Tan, and X. Kuang, “Decision-based black-box attack against vision transformers via patch-wise adversarial removal,”
2022
Earlier work this paper cites.
B. Wu, J. Gu, Z. Li, D. Cai, X. He, and W. Liu, “Towards efficient adversarial training on vision transformers,” in
2022
Earlier work this paper cites.
J. Li, “Patch vestiges in the adversarial examples against vision transformer can be leveraged for adversarial detection,” in
2022
Earlier work this paper cites.
J. Gu, V. Tresp, and Y. Qin, “Are vision transformers robust to patch perturbations?” in
2022
Earlier work this paper cites.
Y. Mo, D. Wu, Y. Wang, Y. Guo, and Y. Wang, “When adversarial training meets vision transformers: Recipes from training to architecture,”
2022
Earlier work this paper cites.
W. Nie, B. Guo, Y. Huang, C. Xiao, A. Vahdat, and A. Anandkumar, “Diffusion models for adversarial purification,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
N. Boucher, I. Shumailov, R. Anderson, and N. Papernot, “Bad characters: Imperceptible nlp attacks,” in
2022
Earlier work this paper cites.
E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving, “Red teaming language models with language models,” in
2022
Earlier work this paper cites.
F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” in
2022
Earlier work this paper cites.
X. Cai, H. Xu, S. Xu, Y. Zhang
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
S. Chen, C. Liu, M. Haque, Z. Song, and W. Yang, “Nmtsloth: understanding and testing efficiency degradation of neural machine translation systems,” in
2022
Earlier work this paper cites.
J. Zhang, Q. Yi, and J. Sang, “Towards adversarial attack on vision-language pre-training models,” in
2022
Earlier work this paper cites.
J. Jia, Y. Liu, and N. Z. Gong, “Badencoder: Backdoor attacks to pre-trained encoders in self-supervised learning,” in
2022
Earlier work this paper cites.
N. Carlini and A. Terzis, “Poisoning and backdooring contrastive learning,” in
2022
Earlier work this paper cites.
G. Daras and A. G. Dimakis, “Discovering the hidden vocabulary of dalle-2,”
2022
Earlier work this paper cites.
R. Millière, “Adversarial attacks on image generation with made-up words,”
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Goldblum, D. Tsipras, C. Xie, X. Chen, A. Schwarzschild, D. Song, A. Mądry, B. Li, and T. Goldstein, “Dataset security for machine learning: Data poisoning, backdoor attacks, and defenses,”
2022
Earlier work this paper cites.
S. Lin, J. Hilton, and O. Evans, “Truthfulqa: Measuring how models mimic human falsehoods,” in
2022
Earlier work this paper cites.
J. Yang, J. Duan, S. Tran, Y. Xu, S. Chanda, L. Chen, B. Zeng, T. Chilimbi, and J. Huang, “Vision-language pre-training with triple contrastive learning,” in
2022
Earlier work this paper cites.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Conditional prompt learning for vision-language models,” in
2022
Earlier work this paper cites.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision
2022
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in
2022
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. N. M. Pinkney, “Pokemon blip captions,” 2022, hugging Face Dataset
2022
Earlier work this paper cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”
2022
Earlier work this paper cites.
X. Wei and S. Zhao, “Boosting adversarial transferability with learnable patch-wise masks,”
2023
Earlier work this paper cites.
W. Ma, Y. Li, X. Jia, and W. Xu, “Transferable adversarial attack for both vision transformers and convolutional networks via momentum integrated gradients,” in
2023
Earlier work this paper cites.
J. Zhang, Y. Huang, W. Wu, and M. R. Lyu, “Transferable adversarial attacks on vision transformers with token gradient regularization,” in
2023
Earlier work this paper cites.
Z. Chen, C. Xu, H. Lv, S. Liu, and Y. Ji, “Understanding and improving adversarial transferability of vision transformers and convolutional neural networks,”
2023
Earlier work this paper cites.
Z. Wei, J. Chen, M. Goldblum, Z. Wu, T. Goldstein, Y.-G. Jiang, and L. S. Davis, “Towards transferable adversarial attacks on image and video transformers,”
2023
Earlier work this paper cites.
L. Liu, Y. Guo, Y. Zhang, and J. Yang, “Understanding and defending patched-based adversarial attacks for vision transformer,” in
2023
Earlier work this paper cites.
Y. Guo, D. Stutz, and B. Schiele, “Robustifying token attention for vision transformers,” in
2023
Earlier work this paper cites.
Y. Y. Guo, D. L. Stutz, and B. T. Schiele, “Improving robustness of vision transformers by reducing sensitivity to patch corruptions,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Yuan, P. Zhou, K. Zou, and Y. Cheng, “You are catching my attention: Are vision transformers bad learners under backdoor attacks?” in
2023
Earlier work this paper cites.
M. Zheng, Q. Lou, and L. Jiang, “Trojvit: Trojan insertion in vision transformers,” in
2023
Earlier work this paper cites.
P. Lv, H. Ma, J. Zhou, R. Liang, K. Chen, S. Zhang, and Y. Yang, “Dbia: Data-free backdoor attack against transformer networks,” in
2023
Earlier work this paper cites.
K. D. Doan, Y. Lao, P. Yang, and P. Li, “Defending backdoor attacks on vision transformer via patch processing,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
S. Zheng and C. Zhang, “Black-box targeted adversarial attack on segment anything (sam),”
2023
Earlier work this paper cites.
D. Han, S. Zheng, and C. Zhang, “Segment anything meets universal adversarial perturbation,”
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Liu, C. Cai, and Y. Qi, “Expanding scope: Adapting english adversarial attacks to chinese,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
B. Liu, B. Xiao, X. Jiang, S. Cen, X. He, and W. Dou, “Adversarial attacks on large language model-based system and mitigating strategies: A case study on chatgpt,”
2023
Earlier work this paper cites.
A. Koleva, M. Ringsquandl, and V. Tresp, “Adversarial attacks on tables with entity swap,”
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z.-X. Yong, C. Menghini, and S. H. Bach, “Low-resource languages jailbreak gpt-4,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong, “Jailbreaking black box large language models in twenty queries,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
B. Chen, A. Paliwal, and Q. Yan, “Jailbreaker in jail: Moving target defense for large language models,” in
2023
Earlier work this paper cites.
Y. Liu, G. Deng, Y. Li, K. Wang, Z. Wang, X. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng
2023
Earlier work this paper cites.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection,” in
2023
Earlier work this paper cites.
B. Deng, W. Wang, F. Feng, Y. Deng, Q. Wang, and X. He, “Attack prompt generation for red teaming and defending large language models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Yan, V. Gupta, and X. Ren, “Bite: Textual backdoor attacks with iterative trigger injection,” in
2023
Earlier work this paper cites.
S. Zhao, J. Wen, L. A. Tuan, J. Zhao, and J. Fu, “Prompt as triggers for backdoor attack: Examining the vulnerability in language models,” in
2023
Earlier work this paper cites.
N. Kandpal, M. Jagielski, F. Tramèr, and N. Carlini, “Backdoor attacks for in-context learning with language models,” in
2023
Earlier work this paper cites.
N. Gu, P. Fu, X. Liu, Z. Liu, Z. Lin, and W. Wang, “A gradient control method for backdoor attacks on parameter-efficient tuning,” in
2023
Earlier work this paper cites.
X. He, J. Wang, B. Rubinstein, and T. Cohn, “Imbert: Making bert immune to insertion-based backdoor attacks,” in
2023
Earlier work this paper cites.
J. Li, Z. Wu, W. Ping, C. Xiao, and V. Vydiswaran, “Defending against insertion-based textual backdoor attacks via attribution,” in
2023
Earlier work this paper cites.
X. Sun, X. Li, Y. Meng, X. Ao, L. Lyu, J. Li, and T. Zhang, “Defending against backdoor attacks in natural language generation,” in
2023
Earlier work this paper cites.
R. R. Tang, J. Yuan, Y. Li, Z. Liu, R. Chen, and X. Hu, “Setting the trap: Capturing and defeating backdoors in pretrained language models through honeypots,”
2023
Earlier work this paper cites.
Z. Liu, B. Shen, Z. Lin, F. Wang, and W. Wang, “Maximum entropy loss, the silver bullet targeting backdoor attacks in pre-trained language models,” in
2023
Earlier work this paper cites.
G. An, J. Lee, X. Zuo, N. Kosaka, K.-M. Kim, and H. O. Song, “Direct preference-based policy optimization without reward modeling,”
2023
Earlier work this paper cites.
Y. Chen, S. Chen, Z. Li, W. Yang, C. Liu, R. Tan, and H. Li, “Dynamic transformers provide a false sense of efficiency,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Jiang, C. Chan, M. Chen, and W. Wang, “Lion: Adversarial distillation of proprietary large language models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Yu, T. Pang, Q. Liu, C. Du, B. Kang, Y. Huang, M. Lin, and S. Yan, “Bag of tricks for training data extraction from language models,” in
2023
Earlier work this paper cites.
Z. Zhou, S. Hu, M. Li, H. Zhang, Y. Zhang, and H. Jin, “Advclip: Downstream-agnostic adversarial examples in multimodal contrastive learning,” in
2023
Earlier work this paper cites.
D. Lu, Z. Wang, T. Wang, W. Guan, H. Gao, and F. Zheng, “Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Yin, M. Ye, T. Zhang, T. Du, J. Zhu, H. Liu, J. Chen, T. Wang, and F. Ma, “Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Azuma and Y. Matsui, “Defense-prefix for preventing typographic attacks on clip,” in
2023
Earlier work this paper cites.
C. Mao, S. Geng, J. Yang, X. Wang, and C. Vondrick, “Understanding zero-shot adversarial robustness for large-scale models,” in
2023
Earlier work this paper cites.
Z. Yang, X. He, Z. Li, M. Backes, M. Humbert, P. Berrang, and Y. Zhang, “Data poisoning attacks against multimodal encoders,” in
2023
Earlier work this paper cites.
H. Bansal, N. Singhi, Y. Yang, F. Yin, A. Grover, and K.-W. Chang, “Cleanclip: Mitigating data poisoning attacks in multimodal contrastive learning,” in
2023
Earlier work this paper cites.
S. Feng, G. Tao, S. Cheng, G. Shen, X. Xu, Y. Liu, K. Zhang, S. Ma, and X. Zhang, “Detecting backdoors in pre-trained encoders,” in
2023
Earlier work this paper cites.
I. Sur, K. Sikka, M. Walmer, K. Koneripalli, A. Roy, X. Lin, A. Divakaran, and S. Jha, “Tijo: Trigger inversion with joint optimization for defending multimodal backdoored models,” in
2023
Earlier work this paper cites.
C. Schlarmann and M. Hein, “On the adversarial robustness of multi-modal foundation models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Dong, H. Chen, J. Chen, Z. Fang, X. Yang, Y. Zhang, Y. Tian, H. Su, and J. Zhu, “How robust is google’s bard to adversarial image attacks?” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, “Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Zhuang, Y. Zhang, and S. Liu, “A pilot study of query-free adversarial attack against stable diffusion,” in
2023
Earlier work this paper cites.
L. Struppek, D. Hintersdorf, F. Friedrich, P. Schramowski, K. Kersting
2023
Earlier work this paper cites.
Z. Kou, S. Pei, Y. Tian, and X. Zhang, “Character as pixels: A controllable prompt adversarial attacking framework for black-box text guided image generation models.” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
N. Maus, P. Chao, E. Wong, and J. R. Gardner, “Black box adversarial prompting for foundation models,” in
2023
Earlier work this paper cites.
H. Liu, Y. Wu, S. Zhai, B. Yuan, and N. Zhang, “Riatig: Reliable and imperceptible adversarial text-to-image generation with natural prompts,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Qu, X. Shen, X. He, M. Backes, S. Zannettou, and Y. Zhang, “Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
R. Gandikota, J. Materzynska, J. Fiotto-Kaufman, and D. Bau, “Erasing concepts from diffusion models,” in
2023
Earlier work this paper cites.
S. Kim, S. Jung, B. Kim, M. Choi, J. Shin, and J. Lee, “Towards safe self-distillation of internet-scale text-to-image diffusion models,” in
2023
Earlier work this paper cites.
N. Kumari, B. Zhang, S.-Y. Wang, E. Shechtman, R. Zhang, and J.-Y. Zhu, “Ablating concepts in text-to-image diffusion models,” in
2023
Earlier work this paper cites.
Z. u. Ni, L. p. Wei, J. u. Li, S. Tang, Y. Zhuang, and Q. m. Tian, “Degeneration-tuning: Using scrambled grid shield unwanted concepts from stable diffusion,” in
2023
Earlier work this paper cites.
H. Orgad, B. Kawar, and Y. Belinkov, “Editing implicit assumptions in text-to-image diffusion models,” in
2023
Earlier work this paper cites.
P. Schramowski, M. Brack, B. Deiseroth, and K. Kersting, “Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models,” in
2023
Earlier work this paper cites.
S.-Y. Chou, P.-Y. Chen, and T.-Y. Ho, “How to backdoor diffusion models?” in
2023
Earlier work this paper cites.
W. Chen, D. Song, and B. Li, “Trojdiff: Trojan attacks on diffusion models with diverse targets,” in
2023
Earlier work this paper cites.
L. Struppek, D. Hintersdorf, and K. Kersting, “Rickrolling the artist: Injecting backdoors into text encoders for text-to-image synthesis,” in
2023
Earlier work this paper cites.
S. Zhai, Y. Dong, Q. Shen, S. Pu, Y. Fang, and H. Su, “Text-to-image diffusion models can be easily backdoored through multimodal data poisoning,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Matsumoto, T. Miura, and N. Yanai, “Membership inference attacks against diffusion models,” in
2023
Earlier work this paper cites.
H. Hu and J. Pang, “Loss and likelihood based membership inference of diffusion models,” in
2023
Earlier work this paper cites.
J. Duan, F. Kong, S. Wang, X. Shi, and K. Xu, “Are diffusion models vulnerable to membership inference attacks?” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Pang and T. Wang, “Black-box membership inference attacks against fine-tuned diffusion models,”
2023
Earlier work this paper cites.
N. Carlini, J. Hayes, M. Nasr, M. Jagielski, V. Sehwag, F. Tramer, B. Balle, D. Ippolito, and E. Wallace, “Extracting training data from diffusion models,” in
2023
Earlier work this paper cites.
R. Webster, “A reproducible extraction of training images from diffusion models,”
2023
Earlier work this paper cites.
C. Liang, X. Wu, Y. Hua, J. Zhang, Y. Xue, T. Song, Z. Xue, R. Ma, and H. Guan, “Adversarial example does good: Preventing painting imitation from diffusion models via adversarial examples,” in
2023
Earlier work this paper cites.
T. Van Le, H. Phung, T. H. Nguyen, Q. Dao, N. N. Tran, and A. Tran, “Anti-dreambooth: Protecting users from personalized text-to-image synthesis,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Cui, J. Ren, H. Xu, P. He, H. Liu, L. Sun, Y. Xing, and J. Tang, “Diffusionshield: A watermark for copyright protection against generative diffusion models,” in
2023
Earlier work this paper cites.
Z. Wang, C. Chen, L. Lyu, D. N. Metaxas, and S. Ma, “Diagnosis: Detecting unauthorized data usages in text-to-image diffusion models,” in
2023
Earlier work this paper cites.
P. Fernandez, G. Couairon, H. Jégou, M. Douze, and T. Furon, “The stable signature: Rooting watermarks in latent diffusion models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Liu, Z. Li, M. Backes, Y. Shen, and Y. Zhang, “Watermarking diffusion model,”
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Wen, J. Kirchenbauer, J. Geiping, and T. Goldstein, “Tree-ring watermarks: Fingerprints for diffusion images that are invisible and robust,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Luo, T. Zhang, Y.-S. Chuang, Y. Gong, Y. Kim, X. Wu, H. Meng, and J. Glass, “Search augmented instruction learning,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang
2023
Earlier work this paper cites.
Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian
2023
Earlier work this paper cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo
2023
Earlier work this paper cites.
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Ding, C. Liu, S. He, X. Jiang, P. H. S. Torr, and S. Bai, “MOSE: A new dataset for video object segmentation in complex scenes,” in
2023
Earlier work this paper cites.
H. Ding, C. Liu, S. He, X. Jiang, and C. C. Loy, “MeViS: A large-scale benchmark for video segmentation with motion expressions,” in
2023
Earlier work this paper cites.
N. Carlini, “A llm assisted exploitation of ai-guardian,”
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
G. Xu, J. Liu, M. Yan, H. Xu, J. Si, Z. Zhou, P. Yi, X. Gao, J. Sang, R. Zhang
2023
Earlier work this paper cites.
M. U. Khattak, H. Rasheed, M. Maaz, S. Khan, and F. S. Khan, “Maple: Multi-modal prompt learning,” in
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo
2023
Cited alongside, same era.
K. Meng, A. S. Sharma, A. Andonian, Y. Belinkov, and D. Bau, “Mass-editing memory in a transformer,” in
2023
Cited alongside, same era.
J. Dubiński, A. Kowalczuk, S. Pawlak, P. Rokita, T. Trzciński, and P. Morawiecki, “Towards more realistic membership inference attacks on large diffusion models,” in
2024
Later among the works it cites.
F. Kong, J. Duan, R. Ma, H. T. Shen, X. Shi, X. Zhu, and K. Xu, “An efficient membership inference attack for the diffusion model by proximal initialization,” in
2024
Later among the works it cites.
S. Zhai, H. Chen, Y. Dong, J. Li, Q. Shen, Y. Gao, H. Su, and Y. Liu, “Membership inference on text-to-image diffusion models via conditional likelihood discrepancy,” in
2024
Later among the works it cites.
Q. Li, X. Fu, X. Wang, J. Liu, X. Gao, J. Dai, and J. Han, “Unveiling structural memorization: Structural membership inference attack for text-to-image diffusion models,” in
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Somepalli, V. Singla, M. Goldblum, J. Geiping, and T. Goldstein, “Understanding and mitigating copying in diffusion models,”
2023
Cited alongside, same era.
X. Gu, C. Du, T. Pang, C. Li, M. Lin, and Y. Wang, “On memorization in diffusion models,”
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Z. Jiang, J. Zhang, and N. Z. Gong, “Evading watermark based detection of ai-generated content,” in
2023
Cited alongside, same era.
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman, “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
K. Navaneet, S. A. Koohpayegani, E. Sleiman, and H. Pirsiavash, “Slowformer: Adversarial attack on compute and energy consumption of efficient vision transformers,” in
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Zhang, N. Yu, R. Wen, M. Backes, and Y. Zhang, “Generated distributions are all you need for membership inference attacks against generative models,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
E. Horwitz, J. Kahana, and Y. Hoshen, “Recovering the pre-fine-tuning weights of generative models,”
2024
Later among the works it cites.
X. Ye, H. Huang, J. An, and Y. Wang, “DUAW: Data-free universal adversarial watermark against stable diffusion customization,” in
2024
Later among the works it cites.
Y. Liu, C. Fan, Y. Dai, X. Chen, P. Zhou, and L. Sun, “Metacloak: Preventing unauthorized subject-driven text-to-image diffusion-based synthesis via meta-learning,” in
2024
Later among the works it cites.
H. Liu, Z. Sun, and Y. Mu, “Countering personalized text-to-image generation with influence watermarks,” in
2024
Later among the works it cites.
F. Wang, Z. Tan, T. Wei, Y. Wu, and Q. Huang, “Simac: A simple anti-customization method for protecting face privacy against text-to-image synthesis of diffusion models,” in
2024
Later among the works it cites.
X. Zhang, R. Li, J. Yu, Y. Xu, W. Li, and J. Zhang, “Editguard: Versatile image watermarking for tamper localization and copyright protection,” in
2024
Later among the works it cites.
R. Min, S. Li, H. Chen, and M. Cheng, “A watermark-conditioned diffusion model for ip protection,” in
2024
Later among the works it cites.
P. Zhu, T. Takahashi, and H. Kataoka, “Watermark-embedded adversarial examples for copyright protection against diffusion models,” in
2024
Later among the works it cites.
V. Asnani, J. Collomosse, T. Bui, X. Liu, and S. Agarwal, “Promark: Proactive diffusion watermarking for causal attribution,” in
2024
Later among the works it cites.
A. Rezaei, M. Akbari, S. R. Alvar, A. Fatemi, and Y. Zhang, “Lawa: Using latent space for in-generation image watermarking,” in
2024
Later among the works it cites.
Z. Ma, G. Jia, B. Qi, and B. Zhou, “Safe-sd: Safe and traceable stable diffusion with text prompt trigger for invisible generative watermarking,” in
2024
Later among the works it cites.
W. Feng, W. Zhou, J. He, J. Zhang, T. Wei, G. Li, T. Zhang, W. Zhang, and N. Yu, “Aqualora: Toward white-box protection for customized stable diffusion models via watermark lora,” in
2024
Later among the works it cites.
Z. Wang, V. Sehwag, C. Chen, L. Lyu, D. N. Metaxas, and S. Ma, “How to trace latent generative model generated images without artificial watermark?” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Liu, Y. Zhou, X. Liu, T. Zhang, S. Liang, J. Wang, Y. Pu, T. Li, J. Zhang, W. Zhou
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Yang, X. Bi, Y. Lin, S. Chen, J. Zhou, and X. Sun, “Watch out for your agents! investigating backdoor threats to llm-based agents,” in
2024
Later among the works it cites.
Y. Wang, D. Xue, S. Zhang, and S. Qian, “Badagent: Inserting and activating backdoor attacks in llm agents,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Geng, S. Wang, D. Dong, H. Wang, G. Li, Z. Jin, X. Mao, and X. Liao, “Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
F. Wu, S. Wu, Y. Cao, and C. Xiao, “Wipi: A new web threat for llm-driven web agents,”
2024
Later among the works it cites.
X. Zhang, H. Xu, Z. Ba, Z. Wang, Y. Hong, J. Liu, Z. Qin, and K. Ren, “Privacyasst: Safeguarding user privacy in tool-using large language model agents,”
2024
Later among the works it cites.
Z. Xiang, L. Zheng, Y. Li, J. Hong, Q. Li, H. Xie, J. Zhang, Z. Xiong, C. Xie, C. Yang
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Zhang, “Attacking vision-language computer agents via pop-ups,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
D. Lee and M. Tiwari, “Prompt infection: Llm-to-llm prompt injection within multi-agent systems,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Gu, X. Zheng, T. Pang, C. Du, Q. Liu, Y. Wang, J. Jiang, and M. Lin, “Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Zeng, Y. Wu, X. Zhang, H. Wang, and Q. Wu, “Autodefense: Multi-agent LLM defense against jailbreak attacks,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Li, S. Zhao, Q. Wang, K. Wang, Y. Zhou, S. Srivastava, C. Gokmen, T. Lee, L. E. Li, R. Zhang, W. Liu, P. Liang, L. Fei-Fei, J. Mao, and J. Wu, “Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making,” in
2024
Later among the works it cites.
Rongwu Xu, Brian Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, and Han Qiu, “The earth is flat because…: Investigating llms’ belief towards misinformation via persuasive conversation,” in
2024
Later among the works it cites.
L. Wu, X. Yang, Y. Dong, L. Xie, H. Su, and J. Zhu, “Embodied Active Defense: Leveraging Recurrent Feedback to Counter Adversarial Patches,” in
2024
Later among the works it cites.
M. Shirasaka, T. Matsushima, S. Tsunashima, Y. Ikeda, A. Horo, S. Ikoma, C. Tsuji, H. Wada, T. Omija, D. Komukai, Y. Matsuo, and Y. Iwasawa, “Self-Recovery Prompting: Promptable General Purpose Service Robot System with Foundation Models and Self-Recovery,” in
2024
Later among the works it cites.
R. Fang, R. Bindu, A. Gupta, Q. Zhan, and D. Kang, “Llm agents can autonomously hack websites,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Wang, A. Zhang, D. T. Nguyen, J. Sun, T.-S. Chua
2024
Later among the works it cites.
Q. Zhan, Z. Liang, Z. Ying, and D. Kang, “Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents,” in
2024
Later among the works it cites.
E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramèr, “Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Guo, X. Liu, C. Xie, A. Zhou, Y. Zeng, Z. Lin, D. Song, and B. Li, “Redcode: Risky code execution and generation benchmark for code agents,” in
2024
Later among the works it cites.
A. to be confirmed], “Vpi-bench: Visual prompt injection attacks for computer-use agents,” 2024, based on available information from Moonlight review
2024
Later among the works it cites.
T. Yuan, Z. He, L. Dong, Y. Wang, R. Zhao, T. Xia, L. Xu, B. Zhou, F. Li, Z. Zhang
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Shao, T. Li, W. Shi, Y. Liu, and D. Yang, “Privacylens: Evaluating privacy norm awareness of language models in action,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
T. Cui, Y. Wang, C. Fu, Y. Xiao, S. Li, X. Deng, Y. Liu, Q. Zhang, Z. Qiu, P. Li
2024
Later among the works it cites.
Y. Gan, Y. Yang, Z. Ma, P. He, R. Zeng, Y. Wang, Q. Li, C. Zhou, S. Li, T. Wang
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
OpenAI, “Introducing openai o1,” 2024, openAI Blog
2024
Later among the works it cites.
2024
Later among the works it cites.
R. Xu, Z. Zhou, T. Zhang, Z. Qi, S. Yao, K. Xu, W. Xu, and H. Qiu, “Walking in others’ shoes: How perspective-taking guides large language models in reducing toxicity and bias,” in
2024
Later among the works it cites.
Y. Wang, H. Li, X. Han, P. Nakov, and T. Baldwin, “Do-not-answer: A dataset for evaluating safeguards in llms,” in
2024
Later among the works it cites.
K. Huang, X. Liu, Q. Guo, T. Sun, J. Sun, Y. Wang, Z. Zhou, Y. Wang, Y. Teng, X. Qiu
2024
Later among the works it cites.
T. Xie, X. Qi, Y. Zeng, Y. Huang, U. M. Sehwag, K. Huang, L. He, B. Wei, D. Li, Y. Sheng
2024
Later among the works it cites.
Z. Zhang, L. Lei, L. Wu, R. Sun, Y. Huang, C. Long, X. Liu, X. Lei, J. Tang, and M. Huang, “Safetybench: Evaluating the safety of large language models,” in
2024
Later among the works it cites.
L. Li, B. Dong, R. Wang, X. Hu, W. Zuo, D. Lin, Y. Qiao, and J. Shao, “Salad-bench: A hierarchical and comprehensive safety benchmark for large language models,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Luo, S. Ma, X. Liu, X. Guo, and C. Xiao, “Jailbreakv-28k: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks,” in
2024
Later among the works it cites.
A. Souly, Q. Lu, D. Bowen, T. Trinh, E. Hsieh, S. Pandey, P. Abbeel, J. Svegliato, S. Emmons, O. Watkins, and S. Toyer, “A strongREJECT for empty jailbreaks,” in
2024
Later among the works it cites.
H. Li, X. Han, Z. Zhai, H. Mu, H. Wang, Z. Zhang, Y. Geng, S. Lin, R. Wang, A. Shelmanov
2024
Later among the works it cites.
Z. Cheng, X. Wu, J. Yu, S. Han, X.-Q. Cai, and X. Xing, “Soft-label integration for robust toxicity classification,” in
2024
Later among the works it cites.
N. Carlini, M. Jagielski, C. A. Choquette-Choo, D. Paleka, W. Pearce, H. Anderson, A. Terzis, K. Thomas, and F. Tramèr, “Poisoning web-scale training datasets is practical,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Tu, C. Cui, Z. Wang, Y. Zhou, B. Zhao, J. Han, W. Zhou, H. Yao, and C. Xie, “How many unicorns are in this image? a safety evaluation benchmark for vision llms,” in
2024
Later among the works it cites.
X. Liu, Y. Zhu, J. Gu, Y. Lan, C. Yang, and Y. Qiao, “Mm-safetybench: A benchmark for safety evaluation of multimodal large language models,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y.-L. Tsai, C.-Y. Hsu, C. Xie, C.-H. Lin, J. Y. Chen, B. Li, P.-Y. Chen, C.-M. Yu, and C.-Y. Huang, “Ring-a-bell! how reliable are concept removal methods for diffusion models?” in
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu
2024
Later among the works it cites.
Y. Wen, Y. Liu, C. Chen, and L. Lyu, “Detecting, explaining, and mitigating memorization in diffusion models,” in
2024
Later among the works it cites.
J. Ren, Y. Li, S. Zeng, H. Xu, L. Lyu, Y. Xing, and J. Tang, “Unveiling and mitigating memorization in text-to-image diffusion models through cross attention,” in
2024
Later among the works it cites.
F. Liu, H. Luo, Y. Li, P. Torr, and J. Gu, “Which model generated this image? a model-agnostic approach for origin attribution,” in
2024
Later among the works it cites.
Y. . Hu, Z. Jiang, M. m. Guo, and N. . Gong, “A transfer attack to image watermarks,”
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Li, R. Wang, M. Cheng, T. Zhou, and C.-J. Hsieh, “Drattack: Prompt decomposition and reconstruction makes powerful llm jailbreakers,” in
2024
Later among the works it cites.
Y. Ding, L. L. Zhang, C. Zhang, Y. Xu, N. Shang, J. Xu, F. Yang, and M. Yang, “Longrope: extending llm context window beyond 2 million tokens,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
L. M. Titus, “Does chatgpt have semantic understanding? a problem with the statistics-of-occurrence strategy,”
2024
Later among the works it cites.
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. Jordan, J. E. Gonzalez
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Zhao, Z. Li, Y. Li, Y. Zhang, and J. Sun, “Defending large language models against jailbreak attacks via layer-specific editing,” in
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Wen, K. Bi, W. Chen, J. Guo, and X. Cheng, “Evaluating implicit bias in large language models by attacking from a psychometric perspective,” in
2025
Closest in time.
Y. Huang, C. Wang, X. Jia, Q. Guo, F. Juefei-Xu, J. Zhang, G. Pu, and Y. Liu, “Semantic-guided prompt organization for universal goal hijacking against llms,” in
2025
Closest in time.
2025
Closest in time.
C. Liang, L. Shen, Y. Deng, X. Zhao, B. Liang, and K.-F. Wong, “Pearl: Towards permutation-resilient llms,” in
2025
Closest in time.
C. Liang, X. Han, L. Shen, J. Bai, and K.-F. Wong, “Vulnerability-aware alignment: Mitigating uneven forgetting in harmful fine-tuning,” in
2025
Closest in time.
C. Yung, H. M. Dolatabadi, S. Erfani, and C. Leckie, “Round trip translation defence against large language model jailbreaking attacks,” in
2025
Closest in time.
2025
Closest in time.
L. Gao, J. Geng, X. Zhang, P. Nakov, and X. Chen, “Shaping the safety boundaries: Understanding and defending against jailbreaks in large language models,” in
2025
Closest in time.
2025
Closest in time.
S. Chen, J. Piet, C. Sitawarin, and D. Wagner, “Struq: Defending against prompt injection with structured queries,”
2025
Closest in time.
2025
Closest in time.
B. Yi, T. Huang, S. Chen, T. Li, Z. Liu, Z. Chu, and Y. Li, “Probe before you talk: Towards black-box defense against backdoor unalignment for large language models,” in
2025
Closest in time.
A. Sheshadri, J. Hughes, J. Michael, A. Mallen, A. Jose, F. Roger
2025
Closest in time.
J. Dong, Z. Zhang, Q. Zhang, T. Zhang, H. Wang, H. Li, Q. Li, C. Zhang, K. Xu, and H. Qiu, “An engorgio prompt makes large language model babble on,” in
2025
Closest in time.
Q. Zhang, H. Qiu, D. Wang, Y. Li, T. Zhang, W. Zhu, H. Weng, L. Yan, and C. Zhang, “A benchmark for semantic sensitive information in llms outputs,” in
2025
Closest in time.
X. Wang, Z. Zhao, and M. Larson, “Typographic attacks in a multi-image setting,” in
2025
Closest in time.
H. Huang, S. Erfani, Y. Li, X. Ma, and J. Bailey, “X-transfer attacks: Towards super transferable adversarial attacks on clip,” in
2025
Closest in time.
X. Wang, K. Chen, J. Zhang, J. Chen, and X. Ma, “Tapt: Test-time adversarial prompt tuning for robust inference in vision-language models,” in
2025
Closest in time.
H. Huang, S. Erfani, Y. Li, X. Ma, and J. Bailey, “Detecting backdoor samples in contrastive language image pretraining,” in
2025
Closest in time.
Y. Gong, D. Ran, J. Liu, C. Wang, T. Cong, A. Wang, S. Duan, and X. Wang, “Figstep: Jailbreaking large vision-language models via typographic visual prompts,” in
2025
Closest in time.
R. Wang, J. Li, Y. Wang, B. Wang, X. Wang, Y. Teng, Y. Wang, X. Ma, and Y.-G. Jiang, “Ideator: Jailbreaking and benchmarking large vision-language models using themselves,”
2025
Closest in time.
Y. Zhao, X. Zheng, L. Luo, Y. Li, X. Ma, and Y.-G. Jiang, “Bluesuffix: Reinforced blue teaming for vision-language models against jailbreak attacks,” in
2025
Closest in time.
Y. Huang, L. Liang, T. Li, X. Jia, R. Wang, W. Miao, G. Pu, and Y. Liu, “Perception-guided jailbreak against text-to-image models,” in
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Guan, M. Hu, S. Li, and A. K. Vullikanti, “Ufid: A unified framework for black-box input-level backdoor detection on diffusion models,” in
2025
Closest in time.
2025
Closest in time.
Q. Zhan, R. Fang, H. S. Panchal, and D. Kang, “Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents,” in
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Jiang, M. Li, G. Yang, J. Wang, Y. Huang, Z. Chang, and Q. Wang, “Mimicking the familiar: Dynamic command generation for information theft attacks in llm tool-learning system,” in
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
D. Kong, S. Lin, Z. Xu, Z. Wang, M. Li, Y. Li, Y. Zhang, Z. Sha, Y. Li, C. Lin
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
C. Chen, Z. Zhang, B. Guo, S. Ma, I. Khalilov, S. A. Gebreegziabher, Y. Ye, Z. Xiao, Y. Yao, T. Li
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
M. Standen, J. Kim, and C. Szabo, “Adversarial machine learning attacks and defences in multi-agent reinforcement learning,”
2025
Closest in time.
A. A. Fime, Z. Hossain, S. Zaman, A. R. Shahid, and A. Imteaj, “Towards Trustworthy Autonomous Vehicles with Vision-Language Models Under Targeted and Untargeted Adversarial Attacks,” in
2025
Closest in time.
2025
Closest in time.
H. Zhang, C. Zhu, X. Wang, Z. Zhou, C. Yin, M. Li, L. Xue, Y. Wang, S. Hu, A. Liu, P. Guo, and L. Y. Zhang, “BadRobot: Jailbreaking Embodied LLMs in the Physical World,” in
2025
Closest in time.
2025
Closest in time.
R. Jiao, S. Xie, J. Yue, T. Sato, L. Wang, Y. Wang, Q. A. Chen, and Q. Zhu, “Can We Trust Embodied Agents? Exploring Backdoor Attacks Against Embodied LLM-Based Decision-Making Systems,” in
2025
Closest in time.
T. Tomilin, M. Fang, and M. Pechenizkiy, “HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents,” in
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Chen, M. Kang, and B. Li, “Shieldagent: Shielding agents via verifiable safety policy reasoning,”
2025
Closest in time.
Z. Cai, “Aegisllm: Scaling agentic systems for self-reflective defense in large language models,”
2025
Closest in time.
S. Barua, “Guardians of the agentic system: Preventing many shot jailbreaking with agentic system,”
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
K. Wang, G. Zhang, Z. Zhou, J. Wu, M. Yu, S. Zhao, C. Yin, J. Fu, Y. Yan, H. Luo
2025
Closest in time.
Y. Ren, Z. Zhao, C. Lin, B. Yang, L. Zhou, Z. Liu, and C. Shen, “Improving adversarial transferability on vision transformers via forward propagation refinement,” in
2025
Closest in time.
N. Nikzad, Y. Liao, Y. Gao, and J. Zhou, “Sata: Spatial autocorrelation token analysis for enhancing the robustness of vision transformers,” in
2025
Closest in time.
J. Long, Z. Xu, T. Jiang, W. Yao, S. Jia, C. Ma, and X. Chen, “Robust sam: On the adversarial robustness of vision foundation models,” in
2025
Closest in time.
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi
2025
Closest in time.
H. Ge, Y. Li, Q. Wang, Y. Zhang, and R. Tang, “When backdoors speak: Understanding llm backdoor attacks through model-generated explanations,” in
2025
Closest in time.
Y. Li, S. Shao, Y. He, J. Guo, T. Zhang, Z. Qin, P.-Y. Chen, M. Backes, P. Torr, D. Tao
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Q. Zhou, D. Wang, T. Li, Y. Lin, Y. Liu, J. S. Dong, and Q. Guo, “Defending LVLMs against vision attacks through partial-perception supervision,” in
2025
Closest in time.
Y. Ding, B. Li, and R. Zhang, “ETA: Evaluating then aligning safety of vision language models at inference time,” in
2025
Closest in time.
2025
Closest in time.
Y. Yao, L. Li, J. Song, C. Chen, Z. He, Y. Wang, X. Wang, T. Gu, J. Li, Y. Teng
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
B. Li, Y. Wei, Y. Fu, Z. Wang, Y. Li, J. Zhang, R. Wang, and T. Zhang, “Towards reliable verification of unauthorized data usage in personalized text-to-image diffusion models,” in
2025
Closest in time.
R. Hu, J. Zhang, Y. Li, J. Li, Q. Guo, H. Qiu, and T. Zhang, “Videoshield: Regulating diffusion-based video generation models via watermarking,” in
2025
Closest in time.
Z. Wang, J. Guo, J. Zhu, Y. Li, H. Huang, M. Chen, and Z. Tu, “Sleepermark: Towards robust watermark against fine-tuning text-to-image diffusion models,” in
2025
Closest in time.
F. Xing, “Designing heterogeneous llm agents for financial sentiment analysis,”
2025
Closest in time.
2025
Closest in time.
LangChain, “Context engineering for agents,” 2025, langChain Blog
2025
Closest in time.
S. Zhang, M. Yin, J. Zhang, J. Liu, Z. Han, J. Zhang, B. Li, C. Wang, H. Wang, Y. Chen
2025
Closest in time.
X. Qi, A. Panda, K. Lyu, X. Ma, S. Roy, A. Beirami, P. Mittal, and P. Henderson, “Safety alignment should be made more than just a few tokens deep,” in
2025
Closest in time.