Fetching the paper…
Reading the bibliography…
As Large Language Models (LLMs) increasingly become key components in various AI applications, understanding their security vulnerabilities and the effectiveness of defense mechanisms is crucial.
hdbscan: Hierarchical density based clustering
Leland McInnes, John Healy, Steve Astels, et al · 2017
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Adversarial attacks and defenses in deep learning
Kui Ren, Tianhang Zheng, Zhan Qin, and Xue Liu · 2020
Earlier work this paper cites.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li · 2020
Earlier work this paper cites.
Backdoor attacks on pre-trained models by layerwise weight poisoning
Linyang Li, Demin Song, Xiaonan Li, Jiehang Zeng, Ruotian Ma, and Xipeng Qiu · 2021
Earlier work this paper cites.
Onion: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun · 2021
Earlier work this paper cites.
Concealed data poisoning attacks on nlp models
Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh · 2021
Earlier work this paper cites.
Rap: Robustness-aware perturbations for defending against backdoor attacks on nlp models
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, and Xu Sun · 2021
Earlier work this paper cites.
Trojaning language models for fun and profit
Xinyang Zhang, Zheng Zhang, Shouling Ji, and Ting Wang · 2021
Earlier work this paper cites.
A unified evaluation of textual backdoor learning: Frameworks and benchmarks
Ganqu Cui, Lifan Yuan, Bingxiang He, Yangyi Chen, Zhiyuan Liu, and Maosong Sun · 2022
Earlier work this paper cites.
Demystifying prompts in language models via perplexity estimation
Hila Gonen, Srini Iyer, Terra Blevins, Noah A Smith, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
Fine-mixing: Mitigating backdoors in fine-tuned language models
Zhiyuan Zhang, Lingjuan Lyu, Xingjun Ma, Chenguang Wang, and Xu Sun · 2022
Earlier work this paper cites.
Ai risk skepticism, a comprehensive survey
Vemir Michael Ambartsoumean and Roman V Yampolskiy · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Machine-generated text: A comprehensive survey of threat models and detection methods
Evan Crothers, Nathalie Japkowicz, and Herna L Viktor · 2023
Earlier work this paper cites.
Jailbreaker: Automated jailbreak across multiple large language model chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2023
Cited alongside, same era.
Composite backdoor attacks against large language models
Hai Huang, Zhengyu Zhao, Michael Backes, Yun Shen, and Yang Zhang · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Prompt packer: Deceiving llms through compositional instruction with hidden attacks
Shuyu Jiang, Xingshu Chen, and Rui Tang · 2023
Cited alongside, same era.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu · 2023
Later among the works it cites.
Instruction tuning for large language models: A survey
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al · 2023
Later among the works it cites.
Prompt as triggers for backdoor attack: Examining the vulnerability in language models
Shuai Zhao, Jinming Wen, Anh Luu, Junbo Zhao, and Jie Fu · 2023
Later among the works it cites.
Trojanpuzzle: Covertly poisoning code-suggestion models
Hojjat Aghakhani, Wei Dai, Andre Manoel, Xavier Fernandes, Anant Kharkar, Christopher Kruegel, Giovanni Vigna, David Evans, Ben Zorn, and Robert Sim · 2024
Closest in time.
Introducing the next generation of claude
Anthropic · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Cited alongside, same era.
Jailbreaking chatgpt via prompt engineering: An empirical study
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, Kailong Wang, and Yang Liu · 2023
Cited alongside, same era.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Cited alongside, same era.
Chatgpt, version 4
OpenAI · 2023
Cited alongside, same era.
Visual adversarial examples jailbreak large language models
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Mengdi Wang, and Prateek Mittal · 2023
Cited alongside, same era.
Smoothllm: Defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Badhan Chandra Das, M Hadi Amini, and Yanzhao Wu · 2024
Closest in time.
Propile: Probing privacy leakage in large language models
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh · 2024
Closest in time.
Prompt hacking: Jailbreaking, 2024
Learn Prompting · 2024
Closest in time.
Backdoor attacks and defenses in federated learning: Survey, challenges and future research directions
Thuy Dung Nguyen, Tuan Nguyen, Phi Le Nguyen, Hieu H Pham, Khoa D Doan, and Kok-Seng Wong · 2024
Closest in time.
Hello gpt-4o, 2024
OpenAI · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2024
Closest in time.
Defending pre-trained language models as few-shot learners against backdoor attacks
Zhaohan Xi, Tianyu Du, Changjiang Li, Ren Pang, Shouling Ji, Jinghui Chen, Fenglong Ma, and Ting Wang · 2024
Closest in time.
A comprehensive overview of backdoor attacks in large language models within communication networks
Haomiao Yang, Kunlan Xiang, Mengyu Ge, Hongwei Li, Rongxing Lu, and Shui Yu · 2024
Closest in time.