Fetching the paper…
Reading the bibliography…
This study sheds light on the imperative need to bolster safety and privacy measures in large language models (LLMs), such as GPT-4 and LLaMA-2, by identifying and mitigating their vulnerabilities through explainable analysis of prompt attacks.
Counterfactual thinking
Neal J Roese · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Black-box generation of adversarial text sequences to evade deep learning classifiers
Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi · 2018
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang · 2018
Earlier work this paper cites.
BERT-ATTACK: Adversarial attack against BERT using BERT
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu · 2020
Earlier work this paper cites.
Counterfactual explanations and algorithmic recourses for machine learning: A review
Sahil Verma, Varich Boonsanong, Minh Hoang, Keegan E Hines, John P Dickerson, and Chirag Shah · 2020
Earlier work this paper cites.
A survey on adversarial attacks for malware analysis
Kshitiz Aryal, Maanak Gupta, and Mahmoud Abdelsalam · 2021
Earlier work this paper cites.
Using gpt-eliezer against chatgpt jailbreaking
Stuart Armstrong and R Gorman · 2022
Earlier work this paper cites.
Meta. introducing llama: A foundational, 65-billion-parameter large language model
Facebook · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models, 2022
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
Jailbreak chat
Alex Albert · 2023
Cited alongside, same era.
The dark side of explanations: Poisoning recommender systems with counterfactual examples
Ziheng Chen, Fabrizio Silvestri, Jia Wang, Yongfeng Zhang, and Gabriele Tolomei · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection, 2023
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
Pre-training to learn in context, 2023
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang · 2023
Cited alongside, same era.
Baseline defenses for adversarial attacks against aligned language models, 2023
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Later among the works it cites.
Stanford corenlp
Stanford NLP Group · 2023
Later among the works it cites.
Tensor trust: Interpretable prompt injection attacks from an online game
Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, et al · 2023
Later among the works it cites.
Aligning large language models with human: A survey
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?, 2023
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, and Yangqiu Song · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2023
Cited alongside, same era.
Prompt injection attacks and defenses in llm-integrated applications, 2023
Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong · 2023
Cited alongside, same era.
Behavioral analysis of cybercrime: Paving the way for effective policing strategies
Gargi Sarkar and Sandeep K Shukla · 2023
Cited alongside, same era.
Large language model alignment: A survey
Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong · 2023
Cited alongside, same era.
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Eric Sun, and Yue Zhang · 2023
Later among the works it cites.
Assessing prompt injection risks in 200+ custom gpts
Jiahao Yu, Yuhang Wu, Dong Shu, Mingyu Jin, and Xinyu Xing · 2023
Later among the works it cites.
Effective prompt extraction from language models
Yiming Zhang and Daphne Ippolito · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.