Fetching the paper…
Reading the bibliography…
Creating secure and resilient applications with large language models (LLM) requires anticipating, adjusting to, and countering unforeseen threats.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Fine-Tuning Language Models from Human Preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
Membership Inference Attacks From First Principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, A. Terzis, and Florian Tramèr · 1914
Earlier work this paper cites.
Simulation of Decisionmaking in Crises: Three Manual Gaming Experiments
Harvey A Averch and MM Lavin · 1964
Earlier work this paper cites.
Human error
James Reason · 1990
Earlier work this paper cites.
Red Teaming of advanced information assurance concepts
Ruth A. Duggan and Bradley Wood · 1999
Earlier work this paper cites.
SQLrand: Preventing SQL Injection Attacks
Stephen W. Boyd and Angelos Dennis Keromytis · 2004
Earlier work this paper cites.
Red Alert: Are you ready for every airport manager’s worst nightmare?
Jeff Price · 2004
Earlier work this paper cites.
Breaking Free: The Fight for User Control and the Practices of Jailbreaking
Devon C Fitzgerald · 2005
Earlier work this paper cites.
Side-channel attacks: Ten years after its publication and the impacts on cryptographic module security testing
YongBin Zhou and DengGuo Feng · 2005
Earlier work this paper cites.
A Classification of SQL-Injection Attacks and Countermeasures
William G. J. Halfond, Jeremy Viegas, and Alessandro Orso · 2006
Earlier work this paper cites.
A complete guide to the common vulnerability scoring system version 2.0, 2007-07-30 2007
Peter Mell, Karen Scarfone, and Sasha Romanosky · 2007
Earlier work this paper cites.
SQL Injection Attacks and Defense
Justin Clarke, Kevvie Fowler, Erlend Oftedal, Rodrigo Marcos Alvarez, David F. Hartley, Alexander Kornbrust, Gary O’leary-Steele, Alberto Revelli, Sumit Siddharth, and Marco Slaviero · 2009
Earlier work this paper cites.
Side channel attack-survey
G Joy Persial, M Prabhu, and R Shanmugalakshmi · 2011
Earlier work this paper cites.
Extracting Training Data from Large Language Models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel · 2012
Earlier work this paper cites.
Stochastic gradient descent with differentially private updates
Shuang Song, Kamalika Chaudhuri, and Anand D. Sarwate · 2013
Earlier work this paper cites.
Private Empirical Risk Minimization: Efficient Algorithms and Tight Error Bounds
Raef Bassily, Adam D. Smith, and Abhradeep Thakurta · 2014
Earlier work this paper cites.
The Algorithmic Foundations of Differential Privacy
Cynthia Dwork and Aaron Roth · 2014
Earlier work this paper cites.
Usable Security: History, Themes, and Challenges
Simson Garfinkel and Heather Richter Lipford · 2014
Earlier work this paper cites.
Generative Adversarial Nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Inverting Visual Representations with Convolutional Networks
Alexey Dosovitskiy and Thomas Brox · 2015
Earlier work this paper cites.
Model inversion attacks that exploit confidence information and basic countermeasures
Matt Fredrikson, Somesh Jha, and Thomas Ristenpart · 2015
Earlier work this paper cites.
Deep Learning with Differential Privacy
Martín Abadi, Andy Chu, Ian J. Goodfellow, H. B. McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Earlier work this paper cites.
Categorical Reparameterization with Gumbel-Softmax
Eric Jang, Shixiang Shane Gu, and Ben Poole · 2016
Earlier work this paper cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Earlier work this paper cites.
Practical Black-Box Attacks against Machine Learning
Nicolas Papernot, Patrick Mcdaniel, Ian J. Goodfellow, Somesh Jha, Z. Berkay Celik, and Ananthram Swami · 2016
Earlier work this paper cites.
Membership Inference Attacks Against Machine Learning Models
R. Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2016
Earlier work this paper cites.
Stealing Machine Learning Models via Prediction APIs
Florian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter, and Thomas Ristenpart · 2016
Earlier work this paper cites.
A methodology for formalizing model-inversion attacks
Xi Wu, Matthew Fredrikson, Somesh Jha, and Jeffrey F Naughton · 2016
Earlier work this paper cites.
BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg · 2017
Earlier work this paper cites.
Security and privacy controls for information systems and organizations
JointTaskForce · 2017
Earlier work this paper cites.
Towards Deep Learning Models Resistant to Adversarial Attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Towards Reverse-Engineering Black-Box Neural Networks
Seong Joon Oh, Maximilian Augustin, Mario Fritz, and Bernt Schiele · 2017
Earlier work this paper cites.
Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen · 2017
Earlier work this paper cites.
Privacy Risk in Machine Learning: Analyzing the Connection to Overfitting
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha · 2017
Earlier work this paper cites.
The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Xiaodong Song · 2018
Earlier work this paper cites.
Machine Learning and Security: Protecting Systems with Data and Algorithms
C. Chio and D. Freeman · 2018
Earlier work this paper cites.
Stealing Neural Networks via Timing Side Channels
Vasisht Duddu, Debasis Samanta, D. Vijay Rao, and Valentina Emilia Balas · 2018
Earlier work this paper cites.
Security Analysis of Deep Neural Networks Operating in the Presence of Cache Side-Channel Attacks
Sanghyun Hong, Michael Davinroy, Yigitcan Kaya, Stuart Nevans Locke, Ian Rackow, Kevin Kulda, Dana Dachman-Soled, and Tudor Dumitras · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
The art of cybersecurity: Defense in depth strategy for robust protection
Arif Ali Mughal · 2018
Earlier work this paper cites.
SoK: Security and Privacy in Machine Learning
Nicolas Papernot, Patrick Mcdaniel, Arunesh Sinha, and Michael P. Wellman · 2018
Earlier work this paper cites.
I Know What You See: Power Side-Channel Attack on Convolutional Neural Network Accelerators
Lingxiao Wei, Yannan Liu, Bo Luo, Yu LI, and Qiang Xu · 2018
Earlier work this paper cites.
CSI NN: Reverse Engineering of Neural Network Architectures Through Electromagnetic Side Channel
Lejla Batina, Shivam Bhasin, Dirmanto Jap, and Stjepan Picek · 2019
Earlier work this paper cites.
A Backdoor Attack Against LSTM-Based Text Classification Systems
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li · 2019
Earlier work this paper cites.
Neural network model extraction attacks in edge devices by hearing architectural hints
Xing Hu, Ling Liang, Lei Deng, Shuangchen Li, Xinfeng Xie, Yu Ji, Yufei Ding, Chang Liu, Timothy Sherwood, and Yuan Xie · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Earlier work this paper cites.
Open DNN Box by Power Side-Channel Attack
Yun Xiang, Zhuangzhi Chen, Zuohui Chen, Zebin Fang, Haiyang Hao, Jinyin Chen, Yi Liu, Zhefu Wu, Qi Xuan, and Xiaoniu Yang · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Label-Only Membership Inference Attacks
Christopher A. Choquette-Choo, Florian Tramèr, Nicholas Carlini, and Nicolas Papernot · 2020
Earlier work this paper cites.
Weight Poisoning Attacks on Pretrained Models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Earlier work this paper cites.
Backdoor Learning: A Survey
Yiming Li, Baoyuan Wu, Yong Jiang, Zhifeng Li, and Shutao Xia · 2020
Earlier work this paper cites.
Explainability matters: Backdoor attacks on medical imaging
Munachiso Nwadike, Takumi Miyawaki, Esha Sarkar, Michail Maniatakos, and Farah Shamout · 2020
Earlier work this paper cites.
ONION: A Simple and Effective Defense Against Textual Backdoor Attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Zhiyuan Liu, and Maosong Sun · 2020
Earlier work this paper cites.
Training Production Language Models without Memorizing User Data
Swaroop Indra Ramaswamy, Om Thakkar, Rajiv Mathews, Galen Andrew, H. B. McMahan, and Franccoise Beaufays · 2020
Earlier work this paper cites.
Information Leakage in Embedding Models
Congzheng Song and Ananth Raghunathan · 2020
Earlier work this paper cites.
Red Team Development and Operations–A practical Guide
Joe Vest and James Tubberville · 2020
Earlier work this paper cites.
Leaky DNN: Stealing Deep-Learning Model Secret with GPU Context-Switching Side-Channel
Junyin Wei, Yicheng Zhang, Zhe Zhou, Zhou Li, and Mohammad Abdullah Al Faruque · 2020
Earlier work this paper cites.
SAFER: A Structure-free Approach for Certified Robustness to Adversarial Word Substitutions
Mao Ye, Chengyue Gong, and Qiang Liu · 2020
Earlier work this paper cites.
Trojaning Language Models for Fun and Profit
Xinyang Zhang, Zheng Zhang, and Ting Wang · 2020
Earlier work this paper cites.
Large-Scale Differentially Private BERT
Rohan Anil, Badih Ghazi, Vineet Gupta, Ravi Kumar, and Pasin Manurangsi · 2021
Earlier work this paper cites.
A General Language Assistant as a Laboratory for Alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, T. J. Henighan, Andy Jones, Nicholas Joseph, Benjamin Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, John Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Christopher Olah, and Jared Kaplan · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Retiring adult: New datasets for fair machine learning
Frances Ding, Moritz Hardt, John Miller, and Ludwig Schmidt · 2021
Earlier work this paper cites.
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten Bosma, Zongwei Zhou, Tao Wang, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen S. Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc V. Le, Yonghui Wu, Z. Chen, and Claire Cui · 2021
Earlier work this paper cites.
Model Extraction and Adversarial Transferability, Your BERT is Vulnerable!
Xuanli He, L. Lyu, Qiongkai Xu, and Lichao Sun · 2021
Earlier work this paper cites.
Deduplicating Training Data Makes Language Models Better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Earlier work this paper cites.
Hidden Backdoors in Human-Centric Language Models
Shaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao, Minhui Xue, Haojin Zhu, and Jialiang Lu · 2021
Earlier work this paper cites.
Backdoor Pre-trained Models Can Transfer to All
Lujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li, Jing Chen, Jie Shi, Chengfang Fang, Jianwei Yin, and Ting Wang · 2021
Earlier work this paper cites.
Understanding the capabilities, limitations, and societal impact of large language models
Alex Tamkin, Miles Brundage, Jack Clark, and Deep Ganguli · 2021
Earlier work this paper cites.
Concealed Data Poisoning Attacks on NLP Models
Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh · 2021
Earlier work this paper cites.
Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and B. Li · 2021
Earlier work this paper cites.
Ethical and social risks of harm from Language Models
Laura Weidinger, John F. J. Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zachary Kenton, Sande Minnich Brown, William T. Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William S. Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Earlier work this paper cites.
Red Alarm for Pre-trained Models: Universal Vulnerability to Neuron-level Backdoor Attacks
Zhengyan Zhang, Guangxuan Xiao, Yongwei Li, Tian Lv, Fanchao Qi, Yasheng Wang, Xin Jiang, Zhiyuan Liu, and Maosong Sun · 2021
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, John Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, E Perez, Jamie Kerr, Jared Mueller, Jeff Ladish, J Landau, Kamal Ndousse, and et al · 2022
Earlier work this paper cites.
Reconstructing Training Data with Informed Adversaries
Borja Balle, Giovanni Cherubin, and Jamie Hayes · 2022
Earlier work this paper cites.
Spinning Coherent Interactive Fiction through Foundation Model Prompts
Alex Calderwood, Noah Wardrip-Fruin, and Michael Mateas · 2022
Earlier work this paper cites.
Quantifying Memorization Across Neural Language Models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang · 2022
Earlier work this paper cites.
Softmax bottleneck makes language models unable to represent multi-mode word distributions
Haw-Shiuan Chang and Andrew McCallum · 2022
Earlier work this paper cites.
RLPrompt: Optimizing discrete text prompts with reinforcement learning
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric Xing, and Zhiting Hu · 2022
Earlier work this paper cites.
Predictability and Surprise in Large Generative Models
Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, Sheer El Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Scott Johnston, Andy Jones, Nicholas Joseph, Jackson Kernian, Shauna Kravec, Ben Mann, Neel Nanda, Kamal Ndousse, Catherine Olsson, Daniela Amodei, Tom Brown, Jared Kaplan, Sam McCandlish, Christopher Olah, Dario Amodei, and Jack Clark · 2022
Earlier work this paper cites.
Google Bard
Google · 2022
Earlier work this paper cites.
Deduplicating Training Data Mitigates Privacy Risks in Language Models
Nikhil Kandpal, Eric Wallace, and Colin Raffel · 2022
Earlier work this paper cites.
A New Generation of Perspective API: Efficient Multilingual Character-level Transformers
Alyssa Lees, Vinh Q. Tran, Yi Tay, Jeffrey Scott Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman · 2022
Earlier work this paper cites.
Backdoor learning: A survey
Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia · 2022
Earlier work this paper cites.
Piccolo: Exposing Complex Backdoors in NLP Transformer Models
Yingqi Liu, Guangyu Shen, Guanhong Tao, Shengwei An, Shiqing Ma, and X. Zhang · 2022
Earlier work this paper cites.
Quantifying privacy risks of masked language models using membership inference attacks
Fatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, and Reza Shokri · 2022
Earlier work this paper cites.
NIST CVE
NIST CVE · 2022
Earlier work this paper cites.
Introducing ChatGPT
OpenAI · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike · 2022
Earlier work this paper cites.
A survey on backdoor attack and defense in natural language processing
Xuan Sheng, Zhaoyang Han, Piji Li, and Xiangmao Chang · 2022
Earlier work this paper cites.
AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model
Saleh Soltan, Shankar Ananthakrishnan, Jack FitzGerald, Rahul Gupta, Wael Hamza, Haidar Khan, Charith Peris, Stephen Rawls, Andy Rosenbaum, Anna Rumshisky, Chandana Satya Prakash, Mukund Sridhar, Fabian Triefenbach, Apurv Verma, Gökhan Tür, and Prem Natarajan · 2022
Earlier work this paper cites.
Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models
Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2022
Earlier work this paper cites.
Secure software development framework (ssdf) version 1.1
Murugiah Souppaya, Karen Scarfone, and Donna Dodson · 2022
Earlier work this paper cites.
Emergent Abilities of Large Language Models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Earlier work this paper cites.
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2022
Earlier work this paper cites.
One parameter defense—defending against data inference attacks via differential privacy
Dayong Ye, Sheng Shen, Tianqing Zhu, Bo Liu, and Wanlei Zhou · 2022
Earlier work this paper cites.
Certified robustness against natural language attacks by causal intervention
Haiteng Zhao, Chang Ma, Xinshuai Dong, Anh Tuan Luu, Zhi-Hong Deng, and Hanwang Zhang · 2022
Earlier work this paper cites.
Adversarial Training for High-Stakes Reliability
Daniel M. Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel Haas, Buck Shlegeris, and Nate Thomas · 2022
Earlier work this paper cites.
Adversarial Prompting in LLMs
Adversarial Prompting · 2023
Earlier work this paper cites.
AI Vulnerability Database
AI Vulnerability Database · 2023
Earlier work this paper cites.
Detecting Language Model Attacks with Perplexity
Gabriel Alon and Michael Kamfonas · 2023
Earlier work this paper cites.
ATLAS Matrix
ATLAS Matrix · 2023
Earlier work this paper cites.
Avid Taxonomy Matrix
Avid Taxonomy Matrix · 2023
Earlier work this paper cites.
MorseCode
Boaz Barak · 2023
Earlier work this paper cites.
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al · 2023
Earlier work this paper cites.
Model Leeching: An Extraction Attack Targeting LLMs
Lewis Birch, William Hackett, Stefan Trawicki, Neeraj Suri, and Peter Garraghan · 2023
Earlier work this paper cites.
AI Safety Summit
Bletchley Declaration · 2023
Earlier work this paper cites.
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
Yuanpu Cao, Bochuan Cao, and Jinghui Chen · 2023
Earlier work this paper cites.
Explore, Establish, Exploit: Red Teaming Language Models from Scratch
Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp, and Dylan Hadfield-Menell · 2023
Earlier work this paper cites.
AI Village at DEF CON announces largest-ever public Generative AI Red Team
Sven Cattell, Rumman Chowdhury, and Austin Carson · 2023
Earlier work this paper cites.
Jailbreaking Black Box Large Language Models in Twenty Queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Earlier work this paper cites.
Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity Even better
David Dale, Elena Voita, Loic Barrault, and Marta R. Costa-jussà · 2023
Earlier work this paper cites.
DAN is my new friend
DAN · 2023
Earlier work this paper cites.
Collaborating with language models for embodied reasoning
Ishita Dasgupta, Christine Kaeser-Chen, Kenneth Marino, Arun Ahuja, Sheila Babayan, Felix Hill, and Rob Fergus · 2023
Earlier work this paper cites.
Privacy Side Channels in Machine Learning Systems
Edoardo Debenedetti, Giorgio Severi, Nicholas Carlini, Christopher A. Choquette-Choo, Matthew Jagielski, Milad Nasr, Eric Wallace, and Florian Tramèr · 2023
Earlier work this paper cites.
MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu · 2023
Earlier work this paper cites.
How Ready are Pre-trained Abstractive Models and LLMs for Legal Case Judgement Summarization?
Aniket Deroy, Kripabandhu Ghosh, and Saptarshi Ghosh · 2023
Earlier work this paper cites.
Anthropomorphization of AI: Opportunities and Risks
A. Deshpande, Tanmay Rajpurohit, Karthik Narasimhan, and A. Kalyan · 2023
Earlier work this paper cites.
Chain-of-Verification Reduces Hallucination in Large Language Models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston · 2023
Earlier work this paper cites.
Sok: Model inversion attack landscape: Taxonomy, challenges, and future roadmap
Sayanton V Dibbo · 2023
Cited alongside, same era.
Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang · 2023
Cited alongside, same era.
Unleashing Cheapfakes through Trojan Plugins of Large Language Models
Tian Dong, Guoxing Chen, Shaofeng Li, Minhui Xue, Rayne Holland, Yan Meng, Zhen Liu, and Haojin Zhu · 2023
Cited alongside, same era.
Improving Factuality and Reasoning in Language Models through Multiagent Debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mordatch · 2023
Cited alongside, same era.
Google’s AI Red Team: the ethical hackers making AI safer
Daniel Fabian · 2023
Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing · 2023
Later among the works it cites.
Can Large Language Models Transform Computational Social Science?
Caleb Ziems, William B. Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
Sara Abdali, Richard Anarfi, C J Barberan, and Jia He · 2024
Closest in time.
Openai’s approach to external red teaming for ai models and systems
Lama Ahmad, Sandhini Agarwal, Michael Lampe, and Pamela Mishkin · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Emilio Ferrara · 2023
Cited alongside, same era.
Wenjie Fu, Huandong Wang, Chen Gao, Guanghua Liu, Yong Li, and Tao Jiang · 2023
Cited alongside, same era.
Retrieval-Augmented Generation for Large Language Models: A Survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang · 2023
Cited alongside, same era.
MART: Improving LLM Safety with Multi-round Automatic Red-Teaming
Suyu Ge, Chunting Zhou, Rui Hou, Madian Khabsa, Yi-Chia Wang, Qifan Wang, Jiawei Han, and Yuning Mao · 2023
Cited alongside, same era.
ChatGPT Perpetuates Gender Bias in Machine Translation and Ignores Non-Gendered Pronouns: Findings across Bengali and Five other Low-Resource Languages
Sourojit Ghosh and Aylin Caliskan · 2023
Cited alongside, same era.
LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?
David Glukhov, Ilia Shumailov, Yarin Gal, Nicolas Papernot, and Vardan Papyan · 2023
Cited alongside, same era.
Josh A. Goldstein, Girish Sastry, Micah Musser, Renée DiResta, Matthew Gentzel, and Katerina Sedova · 2023
Cited alongside, same era.
Closest in time.
Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, and Sara Hooker · 2024
Closest in time.
OrderBkd: Textual backdoor attack through repositioning
Irina Alekseevskaia and Konstantin Arkhipenko · 2024
Closest in time.
MultiAgent Collaboration Attack: Investigating Adversarial Attacks in Large Language Model Collaborations via Debate
Alfonso Amayuelas, Xianjun Yang, Antonis Antoniades, Wenyue Hua, Liangming Pan, and William Yang Wang · 2024
Closest in time.
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Closest in time.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic · 2024
Closest in time.
Refusal in Language Models Is Mediated by a Single Direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda · 2024
Closest in time.
Here’s a Free Lunch: Sanitizing Backdoored Models with Model Merge
Ansh Arora, Xuanli He, Maximilian Mozes, Srinibas Swain, Mark Dras, and Qiongkai Xu · 2024
Closest in time.
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
Anthony M. Barrett, Krystal Jackson, Evan R. Murphy, Nada Madkour, and Jessica Newman · 2024
Closest in time.
Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
Tim Baumgärtner, Yang Gao, Dana Alon, and Donald Metzler · 2024
Closest in time.
Diverse and effective red teaming with auto-generated rewards and multi-step reinforcement learning
Alex Beutel, Kai Xiao, Jo hannes Heidecke, and Lilian Weng · 2024
Closest in time.
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
Federico Bianchi and James Zou · 2024
Closest in time.
Bye Bye Bye…: Evolution of repeated token attacks on ChatGPT models
Mark Breitenbach and Adrian Wood · 2024
Closest in time.
Scalable AI safety via doubly-efficient debate
Jonah Brown-Cohen, Geoffrey Irving, and Georgios Piliouras · 2024
Closest in time.
Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation
Riccardo Cantini, Giada Cosenza, Alessio Orsino, and Domenico Talia · 2024
Closest in time.
Stealing Part of a Production Language Model
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Eric Wallace, David Rolnick, and Florian Tramèr · 2024
Closest in time.
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models, 2024
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramer, Hamed Hassani, and Eric Wong · 2024
Closest in time.
ChatGPTJailbreak
ChatGPTJailbreak · 2024
Closest in time.
Quantitative Certification of Bias in Large Language Models
Isha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi, Rahul Gupta, and Gagandeep Singh · 2024
Closest in time.
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
Yixin Cheng, Markos Georgopoulos, Volkan Cevher, and Grigorios G. Chrysos · 2024
Closest in time.
Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models
Arijit Ghosh Chowdhury, Md Mofijul Islam, Vaibhav Kumar, Faysal Hossain Shezan, Vinija Jain, and Aman Chadha · 2024
Closest in time.
Keras Vulnerability
CISA · 2024
Closest in time.
Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications
Stav Cohen, Ron Bitton, and Ben Nassi · 2024
Closest in time.
garak: A Framework for Security Probing Large Language Models
Leon Derczynski, Erick Galinkin, Jeffrey Martin, Subho Majumdar, and Nanna Inie · 2024
Closest in time.
Do Membership Inference Attacks Work on Large Language Models?
Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Closest in time.
Red-Teaming for Generative AI: Silver Bullet or Security Theater?
Michael Feffer, Anusha Sinha, Zachary Chase Lipton, and Hoda Heidari · 2024
Closest in time.
Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
Shanglun Feng and Florian Tramèr · 2024
Closest in time.
Logits of API-Protected LLMs Leak Proprietary Information
Matthew Finlayson, Xiang Ren, and Swabha Swayamdipta · 2024
Closest in time.
Coercing LLMs to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein · 2024
Closest in time.
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien · 2024
Closest in time.
giskard: The Evaluation & Testing framework for LLMs & ML models
Giskard · 2024
Closest in time.
Hugging Face, the GitHub of AI, hosted code that backdoored user devices
Dan Goodin · 2024
Closest in time.
Two Heads are Better than One: Nested PoE for Robust Defense Against Multi-Backdoors
Victoria Graf, Qin Liu, and Muhao Chen · 2024
Closest in time.
Alignment faking in large language models
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Samuel Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, and Evan Hubinger · 2024
Closest in time.
Accelerated Coordinate Gradient (ACG) attack method
HaizeLabs · 2024
Closest in time.
Haize Labs. A trivial jailbreak against LLama 3
HaizeLabs · 2024
Closest in time.
Red Teaming Resistance Benchmark
HaizeLabs · 2024
Closest in time.
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri · 2024
Closest in time.
Jailbreaking Proprietary Large Language Models using Word Substitution Cipher
Divij Handa, Advait Chirmule, Bimal Gajera, and Chitta Baral · 2024
Closest in time.
Risk and Response in Large Language Models: Evaluating Key Threat Categories
Bahareh Harandizadeh, Abel Salinas, and Fred Morstatter · 2024
Closest in time.
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
Adib Hasan, Ileana Rugina, and Alex Wang · 2024
Closest in time.
Query-Based Adversarial Prompt Generation
Jonathan Hayase, Ema Borevkovic, Nicholas Carlini, Florian Tramèr, and Milad Nasr · 2024
Closest in time.
Buffer Overflow in Mixture of Experts
Jamie Hayes, Ilia Shumailov, and Itay Yona · 2024
Closest in time.
What’s in Your"Safe"Data?: Identifying Benign Data that Breaks Safety
Luxi He, Mengzhou Xia, and Peter Henderson · 2024
Closest in time.
Defending Against Indirect Prompt Injection Attacks With Spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman · 2024
Closest in time.
Curiosity-driven Red-teaming for Large Language Models
Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang, Aldo Pareja, James Glass, Akash Srivastava, and Pulkit Agrawal · 2024
Closest in time.
Safe lora: the silver lining of reducing safety risks when fine-tuning large language models
Chia-Yi Hsu, Yu-Lin Tsai, Chih-Hsun Lin, Pin-Yu Chen, Chia-Mu Yu, and Chun-Ying Huang · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Evan Hubinger, Carson E. Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte Stuart MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Kristjanson Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Markus Brauner, Holden Karnofsky, Paul Francis Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez · 2024
Closest in time.
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
Jiabao Ji, Bairu Hou, Alexander Robey, George J Pappas, Hamed Hassani, Yang Zhang, Eric Wong, and Shiyu Chang · 2024
Closest in time.
DART: Deep Adversarial Automated Red Teaming for LLM Safety
Bojian Jiang, Yi Jing, Tianhao Shen, Qing Yang, and Deyi Xiong · 2024
Closest in time.
When Large Language Models Meet Vector Databases: A Survey
Zhi Jing, Yongye Su, Yikun Han, Bo Yuan, Haiyun Xu, Chunjiang Liu, Kehai Chen, and Min Zhang · 2024
Closest in time.
Injecting Undetectable Backdoors in Deep Learning and Language Models
Alkis Kalavasis, Amin Karbasi, Argyris Oikonomou, Katerina Sotiraki, Grigoris Velegkas, and Manolis Zampetakis · 2024
Closest in time.
C-RAG: certified generation risks for retrieval-augmented language models
Mintong Kang, Nezihe Merve Gürel, Ning Yu, Dawn Song, and Bo Li · 2024
Closest in time.
Understanding the Effects of Iterative Prompting on Truthfulness
Satyapriya Krishna, Chirag Agarwal, and Himabindu Lakkaraju · 2024
Closest in time.
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
Sander Land and Max Bartolo · 2024
Closest in time.
Using Hallucinations to Bypass GPT4’s Filter
Benjamin Lemkin · 2024
Closest in time.
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
Jiawei Liang, Siyuan Liang, Man Luo, Aishan Liu, Dongchen Han, Ee-Chien Chang, and Xiaochun Cao · 2024
Closest in time.
AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
Zeyi Liao and Huan Sun · 2024
Closest in time.
Against The Achilles’ Heel: A Survey on Red Teaming for Generative Models
Lizhi Lin, Honglin Mu, Zenan Zhai, Minghan Wang, Yuxia Wang, Renxi Wang, Junjie Gao, Yixuan Zhang, Wanxiang Che, Timothy Baldwin, Xudong Han, and Haonan Li · 2024
Closest in time.
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
Huijie Lv, Xiao Wang, Yuan Zhang, Caishuang Huang, Shihan Dou, Junjie Ye, Tao Gui, Qi Zhang, and Xuanjing Huang · 2024
Closest in time.
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
Neal Mangaokar, Ashish Hooda, Jihye Choi, Shreyas Chandrashekaran, Kassem Fawaz, Somesh Jha, and Atul Prakash · 2024
Closest in time.
Refusal in llms is an affine function
Thomas Marshall, Adam Scherlis, and Nora Belrose · 2024
Closest in time.
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Closest in time.
Jailbreaking Attack against Multimodal Large Language Model
Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin · 2024
Closest in time.
Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks
Dario Pasquini, Martin Strohmeier, and Carmela Troncoso · 2024
Closest in time.
Introducing the Chatbot Guardrails Arena
Sonali Pattnaik, Rohan Karan, Srijan Kumar, and Clémentine Fourrier · 2024
Closest in time.
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian · 2024
Closest in time.
Navigating the safety landscape: Measuring risks in finetuning large language models
Shengyun Peng, Pin-Yu Chen, Matthew Hull, and Duen Horng Chau · 2024
Closest in time.
Perspective API Toxicity Categories
Perspective API · 2024
Closest in time.
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson · 2024
Closest in time.
Learning to Poison Large Language Models During Instruction Tuning
Yao Qiang, Xiangyu Zhou, Saleh Zare Zade, Mohammad Amin Roshani, Douglas Zytko, and Dongxiao Zhu · 2024
Closest in time.
Sok: Prompt hacking of large language models
Baha Rababah, Shang Tommy Wu, Matthew Kwiatkowski, Carson K. Leung, and Cuneyt Gurcan Akcora · 2024
Closest in time.
GUARDIAN: A Multi-Tiered Defense Architecture for Thwarting Prompt Injection Attacks on LLMs
Parijat Rai, Saumil Sood, Vijay Krishna Madisetti, and Arshdeep Bahga · 2024
Closest in time.
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
Vyas Raina, Adian Liusie, and Mark Gales · 2024
Closest in time.
Safetywashing: Do ai safety benchmarks actually measure safety progress?, 2024
Richard Ren, Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, Alexander Pan, Gabriel Mukobi, Ryan H. Kim, Stephen Fitz, and Dan Hendrycks · 2024
Closest in time.
Towards Red Teaming in Multimodal and Multilingual Translation
Christophe Ropers, David Dale, Prangthip Hansanti, Gabriel Mejia Gonzalez, Ivan Evtimov, Corinne Wong, Christophe Touret, Kristina Pereyra, Seohyun Sonia Kim, Cristian Canton-Ferrer, Pierre Andrews, and Marta Ruiz Costa-jussà · 2024
Closest in time.
Mitigating Skeleton Key, a new type of generative AI jailbreak technique
Mark Russinovich · 2024
Closest in time.
Fast Adversarial Attacks on Language Models In One GPU Minute
Vinu Sankar Sadasivan, Shoumik Saha, Gaurang Sriramanan, Priyatham Kattakinda, Atoosa Malemir Chegini, and Soheil Feizi · 2024
Closest in time.
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktaschel, and Roberta Raileanu · 2024
Closest in time.
Pickle Scanning, Safetensors, Social validation features
Omar Sanseviero · 2024
Closest in time.
Bing Webmaster Guidelines
Barry Schwartz · 2024
Closest in time.
Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel, and Stephan Günnemann · 2024
Closest in time.
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
Guangyu Shen, Siyuan Cheng, Kai xian Zhang, Guanhong Tao, Shengwei An, Lu Yan, Zhuo Zhang, Shiqing Ma, and Xiangyu Zhang · 2024
Closest in time.
PAL: Proxy-Guided Black-Box Attack on Large Language Models
Chawin Sitawarin, Norman Mu, David Wagner, and Alexandre Araujo · 2024
Closest in time.
Steering Without Side Effects: Improving Post-Deployment Control of Language Models
Asa Cooper Stickland, Alexander Lyzhov, Jacob Pfau, Salsabila Mahdi, and Samuel R. Bowman · 2024
Closest in time.
TrustLLM: Trustworthiness in Large Language Models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zheng Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chun-Yan Li, Eric P. Xing, Furong Huang, Haodong Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang, Mohit Bansal, James Zou, Jian Pei, Jian Liu, Jianfeng Gao, Jiawei Han, Jieyu Zhao, Jiliang Tang, Jindong Wang, John Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang, Lifang He, Lifu Huang, Michael Backes, Neil Zhenqiang Gong, Philip S. Yu, Pin-Yu Chen, Quanquan Gu, Ran Xu, Rex Ying, Shuiwang Ji, Suman Sekhar Jana, Tian-Xiang Chen, Tianming Liu, Tianying Zhou, William Wang, Xiang Li, Xiang-Yu Zhang, Xiao Wang, Xingyao Xie, Xun Chen, Xuyu Wang, Yan Liu, Yanfang Ye, Yinzhi Cao, and Yue Zhao · 2024
Closest in time.
Xuchen Suo · 2024
Closest in time.
Prioritizing Safeguarding Over Autonomy: Risks of LLM Agents for Science
Xiangru Tang, Qiao Jin, Kunlun Zhu, Tongxin Yuan, Yichi Zhang, Wangchunshu Zhou, Meng Qu, Yilun Zhao, Jian Tang, Zhuosheng Zhang, Arman Cohan, Zhiyong Lu, and Mark B. Gerstein · 2024
Closest in time.
ALERT: A Comprehensive Benchmark for Assessing Large Language Models’ Safety through Red Teaming
Simone Tedeschi, Felix Friedrich, Patrick Schramowski, Kristian Kersting, Roberto Navigli, Huu Nguyen, and Bo Li · 2024
Closest in time.
Guardrails AI: Adding guardrails to large language models
Guardrails AI · 2024
Closest in time.
Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell · 2024
Closest in time.
Preventing jailbreak prompts as malicious tools for cybercriminals: A cyber defense perspective
Jean Marie Tshimula, Xavier Ndona, D’Jeff K. Nkashama, Pierre-Martin Tardif, Froduald Kabanza, Marc Frappier, and Shengrui Wang · 2024
Closest in time.
Hacc-Man: An Arcade Game for Jailbreaking LLMs
Matheus Valentim, Jeanette Falk, and Nanna Inie · 2024
Closest in time.
Introducing v0.5 of the AI Safety Benchmark from MLCommons
Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed, Victor Akinwande, Namir Al-nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Hamelin, Kurt D. Bollacker, Rishi Bomassani, Marisa Ferrara Boston, Sim’eon Campos, Kal Chakra, Canyu Chen, Cody Coleman, Zacharie Delpierre Coudert, Leon Derczynski, Debojyoti Dutta, Ian Eisenberg, James R. Ezick, Heather Frase, Brian Fuller, Ram Gandikota, Agasthya Gangavarapu, Ananya Gangavarapu, James Gealy, Rajat Ghosh, James Goel, Usman Gohar, Sujata Goswami, Scott A. Hale, Wiebke Hutiri, Joseph Marvin Imperial, Surgan Jandial, Nicholas C. Judd, Felix Juefei-Xu, Foutse Khomh, Bhavya Kailkhura, Hannah Rose Kirk, Kevin Klyman, Chris Knotz, Michael Kuchnik, Shachi H. Kumar, Chris Lengerich, Bo Li, Zeyi Liao, Eileen Peters Long, Victor Lu, Yifan Mai, Priyanka Mary Mammen, Kelvin Manyeki, Sean McGregor, Virendra Mehta, Shafee Mohammed, Emanuel Moss, Lama Nachman, Dinesh Jinenhally Naganna, Amin Nikanjam, Besmira Nushi, Luis Oala, Iftach Orr, Alicia Parrish, Çigdem Patlak, William Pietri, Forough Poursabzi-Sangdeh, Eleonora Presani, Fabrizio Puletti, Paul Röttger, Saurav Sahay, Tim Santos, Nino Scherrer, Alice Schoenauer Sebag, Patrick Schramowski, Abolfazl Shahbazi, Vin Sharma, Xudong Shen, Vamsi Sistla, Leonard Tang, Davide Testuggine, Vithursan Thangarasa, Elizabeth Anne Watkins, Rebecca Weiss, Christoper A. Welty, Tyler Wilbers, Adina Williams, Carole-Jean Wu, Poonam Yadav, Xianjun Yang, Yi Zeng, Wenhui Zhang, Fedor Zhdanov, Jiacheng Zhu, Percy Liang, Peter Mattson, and Joaquin Vanschoren · 2024
Closest in time.
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Eric Wallace, Kai Xiao, Reimar H. Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel · 2024
Closest in time.
STAR: SocioTechnical approach to red teaming language models
Laura Weidinger, John F J Mellor, Bernat Guillén Pegueroles, Nahema Marchal, Ravin Kumar, Kristian Lum, Canfer Akbulut, Mark Diaz, A. Stevie Bergman, Mikel D. Rodriguez, Verena Rieser, and William Isaac · 2024
Closest in time.
What Was Your Prompt? A Remote Keylogging Attack on AI Assistants
Roy Weiss, Daniel Ayzenshteyn, Guy Amit, and Yisroel Mirsky · 2024
Closest in time.
Gradient-Based Language Model Red Teaming
Nevan Wichers, Carson E. Denison, and Ahmad Beirami · 2024
Closest in time.
Tastle: Distract Large Language Models for Automatic Jailbreak Attack
Zeguan Xiao, Yan Yang, Guanhua Chen, and Yun Chen · 2024
Closest in time.
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Guo Xingang, Fangxu Yu, Huan Zhang, Lianhui Qin, and Bin Hu · 2024
Closest in time.
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
Chen Xiong, Xiangyu Qi, Pin-Yu Chen, and Tsung-Yi Ho · 2024
Closest in time.
LinkPrompt: Natural and Universal Adversarial Attacks on Prompt-based Language Models
Yue Xu and Wenjie Wang · 2024
Closest in time.
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun · 2024
Closest in time.
A safety realignment framework via subspace-oriented model fusion for large language models
Xin Yi, Shunfan Zheng, Linlin Wang, Xiaoling Wang, and Liang He · 2024
Closest in time.
Stealing user prompts from mixture of experts
Itay Yona, Ilia Shumailov, Jamie Hayes, and Nicholas Carlini · 2024
Closest in time.
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
Zhuowen Yuan, Zidi Xiong, Yi Zeng, Ning Yu, Ruoxi Jia, Dawn Xiaodong Song, and Bo Li · 2024
Closest in time.
Round Trip Translation Defence against Large Language Model Jailbreaking Attacks
Canaan Yung, Hadi Mohaghegh Dolatabadi, Sarah Monazam Erfani, and Christopher Leckie · 2024
Closest in time.
The Shift from Models to Compound AI Systems
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi · 2024
Closest in time.
Ai risk categorization decoded (air 2024): From government regulations to corporate policies
Yi Zeng, Kevin Klyman, Andy Zhou, Yu Yang, Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, and Bo Li · 2024
Closest in time.
Boosting Jailbreak Attack with Momentum
Yihao Zhang and Zeming Wei · 2024
Closest in time.
Generative AI Security: Challenges and Countermeasures
Banghua Zhu, Norman Mu, Jiantao Jiao, and David Wagner · 2024
Closest in time.
PoisonedRAG: Knowledge Poisoning Attacks to Retrieval-Augmented Generation of Large Language Models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia · 2024
Closest in time.
Adversarial tokenization
Renato Lui Geh, Zilei Shao, and Guy Van den Broeck · 2025
Closest in time.
Auditing prompt caching in language model apis
Chenchen Gu, Xiang Lisa Li, Rohith Kuditipudi, Percy Liang, and Tatsunori Hashimoto · 2025
Closest in time.
Virus: Harmful fine-tuning attack for large language models bypassing guardrail moderation
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, and Ling Liu · 2025
Closest in time.
2025 Top 10 Risk and Mitigations for LLMs and Gen AI Apps
OWASP · 2025
Closest in time.
A Few Useful Lessons about AI Red Teaming
Ram Shankar Siva Kumar · 2025
Closest in time.
Compromising honesty and harmlessness in language models via deception attacks
Laurene Vaugrante, Francesca Carlon, Maluna Menke, and Thilo Hagendorff · 2025
Closest in time.