Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are shown to be vulnerable to jailbreaking attacks where adversarial prompts are designed to elicit harmful responses.
Control barrier function based quadratic programs with application to adaptive cruise control
Aaron D Ames, Jessy W Grizzle, and Paulo Tabuada · 2014
Earlier work this paper cites.
Control in a safe set: Addressing safety in human-robot interactions
Changliu Liu and Masayoshi Tomizuka · 2014
Earlier work this paper cites.
Control barrier functions: Theory and applications
Aaron D Ames, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada · 2019
Earlier work this paper cites.
Neural lyapunov control
Ya-Chien Chang, Nima Roohi, and Sicun Gao · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter · 2019
Earlier work this paper cites.
Bridging hamilton-jacobi safety analysis and reinforcement learning
Jaime F Fisac, Neil F Lugovoy, Vicenç Rubies-Royo, Shromona Ghosh, and Claire J Tomlin · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Learning control barrier functions from expert demonstrations
Alexander Robey, Haimin Hu, Lars Lindemann, Hanwen Zhang, Dimos V Dimarogonas, Stephen Tu, and Nikolai Matni · 2020
Earlier work this paper cites.
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2020
Earlier work this paper cites.
Learning stability certificates from data
Nicholas Boffi, Stephen Tu, Nikolai Matni, Jean-Jacques Slotine, and Vikas Sindhwani · 2021
Earlier work this paper cites.
Scalable learning of safety guarantees for autonomous systems using hamilton-jacobi reachability
Sylvia Herbert, Jason J Choi, Suvansh Sanjeev, Marsalis Gibson, Koushil Sreenath, and Claire J Tomlin · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Learning hybrid control barrier functions from data
Lars Lindemann, Haimin Hu, Alexander Robey, Hanwen Zhang, Dimos Dimarogonas, Stephen Tu, and Nikolai Matni · 2021
Earlier work this paper cites.
Learning a better control barrier function
Bolun Dai, Prashanth Krishnamurthy, and Farshad Khorrami · 2022
Earlier work this paper cites.
Safe nonlinear control using robust neural lyapunov-barrier functions
Charles Dawson, Zengyi Qin, Sicun Gao, and Chuchu Fan · 2022
Earlier work this paper cites.
Safety certification for stochastic systems via neural barrier functions
Frederik Baymler Mathiesen, Simeon C Calvert, and Luca Laurenti · 2022
Earlier work this paper cites.
Safety guarantees for neural network dynamic systems via stochastic barrier functions
Rayan Mazouz, Karan Muvvala, Akash Ratheesh Babu, Luca Laurenti, and Morteza Lahijanian · 2022
Earlier work this paper cites.
Safe control with neural network dynamic models
Tianhao Wei and Changliu Liu · 2022
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al · 2023
Earlier work this paper cites.
System-level safety guard: Safe tracking control through uncertain neural network dynamics models
Xiao Li, Yutong Li, Anouck Girard, and Ilya Kolmanovsky · 2023
Earlier work this paper cites.
Prompt template: Llama-2-chat, 2023
Llama-2-Chat · 2023
Earlier work this paper cites.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Earlier work this paper cites.
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi · 2023
Earlier work this paper cites.
Representation biases in sentence transformers
Dmitry Nikolaev and Sebastian Padó · 2023
Earlier work this paper cites.
Gpt-3.5 turbo, 2023
OpenAI · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Earlier work this paper cites.
Smoothllm: Defending large language models against jailbreaking attacks
Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas · 2023
Earlier work this paper cites.
Oswin So, Zachary Serlin, Makai Mann, Jake Gonzales, Kwesi Rutledge, Nicholas Roy, and Chuchu Fan · 2023
Earlier work this paper cites.
Taming ai bots: Controllability of neural states in large language models
Stefano Soatto, Paulo Tabuada, Pratik Chaudhari, and Tian Yu Liu · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Xinyu Wang, Luzia Knoedler, Frederik Baymler Mathiesen, and Javier Alonso-Mora · 2023
Cited alongside, same era.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Cited alongside, same era.
Barriernet: Differentiable control barrier functions for learning of safe robot control
Automated red teaming with goat: the generative offensive agent tester
Maya Pavlova, Erik Brinkman, Krithika Iyer, Vitor Albiero, Joanna Bitton, Hailey Nguyen, Joe Li, Cristian Canton Ferrer, Ivan Evtimov, and Aaron Grattafiori · 2024
Later among the works it cites.
Derail yourself: Multi-turn llm jailbreak attack through self-discovered clues
Qibing Ren, Hao Li, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, and Jing Shao · 2024
Later among the works it cites.
Jailbreaking llm-controlled robots
Alexander Robey, Zachary Ravichandran, Vijay Kumar, Hamed Hassani, and George J Pappas · 2024
Later among the works it cites.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei Xiao, Tsun-Hsuan Wang, Ramin Hasani, Makram Chahine, Alexander Amini, Xiao Li, and Daniela Rus · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Cited alongside, same era.
Vrushabh Zinage, Rohan Chandra, and Efstathios Bakolas · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al · 2024
Cited alongside, same era.
Automatic pseudo-harmful prompt generation for evaluating false refusals in large language models
Bang An, Sicheng Zhu, Ruiyi Zhang, Michael-Andrei Panaitescu-Liess, Yuancheng Xu, and Furong Huang · 2024
Cited alongside, same era.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Cited alongside, same era.
Claude-3.5-sonnet, 2024
Anthropic · 2024
Cited alongside, same era.
Mark Russinovich, Ahmed Salem, and Ronen Eldan · 2024
Later among the works it cites.
Bab-nd: Long-horizon motion planning with branch-and-bound and neural dynamics
Keyi Shen, Jiangwei Yu, Huan Zhang, and Yunzhu Li · 2024
Later among the works it cites.
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al · 2024
Later among the works it cites.
Securing multi-turn conversational language models from distributed backdoor attacks
Terry Tong, Qin Liu, Jiashu Xu, and Muhao Chen · 2024
Later among the works it cites.
Mrj-agent: An effective jailbreak agent for multi-round dialogue
Fengxiang Wang, Ranjie Duan, Peng Xiao, Xiaojun Jia, YueFeng Chen, Chongwen Wang, Jialing Tao, Hang Su, Jun Zhu, and Hui Xue · 2024
Later among the works it cites.
Chain of attack: a semantic-driven contextual multi-turn attacker for llm
Xikang Yang, Xuehai Tang, Songlin Hu, and Jizhong Han · 2024
Later among the works it cites.
Cosafe: Evaluating large language model safety in multi-turn dialogue coreference
Erxin Yu, Jing Li, Ming Liao, Siqi Wang, Zuchen Gao, Fei Mi, and Lanqing Hong · 2024
Later among the works it cites.
Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Jiahao Xu, Tian Liang, Pinjia He, and Zhaopeng Tu · 2024
Later among the works it cites.
Shieldgemma: Generative ai content moderation based on gemma
Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, et al · 2024
Later among the works it cites.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma · 2024
Later among the works it cites.
Speak out of turn: Safety vulnerability of large language models in multi-turn dialogue
Zhenhong Zhou, Jiuyang Xiang, Haopeng Chen, Quan Liu, Zherui Li, and Sen Su · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, J Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Later among the works it cites.
Claude sonnet 4.5 system card, 2025
Anthropic · 2025
Closest in time.
Conformal information pursuit for interactively guiding large language models
Kwan Ho Ryan Chan, Yuyan Ge, Edgar Dobriban, Hamed Hassani, and René Vidal · 2025
Closest in time.
Safeguarding large language models in real-time with tunable safety-performance trade-offs
Joao Fonseca, Andrew Bell, and Julia Stoyanovich · 2025
Closest in time.
One-shot is enough: Consolidating multi-turn attacks into efficient single-turn prompts for llms
Junwoo Ha, Hyunjun Kim, Sangyoon Yu, Haon Park, Ashkan Yousefpour, Yuna Park, and Suhyun Kim · 2025
Closest in time.
Verifiable safety q-filters via hamilton-jacobi reachability and multiplicative q-networks
Jiaxing Li, Hanjiang Hu, Yujie Yang, and Changliu Liu · 2025
Closest in time.
Guardreasoner: Towards reasoning-based llm safeguards
Yue Liu, Hongcheng Gao, Shengfang Zhai, Jun Xia, Tianyi Wu, Zhiwei Xue, Yulin Chen, Kenji Kawaguchi, Jiaheng Zhang, and Bryan Hooi · 2025
Closest in time.
Xiaoya Lu, Dongrui Liu, Yi Yu, Luxin Xu, and Jing Shao · 2025
Closest in time.
Usage policies, 2022
OpenAI · 2025
Closest in time.
Gpt-5 system card, 2025
OpenAI · 2025
Closest in time.
Safety alignment should be made more than just a few tokens deep
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson · 2025
Closest in time.
Data-driven neural certificate synthesis
Luke Rickard, Alessandro Abate, and Kostas Margellos · 2025
Closest in time.
Barrierbench: Evaluating large language models for safety verification in dynamical systems
Ali Taheri, Alireza Taban, Sadegh Soudjani, and Ashutosh Trivedi · 2025
Closest in time.
Online adaptive probabilistic safety certificate with language guidance
Zhuoyuan Wang, Xiyu Deng, Hikaru Hoshino, and Yorie Nakahira · 2025
Closest in time.
Trading inference-time compute for adversarial robustness
Wojciech Zaremba, Evgenia Nitishinskaya, Boaz Barak, Stephanie Lin, Sam Toyer, Yaodong Yu, Rachel Dias, Eric Wallace, Kai Xiao, Johannes Heidecke, et al · 2025
Closest in time.
Safety is not only about refusal: Reasoning-enhanced fine-tuning for interpretable llm safety
Yuyou Zhang, Miao Li, William Han, Yihang Yao, Zhepeng Cen, and Ding Zhao · 2025
Closest in time.