Fetching the paper…
Reading the bibliography…
Multi-turn jailbreak attacks simulate real-world human interactions by engaging large language models (LLMs) in iterative dialogues, exposing critical safety vulnerabilities.
A mathematical theory of communication
Claude Elwood Shannon · 1948
Earlier work this paper cites.
Equilibrium points in n-person games
John F Nash Jr · 1950
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
Arthur L Samuel · 1959
Earlier work this paper cites.
Reasoning about a rule
Peter C Wason · 1968
Earlier work this paper cites.
Psychology of Reasoning: Structure and Content
PC Wason · 1972
Earlier work this paper cites.
Information gain and a general measure of correlation
John T Kent · 1983
Earlier work this paper cites.
Finite state machine based formal methods in protocol conformance testing: from theory to implementation
Barry S Bosik and M Ümit Uyar · 1991
Earlier work this paper cites.
Divergence measures based on the shannon entropy
Jianhua Lin · 1991
Earlier work this paper cites.
Testing finite-state machines: State identification and verification
David Lee and Mihalis Yannakakis · 1994
Earlier work this paper cites.
Introduction to the theory of computation
Michael Sipser · 1996
Earlier work this paper cites.
Introduction to automata theory, languages, and computation
John E Hopcroft, Rajeev Motwani, and Jeffrey D Ullman · 2001
Earlier work this paper cites.
Finding useful questions: on bayesian diagnosticity, probability, impact, and information gain
Jonathan D Nelson · 2005
Earlier work this paper cites.
Finite state machines in hardware: theory and design (with VHDL and SystemVerilog)
Volnei A Pedroni · 2013
Earlier work this paper cites.
An intelligent agent of finite state machine in educational game “flora the explorer”
AF Pukeng, RR Fauzi, R Andrea, E Yulsilviana, S Mallala, et al · 2019
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Earlier work this paper cites.
Maieutic prompting: Logically consistent reasoning with recursive explanations
Jaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman, Chandra Bhagavatula, Ronan Le Bras, and Yejin Choi · 2022
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Earlier work this paper cites.
Chatgpt for good? on opportunities and challenges of large language models for education
Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al · 2023
Earlier work this paper cites.
Large language models in medicine
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting · 2023
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson · 2023
Earlier work this paper cites.
Nba: defensive distillation for backdoor removal via neural behavior alignment
Zonghao Ying and Bin Wu · 2023
Earlier work this paper cites.
Dlp: towards active defense against backdoor attacks with decoupled learning process
Zonghao Ying and Bin Wu · 2023
Earlier work this paper cites.
Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning
Siyuan Liang, Mingli Zhu, Aishan Liu, Baoyuan Wu, Xiaochun Cao, and Ee-Chien Chang · 2023
Earlier work this paper cites.
Open sesame! universal black box jailbreaking of large language models
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Earlier work this paper cites.
On the humanity of conversational ai: Evaluating the psychological portrayal of llms
Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael Lyu · 2023
Cited alongside, same era.
Cladder: Assessing causal reasoning in language models
Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, LYU Zhiheng, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman-Weiner, Mrinmaya Sachan, et al · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Cited alongside, same era.
Look before you leap: An exploratory study of uncertainty measurement for large language models
Yuheng Huang, Jiayang Song, Zhijie Wang, Shengming Zhao, Huaming Chen, Felix Juefei-Xu, and Lei Ma · 2023
Leveraging the context through multi-round interactions for jailbreaking attacks, 2024
Yixin Cheng, Markos Georgopoulos, Volkan Cevher, and Grigorios G. Chrysos · 2024
Later among the works it cites.
Mrj-agent: An effective jailbreak agent for multi-round dialogue, 2024
Fengxiang Wang, Ranjie Duan, Peng Xiao, Xiaojun Jia, YueFeng Chen, Chongwen Wang, Jialing Tao, Hang Su, Jun Zhu, and Hui Xue · 2024
Later among the works it cites.
Derail yourself: Multi-turn llm jailbreak attack through self-discovered clues, 2024
Qibing Ren, Hao Li, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, and Jing Shao · 2024
Later among the works it cites.
Navigate through enigmatic labyrinth A survey of chain of thought reasoning: Advances, frontiers and future
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, and Ting Liu · 2024
Later among the works it cites.
Derail yourself: Multi-turn llm jailbreak attack through self-discovered clues, 2024
Qibing Ren, Hao Li, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, and Jing Shao · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Tree of attacks: Jailbreaking black-box llms automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Cited alongside, same era.
Deepinception: Hypnotize large language model to be jailbreaker
Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han · 2023
Cited alongside, same era.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu · 2023
Cited alongside, same era.
Jailbreak and guard aligned language models with only few in-context demonstrations
Zeming Wei, Yifei Wang, Ang Li, Yichuan Mo, and Yisen Wang · 2023
Cited alongside, same era.
Exploring the relationship between architectural design and adversarially robust generalization
Aishan Liu, Shiyu Tang, Siyuan Liang, Ruihao Gong, Boxi Wu, Xianglong Liu, and Dacheng Tao · 2023
Cited alongside, same era.
Later among the works it cites.
Chain of attack: a semantic-driven contextual multi-turn attacker for llm, 2024
Xikang Yang, Xuehai Tang, Songlin Hu, and Jizhong Han · 2024
Later among the works it cites.
Multi-turn context jailbreak attack on large language models from first principles, 2024
Xiongtao Sun, Deyue Zhang, Dongdong Yang, Quanchen Zou, and Hui Li · 2024
Later among the works it cites.
Speak out of turn: Safety vulnerability of large language models in multi-turn dialogue, 2024
Zhenhong Zhou, Jiuyang Xiang, Haopeng Chen, Quan Liu, Zherui Li, and Sen Su · 2024
Later among the works it cites.
Leveraging the context through multi-round interactions for jailbreaking attacks, 2024
Yixin Cheng, Markos Georgopoulos, Volkan Cevher, and Grigorios G. Chrysos · 2024
Later among the works it cites.
Uncertainty is fragile: Manipulating uncertainty in large language models
Qingcheng Zeng, Mingyu Jin, Qinkai Yu, Zhenting Wang, Wenyue Hua, Zihao Zhou, Guangyan Sun, Yanda Meng, Shiqing Ma, Qifan Wang, et al · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size, 2024
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, et al · 2024
Later among the works it cites.
Qwen2 technical report, 2024
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, et al · 2024
Later among the works it cites.
Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, et al · 2024
Later among the works it cites.
Gpt-4o system card, 2024
OpenAI, :, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, et al · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al · 2024
Later among the works it cites.
Gemini api documentation - thinking mode
Google · 2024
Later among the works it cites.
Openai o1 system card, 2024
OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, et al · 2024
Later among the works it cites.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Later among the works it cites.
Smoothllm: Defending large language models against jailbreaking attacks, 2024
Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas · 2024
Later among the works it cites.
Jailguard: A universal detection framework for llm prompt-based attacks
Xiaoyu Zhang, Cen Zhang, Tianlin Li, Yihao Huang, Xiaojun Jia, Ming Hu, Jie Zhang, Yang Liu, Shiqing Ma, and Chao Shen · 2024
Later among the works it cites.
Llms-as-judges: a comprehensive survey on llm-based evaluation methods
Haitao Li, Qian Dong, Junjie Chen, Huixue Su, Yujia Zhou, Qingyao Ai, Ziyi Ye, and Yiqun Liu · 2024
Later among the works it cites.
Visual adversarial attack on vision-language models for autonomous driving
Tianyuan Zhang, Lu Wang, Xinwei Zhang, Yitong Zhang, Boyi Jia, Siyuan Liang, Shengshan Hu, Qiang Fu, Aishan Liu, and Xianglong Liu · 2024
Later among the works it cites.
Siyuan Liang, Kuanrong Liu, Jiajun Gong, Jiawei Liang, Yuan Xun, Ee-Chien Chang, and Xiaochun Cao · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, et al · 2025
Closest in time.
Black-box adversarial attack on vision language models for autonomous driving
Lu Wang, Tianyuan Zhang, Yang Qu, Siyuan Liang, Yuwei Chen, Aishan Liu, Xianglong Liu, and Dacheng Tao · 2025
Closest in time.