Fetching the paper…
Reading the bibliography…
Despite significant advances in alignment techniques, we demonstrate that state-of-the-art language models remain vulnerable to carefully crafted conversational scenarios that can induce various forms of misalignment without explicit jailbreaking.
Constitutional AI: Harmlessness from AI Feedback
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; Chen, C.; Olsson, C.; Olah, C.; Hernandez, D.; Drain, D.; Ganguli, D.; Li, D.; Tran-Johnson, E.; Perez, E.; Kerr, J.; Mueller, J.; Ladish, J.; Landau, J.; Ndousse, K.; Lukosuite, K.; Lovitt, L.; Sellitto, M.; Elhage, N.; Schiefer, N.; Mercado, N.; DasSarma, N.; Lasenby, R.; Larson, R.; Ringer, S.; Johnston, S.; Kravec, S.; Showk, S. E.; Fort, S.; Lanham, T.; Telleen-Lawton, T.; Conerly, T.; Henighan, T.; Hume, T.; Bowman, S. R.; Hatfield-Dodds, Z.; Mann, B.; Amodei, D.; Joseph, N.; McCandlish, S.; Brown, T.; and Kaplan, J. 2022 · 2022
Earlier work this paper cites.
Defending Against Social Engineering Attacks in the Age of LLMs
Ai, L.; Kumarage, T.; Bhattacharjee, A.; Liu, Z.; Hui, Z.; Davinroy, M.; Cook, J.; Cassani, L.; Trapeznikov, K.; Kirchner, M.; Basharat, A.; Hoogs, A.; Garland, J.; Liu, H.; and Hirschberg, J. 2024 · 2024
Earlier work this paper cites.
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Anwar, U.; Saparov, A.; Rando, J.; Paleka, D.; Turpin, M.; Hase, P.; Lubana, E. S.; Jenner, E.; Casper, S.; Sourbut, O.; Edelman, B. L.; Zhang, Z.; Günther, M.; Korinek, A.; Hernandez-Orallo, J.; Hammond, L.; Bigelow, E.; Pan, A.; Langosco, L.; Korbak, T.; Zhang, H.; Zhong, R.; Héigeartaigh, S.; Recchia, G.; Corsi, G.; Chan, A.; Anderljung, M.; Edwards, L.; Petrov, A.; de Witt, C. S.; Motwani, S. R.; Bengio, Y.; Chen, D.; Torr, P. H.; Albanie, S.; Maharaj, T.; Foerster, J.; Tramer, F.; He, H.; Kasirzade, A.; Choi, Y.; and Krueger, D. 2024 · 2024
Earlier work this paper cites.
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Chao, P.; Debenedetti, E.; Robey, A.; Andriushchenko, M.; Croce, F.; Sehwag, V.; Dobriban, E.; Flammarion, N.; Pappas, G. J.; Tramèr, F.; Hassani, H.; and Wong, E. 2024 · 2024
Earlier work this paper cites.
Strong and weak alignment of large language models with human values
Khamassi, M.; Nahon, M.; and Chatila, R. 2024 · 2024
Earlier work this paper cites.
AgentBench: Evaluating LLMs as Agents
Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; Zhang, S.; Deng, X.; Zeng, A.; Du, Z.; Zhang, C.; Shen, S.; Zhang, T.; Su, Y.; Sun, H.; Huang, M.; Dong, Y.; and Tang, J. 2023 · 2024
Earlier work this paper cites.
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
Mou, Y.; Zhang, S.; and Ye, W. 2024 · 2024
Earlier work this paper cites.
Unveiling Narrative Reasoning Limits of Large Language Models with Trope in Movie Synopses
Su, H.-T.; Hsu, Y.-C.; Lin, X.; Shi, X.-Q.; Niu, Y.; Hsu, H.-Y.; Lee, H.-y.; and Hsu, W. H. 2024 · 2024
Earlier work this paper cites.
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
Williams, M.; Carroll, M.; Narang, A.; Weisser, C.; Murphy, B.; and Dragan, A. 2024 · 2024
Cited alongside, same era.
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors
Xie, T.; Qi, X.; Zeng, Y.; Huang, Y.; Sehwag, U. M.; Huang, K.; He, L.; Wei, B.; Li, D.; Sheng, Y.; Jia, R.; Li, B.; Li, K.; Chen, D.; Henderson, P.; and Mittal, P. 2024 · 2024
Cited alongside, same era.
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Yi, S.; Liu, Y.; Sun, Z.; Cong, T.; He, X.; Song, J.; Xu, K.; and Li, Q. 2024 · 2024
Cited alongside, same era.
Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback
Dahlgren Lindström, A.; Methnani, L.; Krause, L.; Ericson, P.; de Rituerto de Troya, I. M.; Coelho Mollo, D.; and Dobbe, R. 2025 · 2025
Cited alongside, same era.
System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection
Guo, J.; and Cai, H. 2025 · 2025
Security Concerns for Large Language Models: A Survey
Li, M. Q.; and Fung, B. C. M. 2025 · 2025
Closest in time.
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
Naik, A.; Quinn, P.; Bosch, G.; Gouné, E.; Zabala, F. J. C.; Brown, J. R.; and Young, E. J. 2025 · 2025
Closest in time.
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
Pathade, C. 2025 · 2025
Closest in time.
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
Reddy, A.; Zagula Bridgewater, A.; Hs, R.; and Saban, N. 2025 · 2025
Closest in time.
SoK: On the Offensive Potential of AI
Schroer, S. L.; Apruzzese, G.; Human, S.; Laskov, P.; Anderson, H. S.; Bernroider, E. W.; Fass, A.; Nassi, B.; Rimmer, V.; Roli, F.; Salam, S.; Ashley Shen, C. E.; Sunyaev, A.; Wadhwa-Brown, T.; Wagner, I.; and Wang, G. 2024 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
Hackett, W.; Birch, L.; Trawicki, S.; Suri, N.; and Garraghan, P. 2025 · 2025
Cited alongside, same era.
Evaluating the Effectiveness of Psychological Prompt Injection Attacks on Large Language Models for Social Engineering Artifact Generation
Heverin, T.; and Cohen, E. 2025 · 2025
Cited alongside, same era.
A Survey on Progress in LLM Alignment from the Perspective of Reward Design
Ji, M.; Wu, Y.; Wu, Z.; Wang, S.; Yang, J.; Dras, M.; and Naseem, U. 2025 · 2025
Cited alongside, same era.
Agentic Misalignment: How LLMs could be insider threats \ Anthropic
????
Cited in the paper.
A Technical Survey of Reinforcement Learning Techniques for Large Language Models
Srivastava, S. S.; and Aggarwal, V. 2025 · 2025
Closest in time.
StructTransform: A Scalable Attack Surface for Safety-Aligned Large Language Models
Yoosuf, S.; Ali, T.; Lekssays, A.; AlSabah, M.; and Khalil, I. 2025 · 2025
Closest in time.