Fetching the paper…
Reading the bibliography…
Evaluations of large language model (LLM) risks and capabilities are increasingly being incorporated into AI risk management and governance frameworks.
Generalizing the safety factor approach
Jonas Clausen, Sven Ove Hansson, and Fred Nilsson · 2006
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts, 2020
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV au2, Eric Wallace, and Sameer Singh · 2010
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Regularizing deep networks using efficient layerwise adversarial training
Swami Sankaranarayanan, Arpit Jain, Rama Chellappa, and Ser Nam Lim · 2018
Earlier work this paper cites.
Harnessing the vulnerability of latent layers in adversarially trained models, 2019
Mayank Singh, Abhishek Sinha, Nupur Kumari, Harshitha Machiraju, Balaji Krishnamurthy, and Vineeth N Balasubramanian · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Continual learning and private unlearning
Bo Liu, Qiang Liu, and Peter Stone · 2022
Earlier work this paper cites.
Outsider oversight: Designing a third party audit ecosystem for ai governance
Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, and Daniel Ho · 2022
Earlier work this paper cites.
Towards publicly accountable frontier llms: Building an external scrutiny ecosystem under the aspire framework
Markus Anderljung, Everett Thornton Smith, Joe O’Brien, Lisa Soder, Benjamin Bucknall, Emma Bluemke, Jonas Schuett, Robert Trager, Lacey Strahm, and Rumman Chowdhury · 2023
Earlier work this paper cites.
Language model unalignment: Parametric red-teaming to expose hidden harms and biases
Rishabh Bhardwaj and Soujanya Poria · 2023
Earlier work this paper cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
Earlier work this paper cites.
Scaling laws for adversarial attacks on language model activations
Stanislav Fort · 2023
Earlier work this paper cites.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P Dick, Hidenori Tanaka, Edward Grefenstette, Tim Rocktäschel, and David Scott Krueger · 2023
Earlier work this paper cites.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish · 2023
Earlier work this paper cites.
AI Risk Management Framework: AI RMF (1.0), January 2023
NIST · 2023
Earlier work this paper cites.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Vaidehi Patil, Peter Hase, and Mohit Bansal · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Earlier work this paper cites.
Towards best practices in agi safety and governance: A survey of expert opinion
Jonas Schuett, Noemi Dreksler, Markus Anderljung, David McCaffary, Lennart Heim, Emma Bluemke, and Ben Garfinkel · 2023
Earlier work this paper cites.
Adversarial attacks and defenses in large language models: Old and new threats
Leo Schwinn, David Dobre, Stephan Günnemann, and Gauthier Gidel · 2023
Earlier work this paper cites.
Survey of vulnerabilities in large language models revealed by adversarial attacks
Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael Abu-Ghazaleh · 2023
Earlier work this paper cites.
Model evaluation for extreme risks
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al · 2023
Earlier work this paper cites.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
A pro-innovation approach to AI regulation
UK DSIT · 2023
Earlier work this paper cites.
Haoran Wang and Kai Shu · 2023
Earlier work this paper cites.
Shadow alignment: The ease of subverting safely-aligned language models
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin · 2023
Earlier work this paper cites.
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach · 2023
Earlier work this paper cites.
Removing rlhf protections in gpt-4 via fine-tuning
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang · 2023
Earlier work this paper cites.
Adversarial machine learning in latent representations of neural networks
Milin Zhang, Mohammad Abdi, and Francesco Restuccia · 2023
Earlier work this paper cites.
Agieval: A human-centric benchmark for evaluating foundation models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan · 2023
Cited alongside, same era.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Cited alongside, same era.
Many-shot jailbreaking
Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, et al · 2024
Cited alongside, same era.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda · 2024
Cited alongside, same era.
Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues
Openai system card: December 2024, 2024
OpenAI · 2024
Later among the works it cites.
Navigating the safety landscape: Measuring risks in finetuning large language models
ShengYun Peng, Pin-Yu Chen, Matthew Hull, and Duen Horng Chau · 2024
Later among the works it cites.
Open problems in technical ai governance
Anka Reuel, Ben Bucknall, Stephen Casper, Tim Fist, Lisa Soder, Onni Aarne, Lewis Hammond, Lujain Ibrahim, Alan Chan, Peter Wills, et al · 2024
Later among the works it cites.
Representation noising effectively prevents harmful fine-tuning on llms
Domenic Rosati, Jan Wehner, Kai Williams, Łukasz Bartoszcze, David Atanasov, Robie Gonzales, Subhabrata Majumdar, Carsten Maple, Hassan Sajjad, and Frank Rudzicz · 2024
Later among the works it cites.
Fast adversarial attacks on language models in one gpu minute, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, et al · 2024
Cited alongside, same era.
Bill No. 2338 of 2023: Regulating the Use of Artificial Intelligence, Including Algorithm Design and Technical Standards, 2023
Brazil · 2024
Cited alongside, same era.
AI and Data Act: Part of Bill C-27, Digital Charter Implementation Act, 2022, 2022
Canada · 2024
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei W Koh, Daphne Ippolito, Florian Tramer, and Ludwig Schmidt · 2024
Cited alongside, same era.
Defending against unforeseen failure modes with latent adversarial training
Stephen Casper, Lennart Schulze, Oam Patel, and Dylan Hadfield-Menell · 2024
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries, 2024
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2024
Cited alongside, same era.
Interim Measures for the Management of Generative Artificial Intelligence Services, 2023
China · 2024
Cited alongside, same era.
Breaking down the defenses: A comparative survey of attacks on large language models
Arijit Ghosh Chowdhury, Md Mofijul Islam, Vaibhav Kumar, Faysal Hossain Shezan, Vinija Jain, and Aman Chadha · 2024
Cited alongside, same era.
Vinu Sankar Sadasivan, Shoumik Saha, Gaurang Sriramanan, Priyatham Kattakinda, Atoosa Chegini, and Soheil Feizi · 2024
Later among the works it cites.
Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel, and Stephan Gunnemann · 2024
Later among the works it cites.
Latent adversarial training improves robustness to persistent harmful behaviors in llms
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, et al · 2024
Later among the works it cites.
Ununlearning: Unlearning is not sufficient for content regulation in advanced generative ai
Ilia Shumailov, Jamie Hayes, Eleni Triantafillou, Guillermo Ortiz-Jimenez, Nicolas Papernot, Matthew Jagielski, Itay Yona, Heidi Howard, and Eugene Bagdasaryan · 2024
Later among the works it cites.
A strongreject for empty jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al · 2024
Later among the works it cites.
Tamper-resistant safeguards for open-weight llms
Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, et al · 2024
Later among the works it cites.
Ai sandbagging: Language models can strategically underperform on evaluations
Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F Brown, and Francis Rhys Ward · 2024
Later among the works it cites.
Assessing the brittleness of safety alignment via pruning and low-rank modifications
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson · 2024
Later among the works it cites.
Efficient adversarial training in llms with continuous attacks
Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel, and Leo Schwinn · 2024
Later among the works it cites.
On the vulnerability of safety alignment in open-access llms
Jingwei Yi, Rui Ye, Qisi Chen, Bin Zhu, Siheng Chen, Defu Lian, Guangzhong Sun, Xing Xie, and Fangzhao Wu · 2024
Later among the works it cites.
Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Jiahao Xu, Tian Liang, Pinjia He, and Zhaopeng Tu · 2024
Later among the works it cites.
Beear: Embedding-based adversarial removal of safety backdoors in instruction-tuned language models
Yi Zeng, Weiyu Sun, Tran Ngoc Huynh, Dawn Song, Bo Li, and Ruoxi Jia · 2024
Later among the works it cites.
Does your llm truly unlearn? an embarrassingly simple approach to recover unlearned knowledge
Zhiwei Zhang, Fali Wang, Xiaomin Li, Zongyu Wu, Xianfeng Tang, Hui Liu, Qi He, Wenpeng Yin, and Suhang Wang · 2024
Later among the works it cites.
Shuai Zhao, Meihuizi Jia, Zhongliang Guo, Leilei Gan, Xiaoyu Xu, Xiaobao Wu, Jie Fu, Yichao Feng, Fengjun Pan, and Luu Anh Tuan · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Later among the works it cites.
Open problems in machine unlearning for ai safety
Fazl Barez, Tingchen Fu, Ameya Prabhu, Stephen Casper, Amartya Sanyal, Adel Bibi, Aidan O’Gara, Robert Kirk, Ben Bucknall, Tim Fist, et al · 2025
Closest in time.
Towards a science of ai evaluations
Yarin Gal · 2025
Closest in time.
Cascade: Exploring hierarchical inference in language models, 2023
Haize Labs · 2025
Closest in time.
The elicitation game: Evaluating capability elicitation techniques
Felix Hofstätter, Teun van der Weij, Jayden Teoh, Henning Bartsch, and Francis Rhys Ward · 2025
Closest in time.
Act on the protection of personal information, 2025
Korea · 2025
Closest in time.
Safety pretraining: Toward the next generation of safe ai
Pratyush Maini, Sachin Goyal, Dylan Sam, Alex Robey, Yash Savani, Yiding Jiang, Andy Zou, Zacharcy C Lipton, and J Zico Kolter · 2025
Closest in time.
Gauss-newton unlearning for the llm era
Lev E McKinney, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A McIlraith, and Roger Baker Grosse · 2025
Closest in time.
From dormant to deleted: Tamper-resistant unlearning through weight-space regularization, 2025
Shoaib Ahmed Siddiqui, Adrian Weller, David Krueger, Gintare Karolina Dziugaite, Michael Curtis Mozer, and Eleni Triantafillou · 2025
Closest in time.
Modifying llm beliefs with synthetic document finetuning
Rowan Wang, Avery Griffin, Johannes Treutlein, Ethan Perez, Julian Michael, Fabien Roger, and Sam Marks · 2025
Closest in time.
Catastrophic failure of llm unlearning via quantization, 2025
Zhiwei Zhang, Fali Wang, Xiaomin Li, Zongyu Wu, Xianfeng Tang, Hui Liu, Qi He, Wenpeng Yin, and Suhang Wang · 2025
Closest in time.