Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are evolving into autonomous decision-makers, raising concerns about catastrophic risks in high-stakes scenarios, particularly in Chemical, Biological, Radiological and Nuclear (CBRN) domains.
Growing Artificial Societies: Social Science from the Bottom Up
Joshua M Epstein. 1996 · 1996
Earlier work this paper cites.
Approach and avoidance motivation and achievement goals
Andrew J Elliot. 1999 · 1999
Earlier work this paper cites.
Expectancy–value theory of achievement motivation
Allan Wigfield and Jacquelynne S Eccles. 2000 · 2000
Earlier work this paper cites.
Moral emotions and moral behavior
June Price Tangney, Jeff Stuewig, and Debra J Mashek. 2007 · 2007
Earlier work this paper cites.
Agent-based modeling and simulation
Charles M Macal and Michael J North. 2009 · 2009
Earlier work this paper cites.
Robot minds and human ethics: the need for a comprehensive model of moral decision making
Wendell Wallach. 2010 · 2010
Earlier work this paper cites.
Trusted execution environment: What it is, and what it is not
Mohamed Sabt, Mohammed Achemlal, and Abdelmadjid Bouabdallah. 2015 · 2015
Earlier work this paper cites.
Statistical tests, p values, confidence intervals, and power: a guide to misinterpretations
Sander Greenland, Stephen J Senn, Kenneth J Rothman, John B Carlin, Charles Poole, Steven N Goodman, and Douglas G Altman. 2016 · 2016
Earlier work this paper cites.
Ethical decision-making theory: An integrated approach
Mark S Schwartz. 2016 · 2016
Earlier work this paper cites.
Human behavior in the social environment: A social systems approach
Irl Carter. 2017 · 2017
Earlier work this paper cites.
Policy analysis: Concepts and practice
David L Weimer and Aidan R Vining. 2017 · 2017
Earlier work this paper cites.
How might artificial intelligence affect the risk of nuclear war?
Andrew J Lohn and Edward Geist. 2018 · 2018
Earlier work this paper cites.
Liability for ai decision-making: some legal and ethical considerations
Iria Giuffrida. 2019 · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. 2021 · 2021
Earlier work this paper cites.
Moral decision making: From bentham to veil of ignorance via perspective taking accessibility
Rose Martin, Petko Kusev, Joseph Teal, Victoria Baranova, and Bruce Rigal. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Earlier work this paper cites.
Dual use of artificial-intelligence-powered drug discovery
Fabio Urbina, Filippa Lentzos, Cédric Invernizzi, and Sean Ekins. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Managing ai risks in an era of rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, et al. 2023 · 2023
Earlier work this paper cites.
Executive order on the safe, secure, and trustworthy development and use of artificial intelligence
Joseph R Biden. 2023 · 2023
Earlier work this paper cites.
Scheming ais: Will ais fake alignment during training in order to get power?
Joe Carlsmith. 2023 · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023 · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Earlier work this paper cites.
Agentsims: An open-source sandbox for large language model evaluation
Jiaju Lin, Haoran Zhao, Aochi Zhang, Yiting Wu, Huqiuyue Ping, and Qin Chen. 2023 · 2023
Earlier work this paper cites.
AI autonomy: Self-initiated open-world continual learning and adaptation
Bing Liu, Sahisnu Mazumder, Eric Robertson, and Scott Grigsby. 2023 · 2023
Earlier work this paper cites.
Levels of agi: Operationalizing progress on the path to agi
Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, and Shane Legg. 2023 · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023 · 2023
Earlier work this paper cites.
A survey of hallucination in large foundation models
Vipula Rawte, Amit Sheth, and Amitava Das. 2023 · 2023
Earlier work this paper cites.
Provably safe systems: the only path to controllable agi
Max Tegmark and Steve Omohundro. 2023 · 2023
Earlier work this paper cites.
Fake alignment: Are llms really aligned well?
Yixu Wang, Yan Teng, Kexin Huang, Chengqi Lyu, Songyang Zhang, Wenwei Zhang, Xingjun Ma, Yu-Gang Jiang, Yu Qiao, and Yingchun Wang. 2023 · 2023
Earlier work this paper cites.
Defending chatgpt against jailbreak attack via self-reminders
Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. 2023 · 2023
Cited alongside, same era.
Rongwu Xu, Brian S Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, and Han Qiu. 2023 · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Cited alongside, same era.
The eu artificial intelligence act
EU Artificial Intelligence Act. 2024 · 2024
Cited alongside, same era.
Exploring the psychology of llms’ moral and legal reasoning
Guilherme FCF Almeida, José Luiz Nunes, Neele Engelmann, Alex Wiegmann, and Marcelo de Araújo. 2024 · 2024
Cited alongside, same era.
Emergence of social norms in large language model-based agent societies
Siyue Ren, Zhiyao Cui, Ruiqi Song, Zhen Wang, and Shuyue Hu. 2024 · 2024
Later among the works it cites.
Escalation risks from language models in military and diplomatic decision-making
Juan-Pablo Rivera, Gabriel Mukobi, Anka Reuel, Max Lamparth, Chandler Smith, and Jacquelyn Schneider. 2024 · 2024
Later among the works it cites.
Large language models can strategically deceive their users when put under pressure
Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn. 2024 · 2024
Later among the works it cites.
Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments
Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor. 2024 · 2024
Later among the works it cites.
Ai-liedar: Examine the trade-off between utility and truthfulness in llm agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Does refusal training in llms generalize to the past tense?
Maksym Andriushchenko and Nicolas Flammarion. 2024 · 2024
Cited alongside, same era.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al. 2024 · 2024
Cited alongside, same era.
Fairmonitor: A dual-framework for detecting stereotypes and biases in large language models
Yanhong Bai, Jiabao Zhao, Jinxin Shi, Zhentao Xie, Xingjiao Wu, and Liang He. 2024 · 2024
Cited alongside, same era.
Towards evaluations-based safety cases for ai scheming
Mikita Balesni, Marius Hobbhahn, David Lindner, Alex Meinke, Tomek Korbak, Joshua Clymer, Buck Shlegeris, Jérémy Scheurer, Rusheb Shah, Nicholas Goldowsky-Dill, et al. 2024 · 2024
Cited alongside, same era.
Optimizing reasoning abilities in large language models: A step-by-step approach
Zhiyuan Chen, Yaning Li, and Kairui Wang. 2024 · 2024
Cited alongside, same era.
Securing the future of genai: Policy and technology
Mihai Christodorescu, Ryan Craven, Soheil Feizi, Neil Gong, Mia Hoffmann, Somesh Jha, Zhengyuan Jiang, Mehrdad Saberi Kamarposhti, John Mitchell, Jessica Newman, et al. 2024 · 2024
Cited alongside, same era.
Towards guaranteed safe ai: A framework for ensuring robust and reliable ai systems
David Dalrymple, Joar Skalse, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, et al. 2024 · 2024
Cited alongside, same era.
Zhe Su, Xuhui Zhou, Sanketh Rangreji, Anubha Kabra, Julia Mendelsohn, Faeze Brahman, and Maarten Sap. 2024 · 2024
Later among the works it cites.
Qwq: Reflect deeply on the boundaries of the unknown
Qwen Team. 2024 · 2024
Later among the works it cites.
Ai sandbagging: Language models can strategically underperform on evaluations
Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F Brown, and Francis Rhys Ward. 2024 · 2024
Later among the works it cites.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024 · 2024
Later among the works it cites.
Livebench: A challenging, contamination-free llm benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Siddartha Naidu, et al. 2024 · 2024
Later among the works it cites.
Sorry-bench: Systematically evaluating large language model safety refusal behaviors
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al. 2024 · 2024
Later among the works it cites.
Rongwu Xu, Zi’an Zhou, Tianwei Zhang, Zehan Qi, Su Yao, Ke Xu, Wei Xu, and Han Qiu. 2024 · 2024
Later among the works it cites.
Toolsword: Unveiling safety issues of large language models in tool learning across three stages
Junjie Ye, Sixian Li, Guanyu Li, Caishuang Huang, Songyang Gao, Yilong Wu, Qi Zhang, Tao Gui, and Xuanjing Huang. 2024 · 2024
Later among the works it cites.
Yangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng, Yupeng Cao, Zhi Chen, Jordan W Suchow, Rong Liu, Zhenyu Cui, Zhaozhuo Xu, et al. 2024 · 2024
Later among the works it cites.
Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Jiahao Xu, Tian Liang, Pinjia He, and Zhaopeng Tu. 2024 · 2024
Later among the works it cites.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024 · 2024
Later among the works it cites.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024 · 2024
Later among the works it cites.
Agent-safetybench: Evaluating the safety of llm agents
Zhexin Zhang, Shiyao Cui, Yida Lu, Jingzhuo Zhou, Junxiao Yang, Hongning Wang, and Minlie Huang. 2024 · 2024
Later among the works it cites.
Introducing meta llama 3
Meta AI. 2023 · 2025
Closest in time.
Claude 3.5: Sonnet
Anthropic. 2023 · 2025
Closest in time.
Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 sonnet
Anthropic. 2024a · 2025
Closest in time.
Reflections on our responsible scaling policy
Anthropic. 2024b · 2025
Closest in time.
International ai safety report 2025
Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, et al. 2025b · 2025
Closest in time.
Man who exploded tesla cybertruck outside trump hotel in las vegas used generative ai, police say
Mike Catalini. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.
Rules created by symbolic systems cannot constrain a learning system
Shih-Wai Lin, Rongwu Xu, Xiaojian Li, and Wei Xu. 2025 · 2025
Closest in time.
Dall·e 3
OpenAI. 2023 · 2025
Closest in time.
Openai model specification - follow the chain of command
OpenAI. 2024 · 2025
Closest in time.
O3 mini system card
OpenAI. 2025 · 2025
Closest in time.
The case for ensuring that powerful ais are controlled
Nick Ord. 2024 · 2025
Closest in time.
Value creation for healthcare ecosystems through artificial intelligence applied to physician-to-physician communication: A systematic review
Beny Rubinstein and Sergio Matos. 2025 · 2025
Closest in time.
CWMD-DHS CBRN AI EO Report - Public Release
U.S. Department of Homeland Security. 2024 · 2025
Closest in time.