Fetching the paper…
Reading the bibliography…
As AI systems become more advanced, companies and regulators will make difficult decisions about whether it is safe to train and deploy them.
An Analysis of the Principal-Agent Problem
Grossman, Sanford J., & Hart, Oliver D. 1983 · 1983
Earlier work this paper cites.
Space shuttle RTOS bayesian network
Morris, Allan, & Beling, Peter. 2001 (11) · 2001
Earlier work this paper cites.
An independent review into the broader issues surrounding the loss of the RAF Nimrod MR2 Aircraft XV230 in Afghanistan in 2006
Haddon-Cave QC, Charles. 2009 · 2006
Earlier work this paper cites.
Safety, Performance and Interoperability Requirements Document for the In-Trail Procedure in Oceanic Airspace (ATSA-ITP) Application
for Aeronautics, Radio Technical Commission. 2008 · 2008
Earlier work this paper cites.
The Basic AI Drives
Omohundro, Stephen M. 2008 · 2008
Earlier work this paper cites.
The Singularity: A Philosophical Analysis
Chalmers, David J. 2010 · 2010
Earlier work this paper cites.
White Paper on the Use of Safety Cases in Certification and Regulation
Leveson, Nancy. 2011 · 2011
Earlier work this paper cites.
Thinking Inside the Box: Controlling and Using an Oracle AI
Armstrong, Stuart, Sandberg, Anders, & Bostrom, Nick. 2012 · 2012
Earlier work this paper cites.
Leakproofing the Singularity: Artificial Intelligence Confinement Problem
Yampolskiy, Roman. 2012 · 2012
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, Nick. 2014 · 2014
Earlier work this paper cites.
Annex 19 - Safety Management
ICAO. 2016 · 2016
Earlier work this paper cites.
Should healthcare providers do safety cases? Lessons from a cross-industry review of safety case practices
Sujan, Mark A., Habli, Ibrahim, Kelly, Tim P., Pozzi, Simone, & Johnson, Christopher W. 2016 · 2016
Earlier work this paper cites.
Practical Insights and Lessons Learned on Implementing Expert Elicitation
Xing, Jing, & Morrow, Stephanie. 2016 · 2016
Earlier work this paper cites.
Clarifying AI Alignment
Christiano, Paul. 2017 · 2017
Earlier work this paper cites.
5 - Failure Modes and Effects Analysis
Kritzinger, Duane. 2017 · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Christiano, Paul, Shlegeris, Buck, & Amodei, Dario. 2018 · 2018
Earlier work this paper cites.
Goals-based and Rules-Based Approaches to Regulation
Dr. Christopher Decker. May 2018 · 2018
Earlier work this paper cites.
AI safety via debate
Irving, Geoffrey, Christiano, Paul, & Amodei, Dario. 2018 · 2018
Earlier work this paper cites.
Language Models Are Unsupervised Multitask Learners
Radford, Alec, Wu, Jeffrey, Child, Rewon, Luan, David, Amodei, Dario, & Sutskever, Ilya. 2018 · 2018
Earlier work this paper cites.
On the Measure of Intelligence
Chollet, François. 2019 · 2019
Earlier work this paper cites.
Are You Sure Your Software Will Not Kill Anyone? – Communications of the ACM
Leveson, Nancy. 2020 · 2020
Earlier work this paper cites.
The Coup-Proofing Toolbox: Institutional Power, Military Effectiveness, and the Puzzle of Nazi Germany
Reiter, Dan. 2020 · 2020
Earlier work this paper cites.
Safety Assessment Principles for Nuclear Facilities
UK Office for Nuclear Regulation. 2020 · 2020
Earlier work this paper cites.
Learn About FDA Advisory Committees
US Food and Drug Administration. 2020 · 2020
Earlier work this paper cites.
Eliciting latent knowledge: How to tell if your eyes deceive you
Christiano, Paul, Xu, Mark, & Cotra, Ajeya. 2021 · 2021
Cited alongside, same era.
Truthful AI: Developing and governing AI that does not lie
Evans, Owain, Cotton-Barratt, Owen, Finnveden, Lukas, Bales, Adam, Balwit, Avital, Wills, Peter, Righetti, Luca, & Saunders, William. 2021 · 2021
Cited alongside, same era.
Goal Structuring Notation Community Standard, Version 3
Group, Assurance Case Working. 2021 · 2021
Cited alongside, same era.
A Review of Formal Methods applied to Machine Learning
Urban, Caterina, & Miné, Antoine. 2021 · 2021
Cited alongside, same era.
Renewal of a certificate of compliance
US Nuclear Regulatory Commission. 2021 · 2021
Cited alongside, same era.
Challenges in Detoxifying Language Models
Welbl, Johannes, Glaese, Amelia, Uesato, Jonathan, Dathathri, Sumanth, Mellor, John, Hendricks, Lisa Anne, Anderson, Kirsty, Kohli, Pushmeet, Coppin, Ben, & Huang, Po-Sen. 2021 · 2021
Fairness and Bias in Artificial Intelligence: A Brief Survey of Sources, Impacts, and Mitigation Strategies
Ferrara, Emilio. 2023 · 2023
Later among the works it cites.
Geoffrey Hinton tells us why he’s now scared of the tech he helped build
Heaven, Will Douglas. 2023 · 2023
Later among the works it cites.
An Overview of Catastrophic AI Risks
Hendrycks, Dan, Mazeika, Mantas, & Woodside, Thomas. 2023 · 2023
Later among the works it cites.
FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
Kim, Hyunwoo, Sclar, Melanie, Zhou, Xuhui, Bras, Ronan Le, Kim, Gunhee, Choi, Yejin, & Sap, Maarten. 2023 · 2023
Later among the works it cites.
Risk assessment at AGI companies: A review of popular risk assessment techniques from other safety-critical industries
Koessler, Leonie, & Schuett, Jonas. 2023 · 2023
Later among the works it cites.
Towards a Situational Awareness Benchmark for LLMs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Constitutional AI: Harmlessness from AI Feedback
Bai, Yuntao, Kadavath, Saurav, Kundu, Sandipan, Askell, Amanda, Kernion, Jackson, Jones, Andy, Chen, Anna, Goldie, Anna, Mirhoseini, Azalia, McKinnon, Cameron, Chen, Carol, Olsson, Catherine, Olah, Christopher, Hernandez, Danny, Drain, Dawn, Ganguli, Deep, Li, Dustin, Tran-Johnson, Eli, Perez, Ethan, Kerr, Jamie, Mueller, Jared, Ladish, Jeffrey, Landau, Joshua, Ndousse, Kamal, Lukosuite, Kamile, Lovitt, Liane, Sellitto, Michael, Elhage, Nelson, Schiefer, Nicholas, Mercado, Noemi, DasSarma, Nova, Lasenby, Robert, Larson, Robin, Ringer, Sam, Johnston, Scott, Kravec, Shauna, Showk, Sheer El, Fort, Stanislav, Lanham, Tamera, Telleen-Lawton, Timothy, Conerly, Tom, Henighan, Tom, Hume, Tristan, Bowman, Samuel R., Hatfield-Dodds, Zac, Mann, Ben, Amodei, Dario, Joseph, Nicholas, McCandlish, Sam, Brown, Tom, & Kaplan, Jared. 2022 · 2022
Cited alongside, same era.
Discovering Latent Knowledge in Language Models Without Supervision
Burns, Collin, Ye, Haotian, Klein, Dan, & Steinhardt, Jacob. 2022 · 2022
Cited alongside, same era.
Poisoning and Backdooring Contrastive Learning
Carlini, Nicholas, & Terzis, Andreas. 2022 · 2022
Cited alongside, same era.
Formalizing the presumption of independence
Christiano, Paul, Neyman, Eric, & Xu, Mark. 2022 · 2022
Cited alongside, same era.
Discovering Agents
Kenton, Zachary, Kumar, Ramana, Farquhar, Sebastian, Richens, Jonathan, MacDermott, Matt, & Everitt, Tom. 2022 · 2022
Cited alongside, same era.
Defining and Characterizing Reward Hacking
Skalse, Joar, Howe, Nikolaus H. R., Krasheninnikov, Dmitrii, & Krueger, David. 2022 · 2022
Cited alongside, same era.
Laine, Rudolf, Meinke, Alexander, & Evans, Owain. 2023 · 2023
Later among the works it cites.
Measuring Faithfulness in Chain-of-Thought Reasoning
Lanham, Tamera, Chen, Anna, Radhakrishnan, Ansh, Steiner, Benoit, Denison, Carson, Hernandez, Danny, Li, Dustin, Durmus, Esin, Hubinger, Evan, Kernion, Jackson, Lukošiūtė, Kamilė, Nguyen, Karina, Cheng, Newton, Joseph, Nicholas, Schiefer, Nicholas, Rausch, Oliver, Larson, Robin, McCandlish, Sam, Kundu, Sandipan, Kadavath, Saurav, Yang, Shannon, Henighan, Thomas, Maxwell, Timothy, Telleen-Lawton, Timothy, Hume, Tristan, Hatfield-Dodds, Zac, Kaplan, Jared, Brauner, Jan, Bowman, Samuel R., & Perez, Ethan. 2023 · 2023
Later among the works it cites.
Preparedness
OpenAI. 2023 · 2023
Later among the works it cites.
Preventing Language Models From Hiding Their Reasoning
Roger, Fabien, & Greenblatt, Ryan. 2023 · 2023
Later among the works it cites.
Towards best practices in AGI safety and governance: A survey of expert opinion
Schuett, Jonas, Dreksler, Noemi, Anderljung, Markus, McCaffary, David, Heim, Lennart, Bluemke, Emma, & Garfinkel, Ben. 2023 · 2023
Later among the works it cites.
Model evaluation for extreme risks
Shevlane, Toby, Farquhar, Sebastian, Garfinkel, Ben, Phuong, Mary, Whittlestone, Jess, Leung, Jade, Kokotajlo, Daniel, Marchal, Nahema, Anderljung, Markus, Kolt, Noam, Ho, Lewis, Siddarth, Divya, Avin, Shahar, Hawkins, Will, Kim, Been, Gabriel, Iason, Bolina, Vijay, Clark, Jack, Bengio, Yoshua, Christiano, Paul, & Dafoe, Allan. 2023 · 2023
Later among the works it cites.
Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Turpin, Miles, Michael, Julian, Perez, Ethan, & Bowman, Samuel R. 2023 · 2023
Later among the works it cites.
A Survey on Large Language Model based Autonomous Agents
Wang, Lei, Ma, Chen, Feng, Xueyang, Zhang, Zeyu, Yang, Hao, Zhang, Jingsen, Chen, Zhiyuan, Tang, Jiakai, Chen, Xu, Lin, Yankai, Zhao, Wayne Xin, Wei, Zhewei, & Wen, Ji-Rong. 2023 · 2023
Later among the works it cites.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, Jason, Wang, Xuezhi, Schuurmans, Dale, Bosma, Maarten, Ichter, Brian, Xia, Fei, Chi, Ed, Le, Quoc, & Zhou, Denny. 2023 · 2023
Later among the works it cites.
Representation Engineering: A Top-Down Approach to AI Transparency
Zou, Andy, Phan, Long, Chen, Sarah, Campbell, James, Guo, Phillip, Ren, Richard, Pan, Alexander, Yin, Xuwang, Mazeika, Mantas, Dombrowski, Ann-Kathrin, Goel, Shashwat, Li, Nathaniel, Byun, Michael J., Wang, Zifan, Mallen, Alex, Basart, Steven, Koyejo, Sanmi, Song, Dawn, Fredrikson, Matt, Kolter, J. Zico, & Hendrycks, Dan. 2023 · 2023
Later among the works it cites.
SB-1047 Safe and Secure Innovation for Frontier Artificial Intelligence Systems Act
California State Legislature. 2024 · 2024
Closest in time.
Catching AIs red-handed
Greenblatt, Ryan, & Shlegeris, Buck. 2024 · 2024
Closest in time.
AI Control: Improving Safety Despite Intentional Subversion
Greenblatt, Ryan, Shlegeris, Buck, Sachan, Kshitij, & Roger, Fabien. 2024 · 2024
Closest in time.
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
Gu, Xiangming, Zheng, Xiaosen, Pang, Tianyu, Du, Chao, Liu, Qian, Wang, Ye, Jiang, Jing, & Lin, Min. 2024 · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Hubinger, Evan, Denison, Carson, Mu, Jesse, Lambert, Mike, Tong, Meg, MacDiarmid, Monte, Lanham, Tamera, Ziegler, Daniel M., Maxwell, Tim, Cheng, Newton, Jermyn, Adam, Askell, Amanda, Radhakrishnan, Ansh, Anil, Cem, Duvenaud, David, Ganguli, Deep, Barez, Fazl, Clark, Jack, Ndousse, Kamal, Sachan, Kshitij, Sellitto, Michael, Sharma, Mrinank, DasSarma, Nova, Grosse, Roger, Kravec, Shauna, Bai, Yuntao, Witten, Zachary, Favaro, Marina, Brauner, Jan, Karnofsky, Holden, Christiano, Paul, Bowman, Samuel R., Graham, Logan, Kaplan, Jared, Mindermann, Sören, Greenblatt, Ryan, Shlegeris, Buck, Schiefer, Nicholas, & Perez, Ethan. 2024 · 2024
Closest in time.
Evaluating Language-Model Agents on Realistic Autonomous Tasks
Kinniment, Megan, Sato, Lucas Jun Koba, Du, Haoxing, Goodrich, Brian, Hasin, Max, Chan, Lawrence, Miles, Luke Harold, Lin, Tao R., Wijk, Hjalmar, Burget, Joel, Ho, Aaron, Barnes, Elizabeth, & Christiano, Paul. 2024 · 2024
Closest in time.
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Li, Nathaniel, Pan, Alexander, Gopal, Anjali, Yue, Summer, Berrios, Daniel, Gatti, Alice, Li, Justin D., Dombrowski, Ann-Kathrin, Goel, Shashwat, Phan, Long, Mukobi, Gabriel, Helm-Burger, Nathan, Lababidi, Rassin, Justen, Lennart, Liu, Andrew B., Chen, Michael, Barrass, Isabelle, Zhang, Oliver, Zhu, Xiaoyuan, Tamirisa, Rishub, Bharathi, Bhrugu, Khoja, Adam, Zhao, Zhenqi, Herbert-Voss, Ariel, Breuer, Cort B., Zou, Andy, Mazeika, Mantas, Wang, Zifan, Oswal, Palash, Liu, Weiran, Hunt, Adam A., Tienken-Harder, Justin, Shih, Kevin Y., Talley, Kemper, Guan, John, Kaplan, Russell, Steneker, Ian, Campbell, David, Jokubaitis, Brad, Levinson, Alex, Wang, Jean, Qian, William, Karmakar, Kallol Krishna, Basart, Steven, Fitz, Stephen, Levine, Mindy, Kumaraguru, Ponnurangam, Tupakula, Uday, Varadharajan, Vijay, Shoshitaishvili, Yan, Ba, Jimmy, Esvelt, Kevin M., Wang, Alexandr, & Hendrycks, Dan. 2024 · 2024
Closest in time.
The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study
Mouton, Christopher A., Lucas, Caleb, & Guest, Ella. 2024 · 2024
Closest in time.
A Comprehensive Survey of Continual Learning: Theory, Method and Application
Wang, Liyuan, Zhang, Xingxing, Su, Hang, & Zhu, Jun. 2024 · 2024
Closest in time.