Fetching the paper…
Reading the bibliography…
Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose existential risks.
Optimal Policies Tend to Seek Power
Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli. 2023 · 1912
Earlier work this paper cites.
Econometric policy evaluation: A critique
Robert E. Lucas. 1976 · 1976
Earlier work this paper cites.
Assessing the impact of planned social change
Donald T. Campbell. 1979 · 1979
Earlier work this paper cites.
Problems of Monetary Management: The UK Experience
C. A. E. Goodhart. 1984 · 1984
Earlier work this paper cites.
‘Improving ratings’: audit in the British University system
Marilyn Strathern. 1997 · 1997
Earlier work this paper cites.
Goodhart’s Law: its origins, meaning and implications for monetary policy
Alec Chrystal and Paul Mizen. 2003 · 2003
Earlier work this paper cites.
Goodhart’s Law and Performance Indicators in Higher Education
Lewis Elton. 2004 · 2004
Earlier work this paper cites.
The inevitable corruption of indicators and educators through high-stakes testing
David C Berliner and Sharon L Nichols. 2005 · 2005
Earlier work this paper cites.
AI Research Considerations for Human Existential Safety (ARCHES)
Andrew Critch and David Krueger. 2020 · 2006
Earlier work this paper cites.
Measuring Up: What Educational Testing Really Tells Us
Daniel Koretz. 2008 · 2008
Earlier work this paper cites.
The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents
Nick Bostrom. 2012 · 2012
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Nick Bostrom. 2014 · 2014
Earlier work this paper cites.
Concrete Problems in AI Safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016 · 2016
Earlier work this paper cites.
Faulty reward functions in the wild
Dario Amodei Jack Clark. 2016 · 2016
Earlier work this paper cites.
Campbell’s Law: implications for health care
Michael Poku. 2016 · 2016
Earlier work this paper cites.
Why Good Teaching Evaluations May Reward Bad Teaching: On Grade Inflation and Other Unintended Consequences of Student Evaluations
Wolfgang Stroebe. 2016 · 2016
Earlier work this paper cites.
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell. 2017 · 2017
Earlier work this paper cites.
Jan Leike, Miljan Martic, Victoria Krakovna, Pedro A. Ortega, Tom Everitt, Andrew Lefrancq, Laurent Orseau, and Shane Legg. 2017 · 2017
Earlier work this paper cites.
Medicine and the Mcnamara Fallacy
S O’Mahony. 2017 · 2017
Earlier work this paper cites.
Automated Classification of Skin Lesions: From Pixels to Practice
Akhila Narla, Brett Kuprel, Kavita Sarin, Roberto Novoa, and Justin Ko. 2018 · 2018
Earlier work this paper cites.
Growth, degrowth, and the challenge of artificial superintelligence
Salvador Pueyo. 2018 · 2018
Earlier work this paper cites.
Reframing Superintelligence
K Eric Drexler. 2019 · 2019
Cited alongside, same era.
Over-optimization of academic publishing metrics: observing Goodhart’s Law in action
Michael Fire and Carlos Guestrin. 2019 · 2019
Cited alongside, same era.
Incomplete Contracting and AI Alignment
Dylan Hadfield-Menell and Gillian K. Hadfield. 2019 · 2019
Cited alongside, same era.
Multiparty Dynamics and Failure Modes for Machine Learning and Artificial Intelligence
David Manheim. 2019 · 2019
Cited alongside, same era.
Categorizing Variants of Goodhart’s Law
David Manheim and Scott Garrabrant. 2019 · 2019
Cited alongside, same era.
Dissecting racial bias in an algorithm used to manage the health of populations
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. 2019 · 2019
Discovering Language Model Behaviors with Model-Written Evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
Dataset Shift in Machine Learning
Joaquin Quinonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D. Lawrence. 2022 · 2022
Later among the works it cites.
Defining and Characterizing Reward Gaming
Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. 2022 · 2022
Later among the works it cites.
Reliance on metrics is a fundamental challenge for AI
Rachel L. Thomas and David Uminsky. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human Compatible: AI and the Problem of Control
Stuart Russell. 2019 · 2019
Cited alongside, same era.
An unethical optimization principle
Nicholas Beale, Heather Battey, Anthony C. Davison, and Robert S. MacKay. 2020 · 2020
Cited alongside, same era.
The Alignment Problem – Machine Learning and Human Values
Brian Christian. 2020 · 2020
Cited alongside, same era.
Specification gaming examples in AI
Viktoria Krakovna. 2020 · 2020
Cited alongside, same era.
Specification gaming: the flip side of AI ingenuity
Viktoria Krakovna, Jonathan Uesato, Vlad Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg. 2020 · 2020
Cited alongside, same era.
The Precipice
Toby Ord. 2020 · 2020
Cited alongside, same era.
Parametrically Retargetable Decision-Makers Tend To Seek Power
Alexander Matt Turner and Prasad Tadepalli. 2022 · 2022
Later among the works it cites.
Simon Zhuang and Dylan Hadfield-Menell. 2021 · 2022
Later among the works it cites.
Will AI avoid exploitation? Artificial general intelligence and expected utility theory
Adam Bales. 2023 · 2023
Closest in time.
Taken out of context: On measuring situational awareness in LLMs
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans. 2023 · 2023
Closest in time.
Statement on AI Risk
Centre for AI Safety. 2023 · 2023
Closest in time.
There are no coherence theorems
EJT. 2023 · 2023
Closest in time.
Empirical evidence for existential AI risk factors
Rose Hadshar. 2023 · 2023
Closest in time.
An Overview of Catastrophic AI Risks
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside. 2023 · 2023
Closest in time.
Goodhart’s Law and Machine Learning: A Structural Perspective
Christopher A. Hennessy and Charles A. E. Goodhart. 2023 · 2023
Closest in time.
Dead rats, dopamine, performance metrics, and peacock tails: proxy failure is an inherent risk in goal-oriented systems
Yohan J. John, Leigh Caldwell, Dakota E. McCoy, and Oliver Braganza. 2023 · 2023
Closest in time.
Power-seeking can be probable and predictive for trained agents
Victoria Krakovna and Janos Kramar. 2023 · 2023
Closest in time.
Goal Misgeneralization in Deep Reinforcement Learning
Lauro Langosco, Jack Koch, Lee Sharkey, Jacob Pfau, Laurent Orseau, and David Krueger. 2023 · 2023
Closest in time.
Towards Out-Of-Distribution Generalization: A Survey
Jiashuo Liu, Zheyan Shen, Yue He, Xingxuan Zhang, Renzhe Xu, Han Yu, and Peng Cui. 2023 · 2023
Closest in time.
The alignment problem from a deep learning perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann. 2023 · 2023
Closest in time.
Goal Misgeneralisation: Why Correct Specifications Aren’t Enough For Correct Goals
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. 2023 · 2023
Closest in time.
AI offers significant opportunities but twelve governance challenges must be addressed says science, innovation and technology committee
Innovation UK Parliament’s Science and Technology Committee. 2023 · 2023
Closest in time.