Fetching the paper…
Reading the bibliography…
This report examines whether advanced AIs that perform well in training will be doing so in order to gain power later -- a behavior I call "scheming" (also sometimes called "deceptive alignment").
“Risks from Learned Optimization in Advanced Machine Learning Systems” type: article, 2019
Evan Hubinger et al · 1906
Earlier work this paper cites.
R. McCoy, Junghyun Min and Tal Linzen · 1911
Earlier work this paper cites.
“Is SGD a Bayesian sampler? Well, almost” type: article, 2020
Chris Mingard, Guillermo Valle-Pérez, Joar Skalse and Ard. Louis · 2006
Earlier work this paper cites.
“Algorithmic complexity”
Marcus Hutter · 2008
Earlier work this paper cites.
“The Basic AI Drives”
Stephen. Omohundro · 2008
Earlier work this paper cites.
“Hidden Incentives for Auto-Induced Distributional Shift” type: article, 2020
David Krueger, Tegan Maharaj and Jan Leike · 2009
Earlier work this paper cites.
“Value is Fragile”
Eliezer Yudkowsky · 2009
Earlier work this paper cites.
“Superintelligence: Paths, Dangers, Strategies” Google-Books-ID: 7_H8AwAAQBAJ
Nick Bostrom · 2014
Earlier work this paper cites.
“Visual Information Theory”, 2015
Chris Olah · 2015
Earlier work this paper cites.
“What does the universal prior actually look like?”, 2016
Paul Christiano · 2016
Earlier work this paper cites.
“Parfit’s Hitchhiker”, 2016
Eliezer Yudkowsky · 2016
Earlier work this paper cites.
“Goodhart Taxonomy”
Scott Garrabrant · 2017
Earlier work this paper cites.
“Population based training of neural networks”, 2017
Max Jaderberg · 2017
Earlier work this paper cites.
“Relational inductive biases, deep learning, and graph networks”
Peter. Battaglia et al · 2018
Earlier work this paper cites.
“Deep Reinforcement Learning Doesn’t Work Yet”, 2018
Alex Irpan · 2018
Earlier work this paper cites.
Nils Reimers and Iryna Gurevych · 2018
Earlier work this paper cites.
“What failure looks like”
Paul Christiano · 2019
Earlier work this paper cites.
“Worst-case guarantees”, 2019
Paul Christiano · 2019
Earlier work this paper cites.
“The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks”
Jonathan Frankle and Michael Carbin · 2019
Earlier work this paper cites.
“Gradient hacking”
Evan Hubinger · 2019
Earlier work this paper cites.
“Inductive biases stick around”
Evan Hubinger · 2019
Earlier work this paper cites.
“Understanding “Deep Double Descent””
Evan Hubinger · 2019
Earlier work this paper cites.
“Comment on: Understanding “Deep Double Descent””, 2019
Rohin Shah · 2019
Earlier work this paper cites.
“AlphaStar: Mastering the real-time strategy game StarCraft II”, 2019
TheAlphaStar team · 2019
Earlier work this paper cites.
Guillermo Valle-Pérez, Chico. Camargo and Ard. Louis · 2019
Earlier work this paper cites.
“The Complete Reinforcement Learning Dictionary”, 2019
Shaked Zychlinski · 2019
Earlier work this paper cites.
“How Much Computational Power Does It Take to Match the Human Brain?”, 2020
Joseph Carlsmith · 2020
Earlier work this paper cites.
“Homogeneity vs. heterogeneity in AI takeoff scenarios”
Evan Hubinger · 2020
Earlier work this paper cites.
“Neural networks are fundamentally (almost) Bayesian”, 2020
Chris Mingard · 2020
Earlier work this paper cites.
“Does SGD Produce Deceptive Alignment?”
Mark Xu · 2020
Earlier work this paper cites.
“A General Language Assistant as a Laboratory for Alignment” type: article, 2021
Amanda Askell et al · 2021
Earlier work this paper cites.
“On the limits of idealized values”, 2023
Joe Carlsmith · 2021
Earlier work this paper cites.
“Is Power-Seeking AI an Existential Risk?” type: article, 2021
Joseph Carlsmith · 2021
Earlier work this paper cites.
“On the Universal Distribution”, 2022
Joseph Carlsmith · 2021
Earlier work this paper cites.
“Supplement to “Why AI alignment could be hard””, 2021
Ajeya Cotra · 2021
Earlier work this paper cites.
“Why AI alignment could be hard with modern deep learning”, 2021
Ajeya Cotra · 2021
Earlier work this paper cites.
“Understanding and controlling auto-induced distributional shift”
LRudL · 2021
Earlier work this paper cites.
“Deep Neural Networks are biased, at initialisation, towards simple functions”, 2021
Chris Mingard · 2021
Earlier work this paper cites.
“Comment on: Why Neural Networks Generalise, and Why They Are (Kind of) Bayesian”, 2021
Joar Skalse · 2021
Earlier work this paper cites.
“When Do Curricula Work?”
Xiaoxia Wu, Ethan Dyer and Behnam Neyshabur · 2021
Earlier work this paper cites.
“Strong Evidence is Common”, 2021
Mark Xu · 2021
Earlier work this paper cites.
“Comment on: A positive case for how we might succeed at prosaic AI alignment”, 2021
Eliezer Yudkowsky · 2021
Cited alongside, same era.
“Comment on: Why I’m excited about Debate”, 2021
Eliezer Yudkowsky · 2021
Cited alongside, same era.
“Ngo and Yudkowsky on alignment difficulty”
Eliezer Yudkowsky and Richard Ngo · 2021
Cited alongside, same era.
“The Speed + Simplicity Prior is probably anti-deceptive”
Anonymous · 2022
Cited alongside, same era.
“Propositions Concerning Digital Minds and Society”, 2022
Nick Bostrom and Carl Shulman · 2022
Cited alongside, same era.
“Discovering Latent Knowledge in Language Models Without Supervision”, 2022
Collin Burns, Haotian Ye, Dan Klein and Jacob Steinhardt · 2022
In Wikipedia , 2023
“Evolution of the eye” Page Version ID: 1184555315 · 2023
Closest in time.
“Improving the Welfare of AIs: A Nearcasted Proposal”, 2023
Ryan Greenblatt · 2023
Closest in time.
“When can we trust model evaluations?”
Evan Hubinger · 2023
Closest in time.
“Conditioning Predictive Models: Risks and Strategies” type: article, 2023
Evan Hubinger et al · 2023
Closest in time.
“Conditioning Predictive Models: Making inner alignment as easy as possible”
Evan Hubinger et al · 2023
Closest in time.
“Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research”
Evan Hubinger, Nicholas Schiefer, Carson Denison and Ethan Perez · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Shard Theory in Nine Theses: a Distillation and Critical Appraisal”
Lawrence Chan · 2022
Cited alongside, same era.
“Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover”
Ajeya Cotra · 2022
Cited alongside, same era.
“Clarifying wireheading terminology”, 2022
Leo Geo · 2022
Cited alongside, same era.
“Path dependence in ML inductive biases”
Vivek Hebbar and Evan Hubinger · 2022
Cited alongside, same era.
“A transparency and interpretability tech tree”
Evan Hubinger · 2022
Cited alongside, same era.
“How likely is deceptive alignment?”
Evan Hubinger · 2022
Cited alongside, same era.
In Wikipedia , 2023
“Huffman coding” Page Version ID: 1175907130 · 2023
Closest in time.
In Wikipedia , 2023
“Inclusive fitness” Page Version ID: 1149059235 · 2023
Closest in time.
In Wikipedia , 2023
“Inductive bias” Page Version ID: 1178867299 · 2023
Closest in time.
“How LLMs are and are not myopic”
Janus · 2023
Closest in time.
“3 levels of threat obfuscation”
Holden Karnofsky · 2023
Closest in time.
“Discussion with Nate Soares on a key alignment difficulty”
Holden Karnofsky · 2023
Closest in time.
“Goal Misgeneralization in Deep Reinforcement Learning” type: article, 2023
Lauro Langosco et al · 2023
Closest in time.
“Measuring Faithfulness in Chain-of-Thought Reasoning” type: article, 2023
Tamera Lanham et al · 2023
Closest in time.
“Self-exfiltration is a key dangerous capability”, 2023
Jan Leike · 2023
Closest in time.
In Wikipedia , 2023
“Lempel–Ziv complexity” Page Version ID: 1175701547 · 2023
Closest in time.
In Wikipedia , 2023
“Longtermism” Page Version ID: 1182123936 · 2023
Closest in time.
In Wikipedia , 2023
“Meta-learning (computer science)” Page Version ID: 1176940930 · 2023
Closest in time.
“The alignment problem from a deep learning perspective” type: article, 2023
Richard Ngo, Lawrence Chan and Sören Mindermann · 2023
Closest in time.
In Wikipedia , 2023
“Occam’s razor” Page Version ID: 1184501671 · 2023
Closest in time.
“AI Deception: A Survey of Examples, Risks, and Potential Solutions” type: article, 2023
Peter. Park et al · 2023
Closest in time.
“Carl Shulman (Pt 2) - AI Takeover, Bio & Cyber Attacks, Detecting Deception, & Humanity’s Far Future”, 2023
Dwarkesh Patel · 2023
Closest in time.
“Carl Shulman (Pt 1) - Intelligence Explosion, Primate Evolution, Robot Doublings, & Alignment”, 2023
Dwarkesh Patel and Carl Schulman · 2023
Closest in time.
“Playing the training game”, 2023
Kelsey Piper · 2023
Closest in time.
In Wikipedia , 2023
“Regularization (mathematics)” Page Version ID: 1177271127 · 2023
Closest in time.
“The situational awareness assumption in AI risk discourse, or why people should chill”
José Ricón · 2023
Closest in time.
“Preventing Language Models From Hiding Their Reasoning” type: article, 2023
Fabien Roger and Ryan Greenblatt · 2023
Closest in time.
In Wikipedia , 2023
“Rotating locomotion in living systems” Page Version ID: 1175051727 · 2023
Closest in time.
“GPT-4 architecture, datasets, costs and more leaked”, 2023
Maximilian Schreiner · 2023
Closest in time.
In Wikipedia , 2023
“Semiprime” Page Version ID: 1154324640 · 2023
Closest in time.
“Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy”
Buck Shlegeris and Ryan Greenblatt · 2023
Closest in time.
“Deep Deceptiveness”
Nate Soares · 2023
Closest in time.
“What I mean by "alignment is in large part about making cognition aimable at all"”
Nate Soares · 2023
Closest in time.
In Wiktionary, the free dictionary , 2023
“sphexish” Page Version ID: 75894036 · 2023
Closest in time.
In Wikipedia , 2023
“Universal Turing machine” Page Version ID: 1183200306 · 2023
Closest in time.
In Wikipedia , 2023
“Wabi-sabi” Page Version ID: 1184445631 · 2023
Closest in time.
“LLM Powered Autonomous Agents”, 2023
Lilian Weng · 2023
Closest in time.
“Deceptive Alignment is <1% Likely by Default”
David Wheaton · 2023
Closest in time.
In Wikipedia , 2023
“Wirehead (science fiction)” Page Version ID: 1177998718 · 2023
Closest in time.
In Wikipedia , 2023
“RSA numbers” Page Version ID: 1183413403 · 2048
Closest in time.