Fetching the paper…
Reading the bibliography…
Traditional models of rational action treat the agent as though it is cleanly separated from its environment, and can act on that environment from the outside.
“Risks from Learned Optimization in Advanced Machine Learning Systems”, 2019
Evan Hubinger et al · 1906
Earlier work this paper cites.
“Solution of a Problem of Leon Henkin”
Martin Löb · 1955
Earlier work this paper cites.
“Newcomb’s Problem and Two Principles of Choice”
Robert Nozick · 1969
Earlier work this paper cites.
“Problems of Monetary Management: The UK Experience”
Charles Goodhart · 1975
Earlier work this paper cites.
“Counterfactuals and Two Kinds of Expected Utility”
Allan Gibbard and William. Harper · 1978
Earlier work this paper cites.
“The View from Nowhere”
Thomas Nagel · 1986
Earlier work this paper cites.
“Rational Learning Leads to Nash Equilibrium”
Ehud Kalai and Ehud Lehrer · 1993
Earlier work this paper cites.
“Defining the Turing Jump”
Richard Shore and Theodore Slaman · 1999
Earlier work this paper cites.
“Universal Artificial Intelligence”, Texts in Theoretical Computer Science
Marcus Hutter · 2005
Earlier work this paper cites.
“The Optimizer’s Curse”
James. Smith and Robert. Winkler · 2006
Earlier work this paper cites.
“The Basic AI Drives”
Stephen. Omohundro · 2008
Earlier work this paper cites.
“Towards a New Decision Theory”
Wei Dai · 2009
Earlier work this paper cites.
“Counterfactual Mugging”
Vladimir Nesov · 2009
Earlier work this paper cites.
“Ontological Crises in Artificial Agents’ Value Systems”, 2011
Peter de · 2011
Earlier work this paper cites.
“Learning What to Value”
Daniel Dewey · 2011
Earlier work this paper cites.
“One Decade of Universal Artificial Intelligence”
Marcus Hutter · 2012
Earlier work this paper cites.
“Space-Time Embedded Intelligence”
Laurent Orseau and Mark Ring · 2012
Earlier work this paper cites.
“Building Phenomenological Bridges”
Rob Bensinger · 2013
Earlier work this paper cites.
“The 5-and-10 Problem and the Tiling Agents Formalism”, 2013
Benya Fallenstein · 2013
Earlier work this paper cites.
“Tiling Agents for Self-Modifying AI, and the Löbian Obstacle”, 2013
Eliezer Yudkowsky and Marcello Herreshoff · 2013
Earlier work this paper cites.
“Superintelligence”
Nick Bostrom · 2014
Earlier work this paper cites.
“Generative Adversarial Nets”
Ian Goodfellow et al · 2014
Cited alongside, same era.
“Vingean Reflection”, 2015
Benya Fallenstein and Nate Soares · 2015
Cited alongside, same era.
“Reflective Oracles: A Foundation for Game Theory in Artificial Intelligence”
Benya Fallenstein, Jessica Taylor and Paul. Christiano · 2015
Cited alongside, same era.
“An Introduction to Löb’s Theorem in MIRI Research”, 2015
Patrick LaVictoire · 2015
Cited alongside, same era.
“Formalizing Two Problems of Realistic World-Models”, 2015
Nate Soares · 2015
Cited alongside, same era.
“Questions of Reasoning Under Logical Uncertainty”, 2015
Nate Soares and Benya Fallenstein · 2015
Cited alongside, same era.
“Corrigibility”
“Decisions Are For Making Bad Outcomes Inconsistent”
Rob Bensinger · 2017
Later among the works it cites.
“Stable Pointers to Value: An Agent Embedded in Its Own Utility Function”
Abram Demski · 2017
Later among the works it cites.
“Reinforcement Learning with a Corrupted Reward Channel”
Tom Everitt et al · 2017
Later among the works it cites.
“Logical Updatelessness as a Robust Delegation Problem”
Scott Garrabrant · 2017
Later among the works it cites.
“Two Major Obstacles for Logical Inductor Decision Theory”
Scott Garrabrant · 2017
Later among the works it cites.
“Naturalized Induction—A Challenge for Evidential and Causal Decision Theory”
Caspar Oesterheld · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nate Soares, Benya Fallenstein, Eliezer Yudkowsky and Stuart Armstrong · 2015
Cited alongside, same era.
“Quantilizers: A Safer Alternative to Maximizers for Limited Optimization”
Jessica Taylor · 2015
Cited alongside, same era.
“Complexity of Value”
Eliezer Yudkowsky · 2015
Cited alongside, same era.
“Omnipotence Test for AI Safety”
Eliezer Yudkowsky · 2015
Cited alongside, same era.
“Ontology Identification”
Eliezer Yudkowsky · 2015
Cited alongside, same era.
“Optimization Daemons”
Eliezer Yudkowsky · 2015
Cited alongside, same era.
“Ensuring Smarter-Than-Human Intelligence Has A Positive Outcome”
Nate Soares · 2017
Later among the works it cites.
“Agent Foundations for Aligning Machine Intelligence with Human Interests”
Nate Soares and Benya Fallenstein · 2017
Later among the works it cites.
“Coherent Decisions Imply Consistent Utilities”
Eliezer Yudkowsky · 2017
Later among the works it cites.
“Non-Adversarial Principle”
Eliezer Yudkowsky · 2017
Later among the works it cites.
“Functional Decision Theory: A New Theory of Instrumental Rationality”, 2017
Eliezer Yudkowsky and Nate Soares · 2017
Later among the works it cites.
“Techniques for Optimizing Worst-Case Performance”
Paul Christiano · 2018
Later among the works it cites.
“Supervising Strong Learners by Amplifying Weak Experts”, 2018
Paul Christiano, Buck Shlegeris and Dario Amodei · 2018
Later among the works it cites.
“An Untrollable Mathematician Illustrated”
Abram Demski · 2018
Later among the works it cites.
“Toward a New Technical Explanation of Technical Explanation”
Abram Demski · 2018
Later among the works it cites.
“Counterfactual Mugging Poker Game”
Scott Garrabrant · 2018
Later among the works it cites.
“Optimization Amplifies”
Scott Garrabrant · 2018
Later among the works it cites.
“Robustness to Scale”
Scott Garrabrant · 2018
Later among the works it cites.
“Categorizing Variants of Goodhart’s Law”, 2018
David Manheim and Scott Garrabrant · 2018
Later among the works it cites.
“The Value Learning Problem”
Nate Soares · 2018
Later among the works it cites.
“The Rocket Alignment Problem”
Eliezer Yudkowsky · 2018
Later among the works it cites.