Fetching the paper…
Reading the bibliography…
This report examines what I see as the core argument for concern about existential risk from misaligned artificial intelligence.
“Risks from Learned Optimization in Advanced Machine Learning Systems”
Evan Hubinger et al · 1906
Earlier work this paper cites.
“Risks from Learned Optimization in Advanced Machine Learning Systems” arXiv: 1906.01820
Evan Hubinger et al · 1906
Earlier work this paper cites.
“The Role of Cooperation in Responsible AI Development” arXiv: 1907.04534
Amanda Askell, Miles Brundage and Gillian Hadfield · 1907
Earlier work this paper cites.
“Dota 2 with Large Scale Deep Reinforcement Learning”
Christopher Berner et al · 1912
Earlier work this paper cites.
“AI Research Considerations for Human Existential Safety (ARCHES)”
Andrew Critch and David Krueger · 2006
Earlier work this paper cites.
“The Basic AI Drives”
Stephen. Omohundro · 2008
Earlier work this paper cites.
“Artificial Intelligence as a positive and negative factor in global risk”
Eliezer Yudkowsky · 2008
Earlier work this paper cites.
“Hidden Incentives for Auto-Induced Distributional Shift” arXiv: 2009.09153
David Krueger, Tegan Maharaj and Jan Leike · 2009
Earlier work this paper cites.
“REALab: An Embedded Perspective on Tampering” arXiv: 2011.08820
Ramana Kumar et al · 2011
Earlier work this paper cites.
“Avoiding Tampering Incentives in Deep RL via Decoupled Approval” arXiv: 2011.08827
Jonathan Uesato et al · 2011
Earlier work this paper cites.
“Thoughts on the Singularity Institute (SI) - LessWrong”
Holden Karnofsky · 2012
Earlier work this paper cites.
“Superintelligence: Paths, Dangers, Strategies” Google-Books-ID: 7_H8AwAAQBAJ
Nick Bostrom · 2014
Earlier work this paper cites.
“Stephen Hawking warns artificial intelligence could end mankind”
Rory Cellan-Jones · 2014
Earlier work this paper cites.
“Discontinuous progress investigation” Section: AI Timelines
Katja Grace, Rick Korzekwa, Asya Bergal and Daniel Kokotajlo · 2015
Earlier work this paper cites.
“The Secret of Our Success: How Culture Is Driving Human Evolution, Domesticating Our Species, and Making Us Smarter”
Joseph Henrich · 2015
Earlier work this paper cites.
“Concrete Problems in AI Safety”
Dario Amodei et al · 2016
Earlier work this paper cites.
“Why Tool AIs Want to Be Agent AIs”, 2016
Gwern Branwen · 2016
Earlier work this paper cites.
“Faulty Reward Functions in the Wild”
Jack Clark and Dario Amodei · 2016
Earlier work this paper cites.
“2016 Expert Survey on Progress in AI” Section: AI Timeline Surveys
Allan Dafoe, Baobao Zhang and Owain Evans · 2016
Earlier work this paper cites.
“Some Background on Our Views Regarding Advanced Artificial Intelligence”
Holden Karnofsky · 2016
Earlier work this paper cites.
“Multiple Stage Fallacy?”, 2016
Jeff Kaufman · 2016
Earlier work this paper cites.
“Mastering the game of Go with deep neural networks and tree search” Number: 7587 Publisher: Nature Publishing Group
David Silver et al · 2016
Earlier work this paper cites.
“Deep Reinforcement Learning from Human Preferences”
Paul Christiano et al · 2017
Cited alongside, same era.
“On the Impossibility of Supersized Machines”
Ben Garfinkel et al · 2017
Cited alongside, same era.
“Deal or No Deal? End-to-End Learning of Negotiation Dialogues”
Mike Lewis et al · 2017
Cited alongside, same era.
“2017 Report on Consciousness and Moral Patienthood”, 2018
Luke Muehlhauser · 2017
Cited alongside, same era.
“Multiverse-wide Cooperation via Correlated Decision Making”, 2017
Caspar Oesterheld · 2017
Cited alongside, same era.
“Life 3.0: Being Human in the Age of Artificial Intelligence” Google-Books-ID: 3_otDwAAQBAJ
Max Tegmark · 2017
Cited alongside, same era.
“Designing agent incentives to avoid side effects”
Victoria Krakovna, Ramana Kumar, Laurent Orseau and Alexander Turner · 2019
Later among the works it cites.
“Grandmaster level in StarCraft II using multi-agent reinforcement learning” Number: 7782 Publisher: Nature Publishing Group
Oriol Vinyals et al · 2019
Later among the works it cites.
“Emergent Tool Use From Multi-Agent Autocurricula”
Bowen Baker et al · 2020
Later among the works it cites.
“How Much Computational Power Does It Take to Match the Human Brain?”, 2020
Joseph Carlsmith · 2020
Later among the works it cites.
“The Alignment Problem: Machine Learning and Human Values” Google-Books-ID: VmJIzQEACAAJ
Brian Christian · 2020
Later among the works it cites.
“Draft report on AI timelines - AI Alignment Forum”, 2020
Ajeya Cotra · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Exploration by random network distillation”
Yuri Burda, Harrison Edwards, Amos Storkey and Oleg Klimov · 2018
Cited alongside, same era.
“Supervising strong learners by amplifying weak experts” arXiv: 1810.08575
Paul Christiano, Buck Shlegeris and Dario Amodei · 2018
Cited alongside, same era.
“Ben Garfinkel: How sure are we about this AI stuff?”, 2018
Ben Garfinkel · 2018
Cited alongside, same era.
“Viewpoint: When Will AI Exceed Human Performance? Evidence from AI Experts”
Katja Grace et al · 2018
Cited alongside, same era.
“AI safety via debate” arXiv: 1805.00899
Geoffrey Irving, Paul Christiano and Dario Amodei · 2018
Cited alongside, same era.
“Scalable agent alignment via reward modeling: a research direction” arXiv: 1811.07871
Jan Leike et al · 2018
Cited alongside, same era.
Later among the works it cites.
“The COVID-19 Pandemic and the $16 Trillion Virus”
David. Cutler and Lawrence. Summers · 2020
Later among the works it cites.
“Misalignment and misuse: whose values are manifest?” Section: Blog
Katja Grace · 2020
Later among the works it cites.
“The Flight to Safety-Critical AI”, 2020, pp. 42
Will Hunt · 2020
Later among the works it cites.
“Homogeneity vs. heterogeneity in AI takeoff scenarios”
Evan Hubinger and Kate Woolverton · 2020
Later among the works it cites.
“Specification gaming: the flip side of AI ingenuity”
Victoria Krakovna et al · 2020
Later among the works it cites.
“Zoom In: An Introduction to Circuits”
Chris Olah et al · 2020
Later among the works it cites.
“Andrew Critch on AI Research Considerations for Human Existential Safety”
Lucas Perry · 2020
Later among the works it cites.
“Evan Hubinger on Inner Alignment, Outer Alignment, and Proposals for Building Safe Advanced AI”
Lucas Perry · 2020
Later among the works it cites.
“Modeling the Human Trajectory”, 2020
David Roodman · 2020
Later among the works it cites.
“Mastering Atari, Go, chess and shogi by planning with a learned model” Number: 7839 Publisher: Nature Publishing Group
Julian Schrittwieser et al · 2020
Later among the works it cites.
“Clarifying “AI alignment””
Paul Christiano · 2021
Later among the works it cites.
“The case for aligning narrowly superhuman models - LessWrong”
Ajeya Cotra · 2021
Later among the works it cites.
“Report on Semi-informative Priors”, 2021
Tom Davidson · 2021
Later among the works it cites.
“Multimodal Neurons in Artificial Neural Networks”
Gabriel Goh et al · 2021
Later among the works it cites.
“Precipice”
Toby Ord · 2021
Later among the works it cites.
“Optimal Policies Tend To Seek Power”
Alexander Turner et al · 2021
Later among the works it cites.