Fetching the paper…
Reading the bibliography…
As progress in AI continues to advance, it is important to know how advanced systems will make choices and in what ways they may fail.
The interpretation of interaction in contingency tables
Edward H Simpson · 1951
Earlier work this paper cites.
Solution of a problem of leon henkin1
Martin Hugo Löb · 1955
Earlier work this paper cites.
A formal theory of inductive inference. part i
Ray J Solomonoff · 1964
Earlier work this paper cites.
Specimen Theoriae Novae de Mensura Sortis [by] Daniel Bernouilli: Translated Into German and English
Daniel Bernoulli, Alfred Pringsheim, and Louise Sommer · 1967
Earlier work this paper cites.
Newcomb’s problem and two principles of choice
Robert Nozick · 1969
Earlier work this paper cites.
Newcomb’s problem and prisoners’ dilemma
Steven J Brams · 1975
Earlier work this paper cites.
Conditionalization, observation, and change of preference
Paul Teller · 1976
Earlier work this paper cites.
Prisoners’ dilemma is a newcomb problem
David Lewis · 1979
Earlier work this paper cites.
A subjectivist’s guide to objective chance
David Lewis · 1980
Earlier work this paper cites.
A pragmatic investigation of the necessity of laws, 1980
Causal Necessity · 1980
Earlier work this paper cites.
Lindley’s paradox
Glenn Shafer · 1982
Earlier work this paper cites.
Metatickles and the dynamics of deliberation
Ellery Eells · 1984
Earlier work this paper cites.
Reasons and persons
Derek Parfit · 1984
Earlier work this paper cites.
The logic of decision
Richard C Jeffrey · 1990
Earlier work this paper cites.
Implications of the copernican principle for our future prospects
J Richard Gott · 1993
Earlier work this paper cites.
The two-envelope paradox: A complete analysis?
David Chalmers · 1994
Earlier work this paper cites.
The two-envelope paradox
John Broome · 1995
Earlier work this paper cites.
Probability theory: the logic of science
Edwin T Jaynes · 1996
Earlier work this paper cites.
Richard bellman on the birth of dynamic programming
Stuart Dreyfus · 2002
Earlier work this paper cites.
Are we living in a computer simulation?
Nick Bostrom · 2003
Earlier work this paper cites.
Beauty and the bets
Christopher Hitchcock · 2004
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Marcus Hutter · 2004
Earlier work this paper cites.
Puzzles of anthropic reasoning resolved using full non-indexical conditioning
Radford M Neal · 2006
Earlier work this paper cites.
The basic ai drives
Stephen M Omohundro · 2008
Earlier work this paper cites.
Utility indifference
Stuart Armstrong · 2010
Earlier work this paper cites.
Putting a value on beauty
Rachael Briggs · 2010
Cited alongside, same era.
Stuart Armstrong · 2011
Cited alongside, same era.
Anthropic bias: Observation selection effects in science and philosophy
Nick Bostrom · 2013
Cited alongside, same era.
Evidence, decision and causality
Arif Ahmed · 2014
Cited alongside, same era.
Udt with known search order
Tsvi Benson-Tilsen · 2014
Cited alongside, same era.
Problems of self-reference in self-improving space-time embedded intelligence
Benja Fallenstein and Nate Soares · 2014
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Programmatically interpretable reinforcement learning
Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, and Swarat Chaudhuri · 2018
Later among the works it cites.
Exploring neural networks with activation atlases, 2019
Shan Carter, Zan Armstrong, Ludwig Schubert, Ian Johnson, and Chris Olah · 2019
Later among the works it cites.
Designing preferences, beliefs, and identities for artificial intelligence
Vincent Conitzer · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaj Sotala and Roman V Yampolskiy · 2014
Cited alongside, same era.
A dutch book against sleeping beauties who are evidential decision theorists
Vincent Conitzer · 2015
Cited alongside, same era.
A survey of research questions for robust and beneficial ai. future of life institute, 2015
D Dewey, S Russell, M Tegmark, et al · 2015
Cited alongside, same era.
Research priorities for robust and beneficial artificial intelligence
Stuart Russell, Daniel Dewey, and Max Tegmark · 2015
Cited alongside, same era.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Cited alongside, same era.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Abram Demski and Scott Garrabrant · 2019
Later among the works it cites.
Corrigibility with utility preservation
Koen Holtman · 2019
Later among the works it cites.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Later among the works it cites.
Multiparty dynamics and failure modes for machine learning and artificial intelligence
David Manheim · 2019
Later among the works it cites.
Human compatible: Artificial intelligence and the problem of control
Stuart Russell · 2019
Later among the works it cites.
Adversarial examples: Attacks and defenses for deep learning
Xiaoyong Yuan, Pan He, Qile Zhu, and Xiaolin Li · 2019
Later among the works it cites.
Reinforcement learning in newcomblike environments
James Bell, Linda Linsefors, Caspar Oesterheld, and Joar Skalse · 2020
Closest in time.
Dissolving confusion around functional decision theory, Jan 2020
Stephen Casper · 2020
Closest in time.
Procrastination paradoxes: the good, the bad, and the ugly, Aug 2020
Stephen Casper · 2020
Closest in time.
AI Research Considerations for Human Existential Safety (ARCHES)
Andrew Critch and David Krueger · 2020
Closest in time.
Self-explaining ai as an alternative to interpretable ai
Daniel C Elton · 2020
Closest in time.
Challenges and countermeasures for adversarial attacks on deep reinforcement learning
Inaam Ilahi, Muhammad Usama, Junaid Qadir, Muhammad Umar Janjua, Ala Al-Fuqaha, Dinh Thai Hoang, and Dusit Niyato · 2020
Closest in time.
Compositional explanations of neurons
Jesse Mu and Jacob Andreas · 2020
Closest in time.
Extracting money from causal decision theorists
Caspar Oesterheld and Vincent Conitzer · 2020
Closest in time.
The precipice: existential risk and the future of humanity, 2020
Toby Ord · 2020
Closest in time.
Avoiding negative side effects due to incomplete knowledge of ai systems
Sandhya Saisubramanian, Shlomo Zilberstein, and Ece Kamar · 2020
Closest in time.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Closest in time.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li · 2020
Closest in time.
My current take on counterfactuals, Apr 2021
Abram Demski · 2021
Closest in time.