Fetching the paper…
Reading the bibliography…
We introduce three concepts that describe an agent's incentives: response incentives indicate which variables in the environment, such as sensitive demographic information, affect the decision under the optimal policy.
Understanding agent incentives using causal influence diagrams, part i: single action settings
Tom Everitt, Pedro A Ortega, Elizabeth Barnes, and Shane Legg · 1902
Earlier work this paper cites.
Modeling agi safety frameworks with causal influence diagrams
Tom Everitt, Ramana Kumar, Victoria Krakovna, and Shane Legg · 1906
Earlier work this paper cites.
Information value theory
Ronald A Howard · 1966
Earlier work this paper cites.
Evaluating influence diagrams
Ross D Shachter · 1986
Earlier work this paper cites.
Causal Networks: Semantics and Expressiveness
Thomas Verma and Judea Pearl · 1988
Earlier work this paper cites.
Temporal and modal logic
E Allen Emerson · 1990
Earlier work this paper cites.
From influence to relevance to knowledge
Ronald A Howard · 1990
Earlier work this paper cites.
Using influence diagrams to value information and control
James E Matheson · 1990
Earlier work this paper cites.
A decision-based view of causality
David Heckerman and Ross Shachter · 1994
Earlier work this paper cites.
Decision-theoretic foundations for causal reasoning
David Heckerman and Ross Shachter · 1995
Earlier work this paper cites.
Strong completeness and faithfulness in bayesian networks
C Meek · 1995
Earlier work this paper cites.
Axioms of causal relevance
David Galles and Judea Pearl · 1997
Earlier work this paper cites.
A note about redundancy in influence diagrams
Enrico Fagiuoli and Marco Zaffalon · 1998
Earlier work this paper cites.
Bayes-Ball: The Rational Pastime (for Determining Irrelevance and Requisite Information in Belief Networks and Influence Diagrams)
Ross D Shachter · 1998
Earlier work this paper cites.
Welldefined decision scenarios
Thomas D Nielsen and Finn V Jensen · 1999
Earlier work this paper cites.
Representing and solving decision problems with limited information
Steffen L Lauritzen and Dennis Nilsson · 2001
Earlier work this paper cites.
Direct and indirect effects
Judea Pearl · 2001
Earlier work this paper cites.
Influence diagrams for causal modelling and inference
A Philip Dawid · 2002
Earlier work this paper cites.
Identifiability of path-specific effects
Chen Avin, Ilya Shpitser, and Judea Pearl · 2005
Earlier work this paper cites.
Interventions and causal inference
Frederick Eberhardt and Richard Scheines · 2007
Earlier work this paper cites.
Algorithmic game theory, cambridge univ, 2007
N Nisan, T Roughgarden, E Tardos, and VV Vazirani · 2007
Earlier work this paper cites.
The basic AI drives
Stephen M Omohundro · 2008
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Cited alongside, same era.
Strategy logic
Krishnendu Chatterjee, Thomas A Henzinger, and Nir Piterman · 2010
Cited alongside, same era.
Pearl causality and the value of control
Ross Shachter and David Heckerman · 2010
Cited alongside, same era.
Jin Tian and Judea Pearl · 2013
Cited alongside, same era.
Inference of intention and permissibility in moral decision making
Max Kleiman-Weiner, Tobias Gerstenberg, Sydney Levine, and Joshua B Tenenbaum · 2015
Cited alongside, same era.
Corrigibility
Nate Soares, Benja Fallenstein, Stuart Armstrong, and Eliezer Yudkowsky · 2015
Cited alongside, same era.
Asymptotically unambitious artificial general intelligence
Michael K. Cohen, Badri N. Vellambi, and Marcus Hutter · 2020
Closest in time.
A calculus for stochastic interventions: Causal effect identification and surrogate experiments
Juan Correa and Elias Bareinboim · 2020
Closest in time.
Hidden incentives for auto-induced distributional shift
David Krueger, Tegan Maharaj, and Jan Leike · 2020
Closest in time.
Characterizing optimal mixed policies: Where to intervene and what to observe
Sanghack Lee and Elias Bareinboim · 2020
Closest in time.
Causal imitation learning with unobserved confounders
Junzhe Zhang, Daniel Kumor, and Elias Bareinboim · 2020
Closest in time.
Rational verification: game-theoretic verification of multi-agent systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ross D Shachter · 2016
Cited alongside, same era.
Quantilizers: A safer alternative to maximizers for limited optimization
Jessica Taylor · 2016
Cited alongside, same era.
Rational verification: From model checking to equilibrium checking
Michael Wooldridge, Julian Gutierrez, Paul Harrenstein, Enrico Marchioni, Giuseppe Perelli, and Alexis Toumi · 2016
Cited alongside, same era.
Good and safe uses of AI oracles
Stuart Armstrong · 2017
Cited alongside, same era.
Low impact artificial intelligences
Stuart Armstrong and Benjamin Levinstein · 2017
Cited alongside, same era.
The off-switch game
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart J Russell · 2017
Cited alongside, same era.
Alessandro Abate, Julian Gutierrez, Lewis Hammond, Paul Harrenstein, Marta Kwiatkowska, Muhammad Najib, Giuseppe Perelli, Thomas Steeples, and Michael Wooldridge · 2021
Closest in time.
Pycid: a python library for causal influence diagrams
James Fox, Tom Everitt, Ryan Carey, Eric Langlois, Alessandro Abate, and Michael Wooldridge · 2021
Closest in time.
Rational verification for probabilistic systems
Julian Gutierrez, Lewis Hammond, Anthony W Lin, Muhammad Najib, and Michael Wooldridge · 2021
Closest in time.
How rl agents behave when their actions are modified
Eric Langlois and Tom Everitt · 2021
Closest in time.
Why fair labels can yield unfair predictions: Graphical conditions for introduced unfairness
Carolyn Ashurst, Ryan Carey, Silvia Chiappa, and Tom Everitt · 2022
Closest in time.
Probabilistic evaluation of counterfactual queries
Alexander Balke and Judea Pearl · 2022
Closest in time.
Estimating and penalizing induced preference shifts in recommender systems
Micah D Carroll, Anca Dragan, Stuart Russell, and Dylan Hadfield-Menell · 2022
Closest in time.
Path-specific objectives for safer agent incentives
Sebastian Farquhar, Ryan Carey, and Tom Everitt · 2022
Closest in time.
Probabilistic model checking and autonomy
Marta Kwiatkowska, Gethin Norman, and David Parker · 2022
Closest in time.
Jonathan G Richens, Rory Beard, and Daniel H Thompson · 2022
Closest in time.
A complete criterion for value of information in soluble influence diagrams
Chris Van Merwijk, Ryan Carey, and Tom Everitt · 2022
Closest in time.
Human control: Definitions and algorithms
Ryan Carey and Tom Everitt · 2023
Closest in time.
Reasoning about causality in games
Lewis Hammond, James Fox, Tom Everitt, Ryan Carey, , Alessandro Abate, and Michael Wooldridge · 2023
Closest in time.
Discovering agents
Zachary Kenton, Ramana Kumar, Sebastian Farquhar, Jonathan Richens, Matt MacDermott, and Tom Everitt · 2023
Closest in time.
Personalized decision making–a conceptual introduction
Scott Mueller and Judea Pearl · 2023
Closest in time.
The reasons that agents act: Intention and instrumental goals
Francis Rhys Ward, Matt MacDermott, Francesco Belardinelli, Francesca Toni, and Tom Everitt · 2024
Closest in time.