Fetching the paper…
Reading the bibliography…
Causal models of agents have been used to analyse the safety aspects of machine learning systems.
Cybernetics: Circular causal and feedback mechanisms in biological and social systems. Transactions of the seventh conference
H. von Foerster, M. Mead, and H. Teuber, editors · 1951
Earlier work this paper cites.
An Introduction to Cybernetics
W. R. Ashby · 1956
Earlier work this paper cites.
Cybernetics: Or Control and Communication in the Animal and the Machine
N. Wiener · 1961
Earlier work this paper cites.
The intentional stance
D. C. Dennett · 1987
Earlier work this paper cites.
Intelligent agents: Theory and practice
M. Wooldridge and N. R. Jennings · 1995
Earlier work this paper cites.
Axiomatizing causal reasoning
J. Y. Halpern · 2000
Earlier work this paper cites.
Influence diagrams for causal modelling and inference
A. P. Dawid · 2002
Earlier work this paper cites.
Multi-agent influence diagrams for representing and solving games
D. Koller and B. Milch · 2003
Earlier work this paper cites.
On the number of experiments sufficient and in the worst case necessary to identify all causal relations among n variables
F. Eberhardt, C. Glymour, and R. Scheines · 2005
Earlier work this paper cites.
Bayesian networks and influence diagrams
U. B. Kjaerulff and A. L. Madsen · 2008
Earlier work this paper cites.
Ignorable information in multi-agent scenarios, 2008
B. Milch and D. Koller · 2008
Earlier work this paper cites.
The basic AI drives
S. M. Omohundro · 2008
Earlier work this paper cites.
Artificial intelligence as a positive and negative factor in global risk
E. Yudkowsky et al · 2008
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
Settable systems: An extension of pearl’s causal model with optimization, equilibrium, and learning
H. White and K. Chalak · 2009
Earlier work this paper cites.
Actual causation and the art of modeling
J. Y. Halpern and C. Hitchcock · 2010
Earlier work this paper cites.
Causal inference using the algorithmic markov condition
D. Janzing and B. Schölkopf · 2010
Earlier work this paper cites.
Information-geometric approach to inferring causal directions
D. Janzing, J. Mooij, K. Zhang, J. Lemeire, J. Zscheischler, P. Daniušis, B. Steudel, and B. Schölkopf · 2012
Earlier work this paper cites.
On causal and anticausal learning
B. Schölkopf, D. Janzing, J. Peters, E. Sgouritsa, K. Zhang, and J. Mooij · 2012
Cited alongside, same era.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Cited alongside, same era.
Superintelligence: Paths, Dangers, Strategies
N. Bostrom · 2014
Cited alongside, same era.
A systematic search for transiting planets in the k2 data
D. Foreman-Mackey, B. T. Montet, D. W. Hogg, T. D. Morton, D. Wang, and B. Schölkopf · 2015
Cited alongside, same era.
Graphs for margins of bayesian networks
R. J. Evans · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
D. Hadfield-Menell, S. J. Russell, P. Abbeel, and A. Dragan · 2016
Cited alongside, same era.
The ground of optimization
A. Flint · 2020
Later among the works it cites.
Causal feature learning for utility-maximizing agents
D. Kinney and D. Watson · 2020
Later among the works it cites.
AGI safety from first principles: Goals and Agency
R. Ngo · 2020
Later among the works it cites.
Foundations of structural causal models with cycles and latent variables
S. Bongers, P. Forré, J. Peters, and J. M. Mooij · 2021
Later among the works it cites.
Intelligence and unambitiousness using algorithmic information theory
M. K. Cohen, B. Vellambi, and M. Hutter · 2021
Later among the works it cites.
User tampering in reinforcement learning recommender systems
C. Evans and A. Kasirzadeh · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elements of causal inference: foundations and learning algorithms
J. Peters, D. Janzing, and B. Schölkopf · 2017
Cited alongside, same era.
Network propaganda: Manipulation, disinformation, and radicalization in American politics
Y. Benkler, R. Faris, and H. Roberts · 2018
Cited alongside, same era.
Constraint-based causal discovery for non-linear structural causal models with cycles and latent confounders
P. Forré and J. M. Mooij · 2018
Cited alongside, same era.
Towards formal definitions of blameworthiness, intention, and moral responsibility
J. Y. Halpern and M. Kleiman-Weiner · 2018
Cited alongside, same era.
Agents and devices: A relative definition of agency
L. Orseau, S. M. McGill, and S. Legg · 2018
Cited alongside, same era.
Towards the first adversarially robust neural network model on MNIST
L. Schott, J. Rauber, M. Bethge, and W. Brendel · 2018
Cited alongside, same era.
S. Garrabrant · 2021
Later among the works it cites.
Equilibrium refinements for multi-agent influence diagrams: Theory and practice
L. Hammond, J. Fox, T. Everitt, A. Abate, and M. Wooldridge · 2021
Later among the works it cites.
How RL agents behave when their actions are modified
E. Langlois and T. Everitt · 2021
Later among the works it cites.
Toward causal representation learning
B. Schölkopf, F. Locatello, S. Bauer, N. R. Ke, N. Kalchbrenner, A. Goyal, and Y. Bengio · 2021
Later among the works it cites.
Literature Review on Goal-Directedness
A. Shimi, M. Campolo, and J. Collman · 2021
Later among the works it cites.
What are you optimizing for? aligning recommender systems with human values
J. Stray, I. Vendrov, J. Nixon, S. Adler, and D. Hadfield-Menell · 2021
Later among the works it cites.
Why fair labels can yield unfair predictions: Graphical conditions for introduced unfairness
C. Ashurst, R. Carey, S. Chiappa, and T. Everitt · 2022
Closest in time.
Estimating and penalizing induced preference shifts in recommender systems
M. D. Carroll, A. Dragan, S. Russell, and D. Hadfield-Menell · 2022
Closest in time.
Path-specific objectives for safer agent incentives
S. Farquhar, R. Carey, and T. Everitt · 2022
Closest in time.
J. G. Richens, R. Beard, and D. H. Thompson · 2022
Closest in time.
Causality for machine learning
B. Schölkopf · 2022
Closest in time.