Fetching the paper…
Reading the bibliography…
Within the context of video games the notion of perfectly rational agents can be undesirable as it leads to uninteresting situations, where humans face tough adversarial decision makers.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
A course in game theory
Martin J Osborne and Ariel Rubinstein · 1994
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin Puterman · 1994
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
Michael L Littman and Csaba Szepesvári · 1996
Earlier work this paper cites.
Reinforcement learning
R Sutton and A Barto · 1998
Earlier work this paper cites.
Friend or foe q-learning in general-sum games
Michael L Littman · 2001
Earlier work this paper cites.
The case for dynamic difficulty adjustment in games
Robin Hunicke · 2005
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
Hilbert J Kappen · 2005
Earlier work this paper cites.
Human-robot interaction: a survey
Michael A Goodrich and Alan C Schultz · 2007
Cited alongside, same era.
Reinforcement learning and dynamic programming using function approximators
Lucian Busoniu, Robert Babuska, Bart De Schutter, and Damien Ernst · 2010
Cited alongside, same era.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altün · 2010
Cited alongside, same era.
Risk sensitive path integral control
Bart van den Broek, Wim Wiegerinck, and Bert Kappen · 2010
Cited alongside, same era.
Path integral control and bounded rationality
Daniel A Braun, Pedro A Ortega, Evangelos Theodorou, and Stefan Schaal · 2011
Cited alongside, same era.
Information theory of decisions and actions
Naftali Tishby and Daniel Polani · 2011
Cited alongside, same era.
Thermodynamics as a theory of decision-making with information-processing costs
Pedro A Ortega and Daniel A Braun · 2013
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Later among the works it cites.
Planning with information-processing constraints and model uncertainty in markov decision processes
Jordi Grau-Moya, Felix Leibfried, Tim Genewein, and Daniel A Braun · 2016
Later among the works it cites.
Human decision-making under limited time
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Toward large-scale agent guidance in an urban taxi service
Lucas Agussurja and Hoong Chuin Lau · 2012
Cited alongside, same era.
Burn-in, bias, and the rationality of anchoring
Falk Lieder, Tom Griffiths, and Noah Goodman · 2012
Cited alongside, same era.
Trading value and information in mdps
Jonathan Rubin, Ohad Shamir, and Naftali Tishby · 2012
Cited alongside, same era.
Pedro A Ortega and Alan A Stocker · 2016
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
An information-theoretic optimality principle for deep reinforcement learning
Felix Leibfried, Jordi Grau-Moya, and Haitham Bou-Ammar · 2017
Later among the works it cites.
Decentralised learning in systems with many, many strategic agents
David Mguni, Joel Jennings, and Enrique Munoz de Cote · 2018
Closest in time.