Fetching the paper…
Reading the bibliography…
In a single-agent setting, reinforcement learning (RL) tasks can be cast into an inference problem by introducing a binary random variable o, which stands for the "optimality".
Iterative solution of games by fictitious play
George W. Brown · 1951
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
Rudolf Kalman · 1960
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Rational and convergent learning in stochastic games
Michael Bowling and Manuela Veloso · 2001
Earlier work this paper cites.
Reinforcement learning of coordination in cooperative multi-agent systems
Spiros Kapetanakis and Daniel Kudenko · 2002
Earlier work this paper cites.
Coordination in multiagent reinforcement learning: A bayesian approach
Georgios Chalkiadakis and Craig Boutilier · 2003
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
Hilbert J. Kappen · 2005
Earlier work this paper cites.
Probabilistic inference for solving discrete and continuous state markov decision processes
Marc Toussaint and Amos Storkey · 2006
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
Evangelos Theodorou, Jonas Buchli, and Stefan Schaal · 2010
Earlier work this paper cites.
Policy gradients in linearly-solvable mdps
Emanuel Todorov · 2010
Cited alongside, same era.
Bayesian policy search for multi-agent role discovery
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2010
Cited alongside, same era.
Variational policy search via trajectory optimization
Sergey Levine and Vladlen Koltun · 2013
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2013
Cited alongside, same era.
Taming the Noise in Reinforcement Learning via Soft Updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Cited alongside, same era.
Equivalence Between Policy Gradients and Soft Q-Learning
John Schulman, Xi Chen, and Pieter Abbeel · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rémi Munos, Nicolas Heess, and Martin A. Riedmiller · 2018
Later among the works it cites.
Balancing Two-Player Stochastic Games with Soft Q-Learning
Jordi Grau-Moya, Felix Leibfried, and Haitham Bou-Ammar · 2018
Later among the works it cites.
Composable Deep Reinforcement Learning for Robotic Manipulation
Tuomas Haarnoja, Vitchyr Pong, Aurick Zhou, Murtaza Dalal, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daumé III · 2016
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Cited alongside, same era.
Stochastic Neural Networks for Hierarchical Reinforcement Learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Cited alongside, same era.
Counterfactual Multi-Agent Policy Gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
Sergey Levine · 2018
Later among the works it cites.
Neil C Rabinowitz, Frank Perbet, H Francis Song, Chiyuan Zhang, SM Eslami, and Matthew Botvinick · 2018
Later among the works it cites.
Modeling others using oneself in multi-agent reinforcement learning
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus · 2018
Later among the works it cites.
QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Later among the works it cites.
Zheng Tian, Shihao Zou, Tim Warr, Lisheng Wu, and Jun Wang · 2018
Later among the works it cites.
Multiagent soft q-learning
Ermo Wei, Drew Wicke, David Freelan, and Sean Luke · 2018
Later among the works it cites.
Probabilistic recursive reasoning for multi-agent reinforcement learning
Ying Wen, Yaodong Yang, Rui Luo, Jun Wang, and Wei Pan · 2019
Closest in time.