Fetching the paper…
Reading the bibliography…
We consider planning problems, that often arise in autonomous driving applications, in which an agent should decide on immediate actions so as to optimize a long term objective.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Dynamic programming and lagrange multipliers
Richard Bellman · 1956
Earlier work this paper cites.
Introduction to the mathematical theory of control processes , volume 2
Richard Bellman · 1971
Earlier work this paper cites.
Reinforcement learning in markovian and non-markovian environments
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
A survey of solution techniques for the partially observed markov decision process
Chelsea C White III · 1991
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Learning to play the game of chess
S. Thrun · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Cited alongside, same era.
Adversarial reinforcement learning
William Uther and Manuela Veloso · 1997
Cited alongside, same era.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Cited alongside, same era.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Cited alongside, same era.
A simple adaptive procedure leading to correlated equilibrium
S. HART and A. MAS-COLELL · 2000
Cited alongside, same era.
Reinforcement learning with long short-term memory
Bram Bakker · 2001
Cited alongside, same era.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Later among the works it cites.
If multi-agent learning is the answer, what is the question?
Yoav Shoham, Rob Powers, and Trond Grenager · 2007
Later among the works it cites.
Reinforcement Learning with Recurrent Neural Network
Anton Maximilian Schäfer · 2008
Later among the works it cites.
Algorithms for reinforcement learning
Csaba Szepesvári · 2010
Later among the works it cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Later among the works it cites.
Recurrent models of visual attention
Volodymyr Mnih, Nicolas Heess, Alex Graves, et al · 2014
Later among the works it cites.
Deterministic policy gradient algorithms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Cited alongside, same era.
R-max–a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2003
Cited alongside, same era.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Cited alongside, same era.
Theory and application of reward shaping in reinforcement learning
Adam Daniel Laud · 2004
Cited alongside, same era.
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Later among the works it cites.