Fetching the paper…
Reading the bibliography…
We establish a new connection between value and policy based reinforcement learning (RL) based on a relationship between softmax temporal value consistency and policy optimality under entropy regularization.
Learning from delayed rewards
C. J. Watkins · 1989
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
R. J. Williams and J. Peng · 1991
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Dynamic Programming and Optimal Control
D. P. Bertsekas · 1995
Earlier work this paper cites.
Temporal difference learning and TD-gammon
G. Tesauro · 1995
Earlier work this paper cites.
Algorithms for sequential decision making
M. L. Littman · 1996
Earlier work this paper cites.
Incremental multi-step Q-learning
J. Peng and R. J. Williams · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, Y. Mansour, et al · 1999
Earlier work this paper cites.
Convex Analysis and Nonlinear Optimization
J. Borwein and A. Lewis · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup · 2000
Earlier work this paper cites.
A natural policy gradient
S. Kakade · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
D. Precup, R. S. Sutton, and S. Dasgupta · 2001
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
H. J. Kappen · 2005
Earlier work this paper cites.
Linearly-solvable Markov decision problems
E. Todorov · 2006
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Cited alongside, same era.
Relative entropy policy search
J. Peters, K. Müling, and Y. Altun · 2010
Cited alongside, same era.
Policy gradients in linearly-solvable MDPs
E. Todorov · 2010
Cited alongside, same era.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Cited alongside, same era.
Dynamic policy programming with function approximation
M. G. Azar, V. Gómez, and H. J. Kappen · 2011
Cited alongside, same era.
G-learning: Taming the noise in reinforcement learning via soft updates
R. Fox, A. Pakman, and N. Tishby · 2016
Later among the works it cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
S. Gu, E. Holly, T. Lillicrap, and S. Levine · 2016
Later among the works it cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Later among the works it cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamic policy programming
M. G. Azar, V. Gómez, and H. J. Kappen · 2012
Cited alongside, same era.
Optimal control as a graphical model inference problem
M. G. Azar, V. Gómez, and H. J. Kappen · 2012
Cited alongside, same era.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Approximate maxent inverse optimal control and its application for mental simulation of human interactions
D.-A. Huang, A.-m. Farahmand, K. M. Kitani, and J. A. Bagnell · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Later among the works it cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Z. Wang, N. de Freitas, and M. Lanctot · 2016
Later among the works it cites.
The reactor: A sample-efficient actor-critic architecture
A. Gruslys, M. G. Azar, M. G. Bellemare, and R. Munos · 2017
Closest in time.
Q-Prop: Sample-efficient policy gradient with an off-policy critic
S. Gu, T. Lillicrap, Z. Ghahramani, R. E. Turner, and S. Levine · 2017
Closest in time.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Closest in time.
Improving policy gradient by exploring under-appreciated rewards
O. Nachum, M. Norouzi, and D. Schuurmans · 2017
Closest in time.
PGQ: Combining policy gradient and Q-learning
B. O’Donoghue, R. Munos, K. Kavukcuoglu, and V. Mnih · 2017
Closest in time.
Equivalence between policy gradients and soft Q-learning
J. Schulman, X. Chen, and P. Abbeel · 2017
Closest in time.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 2017
Closest in time.
Sample efficient actor-critic with experience replay
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas · 2017
Closest in time.