Fetching the paper…
Reading the bibliography…
Two main families of reinforcement learning algorithms, Q-learning and policy gradients, have recently been proven to be equivalent when using a softmax relaxation on one part, and an entropic regularization on the other.
Unifying divergence minimization and statistical inference via convex duality
Y. Altun and A. Smola · 2006
Earlier work this paper cites.
Optimal Transport : Old and New
C. Villani · 2008
Earlier work this paper cites.
Large Deviations Techniques and Applications
A. Dembo and O. Zeitouni · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
H. H. Bauschke and P. L. Combettes · 2011
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
A. Pakman R. Fox and N. Tishby · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
D. Silver A. A. Rusu J. Veness M. G. Bellemare A. Graves M. Riedmiller A. K. Fidjeland G. Ostrovski et al. V. Mnih, K. Kavukcuoglu · 2015
Cited alongside, same era.
Information Geometry and Its Applications
S. Amari · 2016
Cited alongside, same era.
Pgq : Combining policy gradient and q-learning
K. Kavukcuoglu B. O’Donoghue, R. Munos and V. Mnih · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
M. Mirza A. Graves T. P Lillicrap T. Harley D. Silver V. Mnih, A. Puigdomenech Badia and K. Kavukcuoglu · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
P. Moritz M. I. Jordan J. Schulman, S. Levine and P. Abbeel
Cited in the paper.
Bridging the gap between value and policy based reinforcement learning
K. Xu O. Nachum, M. Norouzi and D. Schuurmans
Cited in the paper.
Trust-pcl: An off-policy trust region method for continuous control
K. Xu O. Nachum, M. Norouzi and D. Schuurmans
Cited in the paper.
A unified view of entropy-regularized markov decision processes
V. Gomez G. Neu and A. Jonsson · 2017
Closest in time.
Equivalence between policy gradients and soft q-learning
X. Chen J. Schulman and P. Abbeel · 2017
Closest in time.
Deep relaxation: partial differential equations for optimizing deep neural networks
S. Osher S. Soatto P. Chaudhari, A. Oberman and G. Carlier · 2017
Closest in time.
Reinforcement learning with deep energy-based policies
P. Abbeel T. Haarnoja, H. Tang and S. Levine · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…