Fetching the paper…
Reading the bibliography…
Regularized MDPs serve as a smooth version of original MDPs.
Functional approximations and dynamic programming
Richard Bellman and Stuart Dreyfus · 1959
Earlier work this paper cites.
Modified policy iteration algorithms for discounted markov decision problems
Martin L Puterman and Moon Chirl Shin · 1978
Earlier work this paper cites.
Possible generalization of boltzmann-gibbs statistics
Constantino Tsallis · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Dynamic policy programming
Mohammad Gheshlaghi Azar, Vicen c · 2012
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Approximate modified policy iteration and its application to the game of tetris
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, and Matthieu Geist · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicen c · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft q-learning
John Schulman, Xi Chen, and Pieter Abbeel · 2017
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Later among the works it cites.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2019
Later among the works it cites.
On the convergence of approximate and regularized policy iteration schemes
Elena Smirnova and Elvis Dohmatob · 2019
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Softmax exploration strategies for multiobjective reinforcement learning
Peter Vamplew, Richard Dazeley, and Cameron Foale · 2017
Cited alongside, same era.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Sparse markov decision processes with causal sparse tsallis entropy regularization for reinforcement learning
Kyungjae Lee, Sungjoon Choi, and Songhwai Oh · 2018
Cited alongside, same era.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2019
Cited alongside, same era.
Understanding the impact of entropy on policy optimization
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2019
Cited alongside, same era.
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Later among the works it cites.
A regularized approach to sparse optimal policy in reinforcement learning
Wenhao Yang, Xiang Li, and Zhihua Zhang · 2019
Later among the works it cites.
A note on the linear convergence of policy gradient methods
Jalaj Bhandari and Daniel Russo · 2020
Closest in time.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2020
Closest in time.
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Closest in time.
Leverage the average: an analysis of regularization in rl
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Rémi Munos, and Matthieu Geist · 2020
Closest in time.