Fetching the paper…
Reading the bibliography…
We study the sparse entropy-regularized reinforcement learning (ERL) problem in which the entropy term is a special form of the Tsallis entropy.
Dynamic Programming
Bellman, R · 1957
Earlier work this paper cites.
Possible generalization of Boltzmann-Gibbs statistics
Tsallis, C · 1988
Earlier work this paper cites.
Markov Decision Processes
Puterman, M · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. and Tsitsiklis, J · 1996
Earlier work this paper cites.
An Introduction to Reinforcement Learning
Sutton, R. and Barto, A · 1998
Earlier work this paper cites.
Nonlinear programming
Bertsekas, D · 1999
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R., McAllester, D., Singh, S., and Mansour, Y · 2000
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
Kappen, H · 2005
Earlier work this paper cites.
Linearly-solvable Markov decision problems
Todorov, E · 2006
Earlier work this paper cites.
Regularized policy iteration
Farahmand, A. M., Ghavamzadeh, M., Szepesvári, Cs., and Mannor, S · 2008
Earlier work this paper cites.
Regularized fitted Q-iteration for planning in continuous-space Markovian decision problems
Farahmand, A. M., Ghavamzadeh, M., Szepesvári, Cs., and Mannor, S · 2009
Cited alongside, same era.
Regularization and feature selection in least-squares temporal difference learning
Kolter, Z. and Ng, A · 2009
Cited alongside, same era.
Linear complementarity for regularized policy evaluation and improvement
Johns, J., Painter-Wakefield, C., and Parr, R · 2010
Cited alongside, same era.
Relative entropy policy search
Peters, J., Müling, K., and Altun, Y · 2010
Cited alongside, same era.
Policy gradients in linearly-solvable MDPs
Todorov, E · 2010
Cited alongside, same era.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Ziebart, B · 2010
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M., and Abbeel, P · 2015
Later among the works it cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Later among the works it cites.
G-learning: Taming the noise in reinforcement learning via soft update
Fox, R., Pakman, A., and Tishby, N · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
An alternative softmax operator for reinforcement learning
Asadi, K. and Littman, M · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finite-sample analysis of lasso-td
Ghavamzadeh, M., Lazaric, A., Munos, R., and Hoffman, M · 2011
Cited alongside, same era.
Dynamic policy programming
Azar, M., Gómez, V., and Kappen, H · 2012
Cited alongside, same era.
Constrained optimization and Lagrange multiplier methods
Bertsekas, D · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, David, Lever, Guy, Heess, Nicolas, Degris, Thomas, Wierstra, Daan, and Riedmiller, Martin · 2014
Cited alongside, same era.
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Later among the works it cites.
A unified view of entropy-regularized Markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Later among the works it cites.
PGQ: Combining policy gradient and Q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2017
Later among the works it cites.
Sparse Markov decision processes with causal sparse Tsallis entropy regularization for reinforcement learning
Lee, K., Choi, S., and Oh, S · 2018
Closest in time.
Trust-pcl: An off-policy trust region method for continuous control
Nachum, Ofir, Norouzi, Mohammad, Xu, Kelvin, and Schuurmans, Dale · 2018
Closest in time.