Fetching the paper…
Reading the bibliography…
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of the Bellman operator.
Dynamic Programming
Bellman, R · 1957
Earlier work this paper cites.
Convergence and finite-time behavior of simulated annealing
Mitra, D., Romeo, F., and Sangiovanni-Vincentelli, A · 1986
Earlier work this paper cites.
The role of exploration in learning control
Thrun, S. B · 1992
Earlier work this paper cites.
Tight performance bounds on greedy policies based on imperfect value functions
Williams, R. J. and Baird III, L. C · 1993
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Thrun, S. and Schwartz, A · 1994
Earlier work this paper cites.
Algorithms for Sequential Decision Making
Littman, M. L · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Linearly-solvable Markov decision problems
Todorov, E · 2007
Earlier work this paper cites.
Double Q-learning
van Hasselt, H · 2010
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Parameter estimation in softmax decision-making models with linear objective functions
Reverdy, P. and Leonard, N. E · 2016
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Learning to play in a day: Faster deep reinforcement learning by optimality tightening
He, F. S., Liu, Y., Schwing, A. G., and Peng, J · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Later among the works it cites.
A unified view of entropy-regularized Markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Later among the works it cites.
Combining policy gradient and Q-learning
O’Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and De Freitas, N · 2016
Cited alongside, same era.
Averaged-DQN: variance reduction and stabilization for deep reinforcement learning
Anschel, O., Baram, N., and Shimkin, N · 2017
Cited alongside, same era.
An alternative softmax operator for reinforcement learning
Asadi, K. and Littman, M. L · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Learning values across many orders of magnitude
van Hasselt, H. P., Guez, A., Hessel, M., Mnih, V., and Silver, D
Cited in the paper.
Schulman, J., Chen, X., and Abbeel, P · 2017
Later among the works it cites.
Distributional reinforcement learning with quantile regression
Dabney, W., Rowland, M., Bellemare, M. G., and Munos, R · 2018
Closest in time.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Closest in time.
Sparse Markov decision processes with causal sparse Tsallis entropy regularization for reinforcement learning
Lee, K., Choi, S., and Oh, S · 2018
Closest in time.
Deep reinforcement learning with double Q-learning
van Hasselt, H., Guez, A., and Silver, D · 2094
Closest in time.