Fetching the paper…
Reading the bibliography…
This paper investigates estimating the variance of a temporal-difference learning agent's update target.
“The Variance of Discounted Markov Decision Processes”
M.. Sobel · 1982
Earlier work this paper cites.
“Learning to Predict by the Methods of Temporal Differences”
Richard Sutton · 1988
Earlier work this paper cites.
“TD Algorithm for the Variance of Return and Mean-Variance Reinforcement Learning”
M. Sato, H. Kimura and S. Kobayashi · 2001
Earlier work this paper cites.
“Autonomous Helicopter Flight via Reinforcement Learning”
Andrew. Ng, H. Kim, Michael. Jordan and Shankar Sastry · 2004
Earlier work this paper cites.
“Fast Gradient-descent Methods for Temporal-Difference Learning with Linear Function Approximation”
Richard Sutton, Hamid Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári and Eric Wiewiora · 2009
Earlier work this paper cites.
“Gradient Temporal-Difference Learning Algorithms”, 2011
Hamid Maei · 2011
Cited alongside, same era.
“Horde: A Scalable Real-time Architecture for Learning Knowledge from Unsupervised Sensorimotor Interaction”
R.. Sutton, J. Modayil, M. Delp, T. Degris, P.. Pilarski, A. White and D. Precup · 2011
Cited alongside, same era.
“Policy Gradients with Variance Related Risk Criteria”
Aviv Tamar, Dotan Di Castro and Shie Mannor · 2012
Cited alongside, same era.
“ADADELTA: An Adaptive Learning Rate Method”
M.. Zeiler · 2012
Cited alongside, same era.
“Actor-Critic Algorithms for Risk-Sensitive MDPs”
L.. Prashanth and Mohammad Ghavamzadeh · 2013
Cited alongside, same era.
“Variance Adjusted Actor Critic Algorithms”
Aviv Tamar and Shie Mannor · 2013
Later among the works it cites.
“On Convergence of Emphatic Temporal-Difference Learning”
Huizhen Yu · 2015
Later among the works it cites.
“Learning the Variance of the Reward-To-Go”
A. Tamar, D. Di and S. Mannor · 2016
Later among the works it cites.
“A Greedy Approach to Adapting the Trace Parameter for Temporal Difference Learning”
M. White and A. White · 2016
Later among the works it cites.
“Unifying Task Specification in Reinforcement Learning.”
M. White · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…