Fetching the paper…
Reading the bibliography…
General Value Function (GVF) is a powerful tool to represent both the {\em predictive} and {\em retrospective} knowledge in reinforcement learning (RL).
Mutual fund performance
W. F. Sharpe · 1966
Earlier work this paper cites.
The variance of discounted Markov decision processes
M. J. Sobel · 1982
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Residual algorithms: reinforcement learning with function approximation
L. Baird · 1995
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Average cost temporal-difference learning
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
TD algorithm for the variance of return and mean-variance reinforcement learning
M. Sato, H. Kimura, and S. Kobayashi · 2001
Earlier work this paper cites.
Error bounds for approximate policy iteration
R. Munos · 2003
Earlier work this paper cites.
Temporal-difference networks
R. S. Sutton and B. Tanner · 2004
Earlier work this paper cites.
The grand challenge of predictive empirical abstract knowledge
R. S. Sutton · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora · 2009
Earlier work this paper cites.
Derivatives of logarithmic stationary distributions for policy gradient reinforcement learning
T. Morimura, E. Uchibe, J. Yoshimoto, J. Peters, and K. Doya · 2010
Earlier work this paper cites.
Risk-averse dynamic programming for markov decision processes
A. Ruszczyński · 2010
Earlier work this paper cites.
The fixed points of off-policy TD
J. Kolter · 2011
Earlier work this paper cites.
Gradient temporal-difference learning algorithms
H. R. Maei · 2011
Earlier work this paper cites.
Informing sequential clinical decision-making through reinforcement learning: an empirical study
S. M. Shortreed, E. Laber, D. J. Lizotte, T. S. Stroup, J. Pineau, and S. A. Murphy · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
Matrix analysis
R. A. Horn and C. R. Johnson · 2012
Earlier work this paper cites.
Policy gradients with variance related risk criteria
A. Tamar, D. Di Castro, and S. Mannor · 2012
Earlier work this paper cites.
Representation search through generate and test
A. R. Mahmood and R. S. Sutton · 2013
Earlier work this paper cites.
Algorithmic aspects of mean–variance optimization in markov decision processes
S. Mannor and J. N. Tsitsiklis · 2013
Earlier work this paper cites.
Better generalization with forecasts
T. Schaul and M. Ring · 2013
Cited alongside, same era.
Gradient temporal difference networks
D. Silver · 2013
Cited alongside, same era.
Variance adjusted actor critic algorithms
A. Tamar and S. Mannor · 2013
Cited alongside, same era.
H. Yao and D. Schuurmans · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Doubly robust bias reduction in infinite horizon off-policy estimation
Z. Tang, Y. Feng, L. Li, D. Zhou, and Q. Liu · 2019
Later among the works it cites.
Two time-scale off-policy TD learning: Non-asymptotic analysis over markovian samples
T. Xu, S. Zou, and Y. Liang · 2019
Later among the works it cites.
Generalized off-policy actor-critic
S. Zhang, W. Boehmer, and S. Whiteson · 2019
Later among the works it cites.
How to learn a useful critic? model-based action-gradient-estimator policy optimization
P. D’Oro and W. Jaśkowski · 2020
Later among the works it cites.
From importance sampling to doubly robust policy gradient
J. Huang and N. Jiang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Heess, G. Wayne, D. Silver, T. Lillicrap, Y. Tassa, and T. Erez · 2015
Cited alongside, same era.
Finite-sample analysis of proximal gradient TD algorithms
B. Liu, J. Liu, M. Ghavamzadeh, S. Mahadevan, and M. Petrik · 2015
Cited alongside, same era.
Developing a predictive approach to knowledge
A. White et al · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Cited alongside, same era.
Learning to filter with predictive state inference machines
W. Sun, A. Venkatraman, B. Boots, and J. A. Bagnell · 2016
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
R. S. Sutton, A. R. Mahmood, and M. White · 2016
Cited alongside, same era.
Learning the variance of the reward-to-go
A. Tamar, D. Di Castro, and S. Mannor · 2016
Cited alongside, same era.
M. Kaledin, E. Moulines, A. Naumov, V. Tadic, and H.-T. Wai · 2020
Later among the works it cites.
Statistically efficient off-policy policy gradients
N. Kallus and M. Uehara · 2020
Later among the works it cites.
First-order and Stochastic Optimization Methods for Machine Learning
G. Lan · 2020
Later among the works it cites.
Representation and general value functions
C. Sherstan · 2020
Later among the works it cites.
Improving sample complexity bounds for (natural) actor-critic algorithms
T. Xu, Z. Wang, and Y. Liang · 2020
Later among the works it cites.
Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms
T. Xu, Z. Wang, and Y. Liang · 2020
Later among the works it cites.
Variational policy gradient method for reinforcement learning with general utilities
J. Zhang, A. Koppel, A. S. Bedi, C. Szepesvari, and M. Wang · 2020
Later among the works it cites.
Gendice: generalized offline estimation of stationary values
R. Zhang, B. Dai, L. Li, and D. Schuurmans · 2020
Later among the works it cites.
GradientDICE: rethinking generalized offline estimation of stationary values
S. Zhang, B. Liu, and S. Whiteson · 2020
Later among the works it cites.
Provably convergent off-policy actor-critic with function approximation
S. Zhang, B. Liu, H. Yao, and S. Whiteson · 2020
Later among the works it cites.
Provably convergent two-timescale off-policy actor-critic with function approximation
S. Zhang, B. Liu, H. Yao, and S. Whiteson · 2020
Later among the works it cites.
Learning retrospective knowledge with reverse reinforcement learning
S. Zhang, V. Veeriah, and S. Whiteson · 2020
Later among the works it cites.
Universal off-policy evaluation
Y. Chandak, S. Niekum, B. C. da Silva, E. Learned-Miller, E. Brunskill, and P. S. Thomas · 2021
Closest in time.
Variance penalized on-policy and off-policy actor-critic
A. Jain, G. Patil, A. Jain, K. Khetarpal, and D. Precup · 2021
Closest in time.
Sample complexity bounds for two timescale value-based reinforcement learning algorithms
T. Xu and Y. Liang · 2021
Closest in time.
Doubly robust off-policy actor-critic: convergence and optimality, 2021
T. Xu, Z. Yang, Z. Wang, and Y. Liang · 2021
Closest in time.