Meta-gradient reinforcement learning
Z. Xu, H. van Hasselt, and D. Silver · 2018
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
Original
J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, and S. Hochreiter · 2019
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
C.-C. Hung, T. P. Lillicrap, J. Abramson, Y. Wu, M. Mirza, F. Carnevale, A. Ahuja, and G. Wayne · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
H. van Hasselt, M. Hessel, and J. Aslanides · 2019
Later among the works it cites.
Forethought and hindsight in credit assignment
V. Chelu, D. Precup, and H. P. van Hasselt · 2020
Later among the works it cites.
Haiku: Sonnet for JAX, 2020
T. Hennigan, T. Cai, T. Norman, and I. Babuschkin · 2020
Later among the works it cites.
Optax: composable gradient transformation and optimisation, in jax!, 2020
M. Hessel, D. Budden, F. Viola, M. Rosca, E. Sezener, and T. Hennigan · 2020
Later among the works it cites.
Discor: Corrective feedback in reinforcement learning via distribution correction
Original
A. Kumar, A. Gupta, and S. Levine · 2020
Later among the works it cites.
Counterfactual credit assignment in model-free reinforcement learning
Original
T. Mesnard, T. Weber, F. Viola, S. Thakoor, A. Saade, A. Harutyunyan, W. Dabney, T. Stepleton, N. Heess, A. Guez, M. Hutter, L. Buesing, and R. Munos · 2020
Later among the works it cites.
Expected eligibility traces
Original
H. van Hasselt, S. Madjiheurem, M. Hessel, D. Silver, A. Barreto, and D. Borsa · 2020
Later among the works it cites.
A self-tuning actor-critic algorithm
T. Zahavy, Z. Xu, V. Veeriah, M. Hessel, J. Oh, H. P. van Hasselt, D. Silver, and S. Singh · 2020
Later among the works it cites.
Learning retrospective knowledge with reverse reinforcement learning
S. Zhang, V. Veeriah, and S. Whiteson · 2020
Later among the works it cites.
Preferential temporal difference learning, 2021
N. Anand and D. Precup · 2021
Later among the works it cites.
Learning expected emphatic traces for deep RL
Original
R. Jiang, S. Zhang, V. Chelu, A. White, and H. van Hasselt · 2021
Later among the works it cites.
Expected eligibility traces
H. van Hasselt, S. Madjiheurem, M. Hessel, D. Silver, A. Barreto, and D. Borsa · 2021
Later among the works it cites.