Fetching the paper…
Reading the bibliography…
We consider the problem of efficient credit assignment in reinforcement learning.
On transforming a certain class of stochastic processes by absolutely continuous substitution of measures
Igor Vladimirovich Girsanov · 1960
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W Moore and Christopher G Atkeson · 1993
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin Puterman · 1994
Earlier work this paper cites.
Learning without state estimation in partially observable environments
Satinder Singh, Tommi Jaakkola, and Michael I. Jordan · 1994
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Rare event simulation using Monte Carlo methods
Gerardo Rubino, Bruno Tuffin, et al · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Gradient estimation using stochastic computation graphs
John Schulman, Nicolas Heess, Theophane Weber, and Pieter Abbeel · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Woulda, coulda, shoulda: Counterfactually-guided policy search
Lars Buesing, Theophane Weber, Yori Zwols, Sébastien Racanière, Arthur Guez, Jean-Baptiste Lespiau, and Nicolas Heess · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Karol Hausman, Jost Tobias Springenberg, Ziyu Wang, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
Chia-Chun Hung, Timothy Lillicrap, Josh Abramson, Yan Wu, Mehdi Mirza, Federico Carnevale, Arun Ahuja, and Greg Wayne · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Model-based planning with discrete and continuous actions
Mikael Henaff, William F Whitney, and Yann LeCun · 2017
Cited alongside, same era.
Later among the works it cites.
Recall traces: Backtracking models for efficient reinforcement learning
Anirudh Goyal, Philemon Brakel, William Fedus, Soumye Singhal, Timothy Lillicrap, Sergey Levine, Hugo Larochelle, and Yoshua Bengio · 2019
Closest in time.
Hindsight policy gradients
Paulo Rauber, Avinash Ummadisingu, Filipe Mutz, and Jürgen Schmidhuber · 2019
Closest in time.
Credit assignment techniques in stochastic computation graphs
Théophane Weber, Nicolas Heess, Lars Buesing, and David Silver · 2019
Closest in time.