Fetching the paper…
Reading the bibliography…
Learning from off-policy data is essential for sample-efficient reinforcement learning.
Steps toward artificial intelligence
Marvin Minsky · 1961
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Game theory
Drew Fudenberg and Jean Tirole · 1991
Earlier work this paper cites.
Reinforcement learning in continuous time: Advantage updating
Leemon C Baird · 1994
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard S Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2003
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Algorithms for reinforcement learning
Csaba Szepesvári · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan · 2015
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Deepmdp: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare · 2019
Later among the works it cites.
Understanding multi-step deep reinforcement learning: A systematic study of the dqn target
J Fernando Hernandez-Garcia and Richard S Sutton · 2019
Later among the works it cites.
Minatar: An atari-inspired testbed for thorough and reproducible reinforcement learning experiments
Kenny Young and Tian Tian · 2019
Later among the works it cites.
Disentangling causal effects for hierarchical reinforcement learning
Oriol Corcoll and Raul Vicente · 2020
Later among the works it cites.
The value equivalence principle for model-based reinforcement learning
Christopher Grimm, André Barreto, Satinder Singh, and David Silver · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Cited alongside, same era.
The reactor: A fast and sample-efficient actor-critic agent for reinforcement learning
Audrunas Gruslys, Will Dabney, Mohammad Gheshlaghi Azar, Bilal Piot, Marc Bellemare, and Remi Munos · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Multi-step reinforcement learning: A unifying algorithm
Kristopher De Asis, J Hernandez-Garcia, G Holland, and Richard Sutton · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Later among the works it cites.
Adaptive trade-offs in off-policy learning
Mark Rowland, Will Dabney, and Rémi Munos · 2020
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman · 2020
Later among the works it cites.
Planning in stochastic environments with a learned model
Ioannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K Hubert, and David Silver · 2021
Later among the works it cites.
Counterfactual credit assignment in model-free reinforcement learning
Thomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Thomas S Stepleton, Nicolas Heess, Arthur Guez, et al · 2021
Later among the works it cites.
Direct advantage estimation
Hsiao-Ru Pan, Nico Gürtler, Alexander Neitz, and Bernhard Schölkopf · 2022
Later among the works it cites.
The phenomenon of policy churn
Tom Schaul, André Barreto, John Quan, and Georg Ostrovski · 2022
Later among the works it cites.