Fetching the paper…
Reading the bibliography…
Off-policy reinforcement learning aims to leverage experience collected from prior policies for sample-efficient learning.
Striving for simplicity in off-policy deep reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 1907
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Àgata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard · 1907
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?, 1999
Stefan Schaal · 1999
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Error bounds for approximate policy iteration
Rémi Munos · 2003
Earlier work this paper cites.
Error bounds for approximate value iteration
Rémi Munos · 2005
Earlier work this paper cites.
Value-iteration based fitted policy iteration: Learning with a single trajectory
Andräs Antos, Csaa Szepesvari, and Remi Munos · 2007
Earlier work this paper cites.
The netflix prize
James Bennett, Stan Lanning, et al · 2007
Earlier work this paper cites.
Fitted q-iteration in continuous action-space mdps
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard S. Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Earlier work this paper cites.
A kernel two-sample test
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola · 2012
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
The importance of experience replay database composition in deep reinforcement learning
Tim de Bruin, Jens Kober, Karl Tuyls, and Robert Babuska · 2015
Cited alongside, same era.
Emphatic temporal-difference learning
A Rupam Mahmood, Huizhen Yu, Martha White, and Richard S Sutton · 2015
Cited alongside, same era.
Approximate modified policy iteration and its application to the game of tetris
Bruno Scherrer, Mohammad Ghavamzadeh, Victor Gabillon, Boris Lesner, and Matthieu Geist · 2015
Cited alongside, same era.
Trust region policy optimization
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, et al · 2018
Later among the works it cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Cited alongside, same era.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Reinforcement learning from imperfect demonstrations
Yang Gao, Huazhe Xu, Ji Lin, Fisher Yu, Sergey Levine, and Trevor Darrell · 2018
Cited alongside, same era.
BDD100K: A diverse driving video database with scalable annotation tooling
Fisher Yu, Wenqi Xian, Yingying Chen, Fangchen Liu, Mike Liao, Vashisht Madhavan, and Trevor Darrell · 2018
Later among the works it cites.
What is the effect of importance weighting in deep learning?
Jonathon Byrd and Zachary Lipton · 2019
Closest in time.
Diagnosing bottlenecks in deep q-learning algorithms
Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine · 2019
Closest in time.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G. Bellemare · 2019
Closest in time.
Soft q-learning with mutual-information regularization
Jordi Grau-Moya, Felix Leibfried, and Peter Vrancx · 2019
Closest in time.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Remi Tachet Des Combes · 2019
Closest in time.
Domain adaptation with asymmetrically-relaxed distribution alignment
Yifan Wu, Ezra Winston, Divyansh Kaushik, and Zachary Lipton · 2019
Closest in time.