Fetching the paper…
Reading the bibliography…
Effectively leveraging large, previously collected datasets in reinforcement learning (RL) is a key challenge for large-scale real-world applications.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard S Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Learning from scarce experience
Leonid Peshkin and Christian R Shelton · 2002
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2004
Earlier work this paper cites.
Robustness in markov decision problems with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Earlier work this paper cites.
Double q-learning
Hado V Hasselt · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Approximate policy iteration schemes: a comparison
Bruno Scherrer · 2014
Earlier work this paper cites.
Scaling up robust mdps using function approximation
Aviv Tamar, Shie Mannor, and Huan Xu · 2014
Earlier work this paper cites.
High confidence policy improvement
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Earlier work this paper cites.
Increasing the action gap: New operators for reinforcement learning
Marc G Bellemare, Georg Ostrovski, Arthur Guez, Philip S Thomas, and Rémi Munos · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Safe policy improvement by minimizing robust baseline regret
Marek Petrik, Mohammad Ghavamzadeh, and Yinlam Chow · 2016
Cited alongside, same era.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Consistent on-line off-policy evaluation
Assaf Hallak and Shie Mannor · 2017
Cited alongside, same era.
Towards characterizing divergence in deep q-learning
Joshua Achiam, Ethan Knight, and Pieter Abbeel · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2019
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Later among the works it cites.
Diagnosing bottlenecks in deep Q-learning algorithms
Justin Fu, Aviral Kumar, Matthew Soh, and Sergey Levine · 2019
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G Bellemare · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Romain Laroche, Paul Trichelair, and Rémi Tachet des Combes · 2017
Cited alongside, same era.
Variance-based regularization with convex objectives
Hongseok Namkoong and John C Duchi · 2017
Cited alongside, same era.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Mel Vecerik, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller · 2017
Cited alongside, same era.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos · 2018
Cited alongside, same era.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Hausman · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Learning self-correctable policies and value functions from demonstrations with negative sampling
Yuping Luo, Huazhe Xu, and Tengyu Ma · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Later among the works it cites.
Safe policy improvement with soft baseline bootstrapping
Kimia Nadjahi, Romain Laroche, and Rémi Tachet des Combes · 2019
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Later among the works it cites.
Safe policy improvement with an estimated baseline policy
Thiago D Simão, Romain Laroche, and Rémi Tachet des Combes · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
Discor: Corrective feedback in reinforcement learning via distribution correction
Aviral Kumar, Abhishek Gupta, and Sergey Levine · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Closest in time.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Closest in time.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, and Martin Riedmiller · 2020
Closest in time.
Q* approximation schemes for batch reinforcement learning: A eoretical comparison
Tengyang Xie and Nan Jiang · 2020
Closest in time.
Gendice: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Closest in time.