Fetching the paper…
Reading the bibliography…
Offline Reinforcement Learning (RL) is a promising approach for learning optimal policies in environments where direct exploration is expensive or unfeasible.
Learning in Embedded Systems
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Bayesian q-learning
Richard Dearden, Nir Friedman, and Stuart J. Russell · 1998
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Reinforcement learning with gaussian processes
Yaakov Engel, Shie Mannor, and Ron Meir · 2005
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Interval estimation for reinforcement-learning algorithms in continuous-state domains
Martha White and Adam White · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl Rasmussen · 2011
Earlier work this paper cites.
A bayesian sampling approach to exploration in reinforcement learning
John Asmuth, Lihong Li, Michael L. Littman, Ali Nouri, and David Wingate · 2012
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, and Aviv Tamar · 2015
Earlier work this paper cites.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip S. Thomas and Emma Brunskill · 2016
Cited alongside, same era.
A freely accessible critical care database mimic-iii
Johnson A. E. W, Pollard T. J., Shen L., Lehman L.-W. H., Feng M., Ghassemi M., Moody B., Szolovits P., Anthony Celi L., , and R. G. Mark · 2016
Cited alongside, same era.
The uncertainty bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Rémi Munos, and Volodymyr Mnih · 2017
Cited alongside, same era.
Reinforcement Learning: An Introduction
An optimistic perspective on offline reinforcement learning, 2019
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2019
Later among the works it cites.
Propagating uncertainty in reinforcement learning via wasserstein barycenters
Alberto Maria Metelli, Amarildo Likmeta, and Marcello Restelli · 2019
Later among the works it cites.
Benchmarking batch deep reinforcement learning algorithms, 2019
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning, 2019
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Richard S. Sutton and Andrew G. Barto · 2017
Cited alongside, same era.
Continuous state-space models for optimal sepsis treatment - a deep reinforcement learning approach
Aniruddh Raghu, Matthieu Komorowski, Leo Anthony Celi, Peter Szolovits, and Marzyeh Ghassemi · 2017
Cited alongside, same era.
Evaluating reinforcement learning algorithms in observational health settings
Omer Gottesman, Fredrik D. Johansson, Joshua Meier, Jack Dent, Donghun Lee, Srivatsan Srinivasan, Linying Zhang, Yi Ding, David Wihl, Xuefeng Peng, Jiayu Yao, Isaac Lage, Christopher Mosch, Li-Wei H. Lehman, Matthieu Komorowski, Aldo Faisal, Leo Anthony Celi, David A. Sontag, and Finale Doshi-Velez · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Cited alongside, same era.
Guidelines for reinforcement learning in healthcare
Omer Gottesman, Fredrik Johansson, Matthieu Komorowski, Aldo Faisal, David Sontag, Finale Doshi-Velez, and Leo Anthony Celi · 2019
Cited alongside, same era.
Scott Fujimoto, David Meger, and Doina Precup · 2019
Later among the works it cites.
Rl unplugged: Benchmarks for offline reinforcement learning, 2020
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gomez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel Mankowitz, Cosmin Paduraru, Gabriel Dulac-Arnold, Jerry Li, Mohammad Norouzi, Matt Hoffman, Ofir Nachum, George Tucker, Nicolas Heess, and Nando de Freitas · 2020
Closest in time.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Closest in time.
Morel : Model-based offline reinforcement learning, 2020
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Closest in time.
Informed bayesian t-tests
Quentin F. Gronau, Alexander Ly, and Eric-Jan Wagenmakers · 2020
Closest in time.