Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) can be used to learn treatment policies and aid decision making in healthcare.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau · 1910
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Neural fitted Q iteration – first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Truncated importance sampling
Edward L Ionides · 2008
Earlier work this paper cites.
A theoretical and empirical analysis of Expected Sarsa
Harm van Seijen, Hado van Hasselt, Shimon Whiteson, and Marco Wiering · 2009
Earlier work this paper cites.
PAC-Bayesian model selection for reinforcement learning
M Fard and Joelle Pineau · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Model selection in reinforcement learning
Amir-massoud Farahmand and Csaba Szepesvári · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Kevin P Murphy · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan Prescott Adams · 2012
Earlier work this paper cites.
An empirical analysis of off-policy learning in discrete mdps
Cosmin Păduraru, Doina Precup, Joelle Pineau, and Gheorghe Comănici · 2013
Earlier work this paper cites.
Policy evaluation with temporal differences: A survey and comparison
Christoph Dann, Gerhard Neumann, Jan Peters, et al · 2014
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
Travis Mandel, Yun-En Liu, Sergey Levine, Emma Brunskill, and Zoran Popovic · 2014
Earlier work this paper cites.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
The self-normalized estimator for counterfactual learning
Adith Swaminathan and Thorsten Joachims · 2015
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2017
Cited alongside, same era.
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units
Niranjani Prasad, Li-Fang Cheng, Corey Chivers, Michael Draugelis, and Barbara E Engelhardt · 2017
Cited alongside, same era.
Deep reinforcement learning for sepsis treatment
Aniruddh Raghu, Matthieu Komorowski, Imran Ahmed, Leo Celi, Peter Szolovits, and Marzyeh Ghassemi · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Empirical study of off-policy policy evaluation for reinforcement learning
Cameron Voloshin, Hoang M Le, Nan Jiang, and Yisong Yue · 2019
Later among the works it cites.
Reinforcement learning in healthcare: A survey
Chao Yu, Jiming Liu, and Shamim Nemati · 2019
Later among the works it cites.
GenDICE: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2019
Later among the works it cites.
Identifying distinct, effective treatments for acute hypotension with SODA-RL: Safely optimized diverse accurate reinforcement learning
Joseph Futoma, Muhammad A Masood, and Finale Doshi-Velez · 2020
Later among the works it cites.
MOReL: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Evaluating reinforcement learning algorithms in observational health settings
Omer Gottesman, Fredrik Johansson, Joshua Meier, Jack Dent, Donghun Lee, Srivatsan Srinivasan, Linying Zhang, Yi Ding, David Wihl, Xuefeng Peng, et al · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al · 2018
Cited alongside, same era.
The Artificial Intelligence Clinician learns optimal treatment strategies for sepsis in intensive care
Matthieu Komorowski, Leo A Celi, Omar Badawi, Anthony C Gordon, and A Aldo Faisal · 2018
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Cited alongside, same era.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Cited alongside, same era.
An empirical study of representation learning for reinforcement learning in healthcare
Taylor W Killian, Haoran Zhang, Jayakumar Subramanian, Mehdi Fatemi, and Marzyeh Ghassemi · 2020
Later among the works it cites.
Conservative Q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Batch reinforcement learning with hyperparameter gradients
Byungjun Lee, Jongmin Lee, Peter Vrancx, Dongho Kim, and Kee-Eung Kim · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Provably good batch off-policy reinforcement learning without great exploration
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Caglar Gulcehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2020
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
Aravind Rajeswaran, Igor Mordatch, and Vikash Kumar · 2020
Later among the works it cites.
Clinician-in-the-loop decision making: Reinforcement learning with near-optimal set-valued policies
Shengpu Tang, Aditya Modi, Michael Sjoding, and Jenna Wiens · 2020
Later among the works it cites.
Minimax weight and Q-function learning for off-policy evaluation
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Benchmarks for deep off-policy evaluation
Justin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker, Alexander Novikov, Mengjiao Yang, Michael R Zhang, Yutian Chen, Aviral Kumar, Cosmin Paduraru, et al · 2021
Closest in time.
Confident off-policy evaluation and selection through self-normalized importance weighting
Ilja Kuzborskij, Claire Vernade, András György, and Csaba Szepesvári · 2021
Closest in time.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Closest in time.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2062
Closest in time.