Fetching the paper…
Reading the bibliography…
Many reinforcement learning (RL) applications have combinatorial action spaces, where each action is a composition of sub-actions.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau · 1910
Earlier work this paper cites.
An upper bound on the loss from approximate optimal-value functions
Satinder P Singh and Richard C Yee · 1994
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Craig Boutilier · 1996
Earlier work this paper cites.
Computing factored value functions for policies in structured MDPs
Daphne Koller and Ronald Parr · 1999
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Efficient solution algorithms for factored MDPs
Carlos Guestrin, Daphne Koller, Ronald Parr, and Shobha Venkataraman · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Reinforcement learning with factored states and actions
Brian Sallans and Geoffrey E Hinton · 2004
Earlier work this paper cites.
Clinical data based optimal STI strategies for HIV: a reinforcement learning approach
Damien Ernst, Guy-Bart Stan, Jorge Goncalves, and Louis Wehenkel · 2006
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J Walsh, and Michael L Littman · 2006
Earlier work this paper cites.
Efficient structure learning in factored-state MDPs
Alexander L Strehl, Carlos Diuk, and Michael L Littman · 2007
Earlier work this paper cites.
Effect of voriconazole and fluconazole on the pharmacokinetics of intravenous fentanyl
Teijo I Saari, Kari Laine, Mikko Neuvonen, Pertti J Neuvonen, and Klaus T Olkkola · 2008
Earlier work this paper cites.
Early administration of norepinephrine increases cardiac preload and cardiac output in septic patients with life-threatening hypotension
Olfa Hamzaoui, Jean-François Georger, Xavier Monnet, Hatem Ksouri, Julien Maizel, Christian Richard, and Jean-Louis Teboul · 2010
Earlier work this paper cites.
Norepinephrine increases cardiac preload and reduces preload dependency assessed by passive leg raising in septic shock patients
Xavier Monnet, Julien Jabot, Julien Maizel, Christian Richard, and Jean-Louis Teboul · 2011
Earlier work this paper cites.
Efficient solutions to factored MDPs with imprecise transition probabilities
Karina Valdivia Delgado, Scott Sanner, and Leliane Nunes De Barros · 2011
Earlier work this paper cites.
Drug–drug interactions in the medical intensive care unit: an assessment of frequency, severity and the medications involved
Pamela L Smithburger, Sandra L Kane-Gill, and Amy L Seybert · 2012
Earlier work this paper cites.
Planning in factored action spaces with symbolic dynamic programming
Aswin Raghavan, Saket Joshi, Alan Fern, Prasad Tadepalli, and Roni Khardon · 2012
Earlier work this paper cites.
Independent reinforcement learners in cooperative Markov games: a survey regarding coordination problems
Laetitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat · 2012
Earlier work this paper cites.
Symbolic opportunistic policy iteration for factored-action mdps
Aswin Raghavan, Roni Khardon, Alan Fern, and Prasad Tadepalli · 2013
Earlier work this paper cites.
Near-optimal reinforcement learning in factored MDPs
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
Introductory econometrics: A modern approach
Jeffrey M Wooldridge · 2015
Earlier work this paper cites.
Effects of passive leg raising and volume expansion on mean systemic pressure and venous return in shock in humans
Laurent Guérin, Jean-Louis Teboul, Romain Persichini, Martin Dres, Christian Richard, and Xavier Monnet · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Sepsis: pathophysiology and clinical management
Jeffrey Gotts and Michael Matthay · 2016
Cited alongside, same era.
MIMIC-III, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Cited alongside, same era.
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units
Niranjani Prasad, Li Fang Cheng, Corey Chivers, Michael Draugelis, and Barbara E Engelhardt · 2017
Cited alongside, same era.
Sahil Sharma, Aravind Suresh, Rahul Ramesh, and Balaraman Ravindran · 2017
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Later among the works it cites.
Transfer learning from well-curated to less-resourced populations with HIV
Sonali Parbhoo, Mario Wieser, Volker Roth, and Finale Doshi-Velez · 2020
Later among the works it cites.
Clinician-in-the-loop decision making: Reinforcement learning with near-optimal set-valued policies
Shengpu Tang, Aditya Modi, Michael Sjoding, and Jenna Wiens · 2020
Later among the works it cites.
An empirical study of representation learning for reinforcement learning in healthcare
Taylor W Killian, Haoran Zhang, Jayakumar Subramanian, Mehdi Fatemi, and Marzyeh Ghassemi · 2020
Later among the works it cites.
Q-learning in enormous action spaces via amortized approximate maximization
Tom Van de Wiele, David Warde-Farley, Andriy Mnih, and Volodymyr Mnih · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multiagent cooperation and competition with deep reinforcement learning
Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente · 2017
Cited alongside, same era.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni · 2017
Cited alongside, same era.
Hybrid reward architecture for reinforcement learning
Harm Van Seijen, Mehdi Fatemi, Joshua Romoff, Romain Laroche, Tavian Barnes, and Jeffrey Tsang · 2017
Cited alongside, same era.
The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care
Matthieu Komorowski, Leo A Celi, Omar Badawi, Anthony C Gordon, and A Aldo Faisal · 2018
Cited alongside, same era.
Action branching architectures for deep reinforcement learning
Arash Tavakoli, Fabio Pardo, and Petar Kormushev · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z. Leibo, Karl Tuyls, and Thore Graepel · 2018
Cited alongside, same era.
QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
On the Rademacher complexity of linear hypothesis sets
Pranjal Awasthi, Natalie Frank, and Mehryar Mohri · 2020
Later among the works it cites.
Learning to represent action values as a hypergraph on the action vertices
Arash Tavakoli, Mehdi Fatemi, and Petar Kormushev · 2021
Later among the works it cites.
Risk bounds and Rademacher complexity in batch reinforcement learning
Yaqi Duan, Chi Jin, and Zhiyuan Li · 2021
Later among the works it cites.
Combining fluids and vasopressors: A magic potion?
Olfa Hamzaoui · 2021
Later among the works it cites.
Model selection for offline reinforcement learning: Practical considerations for healthcare settings
Shengpu Tang and Jenna Wiens · 2021
Later among the works it cites.
Surviving sepsis campaign: international guidelines for management of sepsis and septic shock 2021
Laura Evans, Andrew Rhodes, Waleed Alhazzani, Massimo Antonelli, Craig M Coopersmith, Craig French, Flávia R Machado, Lauralyn Mcintyre, Marlies Ostermann, Hallie C Prescott, et al · 2021
Later among the works it cites.
Causal Markov decision processes: Learning good interventions efficiently
Yangyi Lu, Amirhossein Meisami, and Ambuj Tewari · 2021
Later among the works it cites.
Factored action spaces in deep reinforcement learning, 2021
Thomas Pierrot, Valentin Macé, Jean-Baptiste Sevestre, Louis Monier, Alexandre Laterre, Nicolas Perrin, Karim Beguir, and Olivier Sigaud · 2021
Later among the works it cites.
Factored policy gradients: Leveraging structure for efficient learning in MOMDPs
Thomas Spooner, Nelson Vadori, and Sumitra Ganesh · 2021
Later among the works it cites.
Medical dead-ends and learning to identify high-risk states and treatments
Mehdi Fatemi, Taylor W. Killian, Jayakumar Subramanian, and Marzyeh Ghassemi · 2021
Later among the works it cites.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Later among the works it cites.
Causally motivated shortcut removal using auxiliary labels
Maggie Makar, Ben Packer, Dan Moldovan, Davis Blalock, Yoni Halpern, and Alexander D’Amour · 2022
Later among the works it cites.
Avoiding overfitting to the importance weights in offline policy optimization, 2022
Yao Liu and Emma Brunskill · 2022
Later among the works it cites.
Off-policy evaluation for large action spaces via embeddings
Yuta Saito and Thorsten Joachims · 2022
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason D Lee · 2022
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2062
Closest in time.