Fetching the paper…
Reading the bibliography…
Standard reinforcement learning methods aim to master one way of solving a task whereas there may exist multiple near-optimal policies.
A new simplified acute physiology score (saps ii) based on a european/north american multicenter study
Jean-Roger Le Gall, Stanley Lemeshow, and Fabienne Saulnier · 1993
Earlier work this paper cites.
The sofa (sepsis-related organ failure assessment) score to describe organ dysfunction/failure
J-L Vincent, Rui Moreno, Jukka Takala, Sheila Willatts, Arnaldo De Mendonça, Hajo Bruining, CK Reinhart, PeterM Suter, and LG Thijs · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
A kernel method for the two-sample-problem
Arthur Gretton, Karsten M Borgwardt, Malte Rasch, Bernhard Schölkopf, and Alex J Smola · 2007
Earlier work this paper cites.
Domain independent approaches for finding diverse plans
Biplav Srivastava, Tuan Anh Nguyen, Alfonso Gerevini, Subbarao Kambhampati, Minh Binh Do, and Ivan Serina · 2007
Earlier work this paper cites.
Non-parametric estimation of integral probability metrics
Bharath K Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, and Gert RG Lanckriet · 2010
Earlier work this paper cites.
Non-deterministic policies in markovian decision processes
Mahdi Milani Fard and Joelle Pineau · 2011
Earlier work this paper cites.
Abandoning objectives: Evolution through the search for novelty alone
Joel Lehman and Kenneth O Stanley · 2011
Earlier work this paper cites.
Optimal kernel choice for large-scale two-sample tests
Arthur Gretton, Dino Sejdinovic, Heiko Strathmann, Sivaraman Balakrishnan, Massimiliano Pontil, Kenji Fukumizu, and Bharath K Sriperumbudur · 2012
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
A new severity of illness scale using a subset of acute physiology and chronic health evaluation data elements shows comparable predictive accuracy
Alistair E. W. Johnson, Andrew A Kramer, and Gari D Clifford · 2013
Earlier work this paper cites.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Earlier work this paper cites.
Multiobjective reinforcement learning: A comprehensive overview
Chunming Liu, Xin Xu, and Dewen Hu · 2014
Cited alongside, same era.
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, Aviv Tamar, et al · 2015
Cited alongside, same era.
Safe reinforcement learning
Philip S Thomas · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Finding diverse high-quality plans for hypothesis generation
Shirin Sohrabi, Anton V Riabov, Octavian Udrea, and Oktie Hassanzadeh · 2016
Cited alongside, same era.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Multi-bound tree search for logic-geometric programming in cooperative manipulation domains
Marc Toussaint and Manuel Lopes · 2017
Later among the works it cites.
Stein Points
W. Y. Chen, L. Mackey, J. Gorham, F.-X. Briol, and C. J. Oates · 2018
Later among the works it cites.
Diverse exploration for fast and safe policy improvement
Andrew Cohen, Lei Yu, and Robert Wright · 2018
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ex2: Exploration with exemplar models for deep reinforcement learning
Justin Fu, John Co-Reyes, and Sergey Levine · 2017
Cited alongside, same era.
Predicting intervention onset in the icu with switching state space models
Marzyeh Ghassemi, Mike Wu, Michael C Hughes, Peter Szolovits, and Finale Doshi-Velez · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
The mimic code repository: enabling reproducibility in critical care research
Alistair EW Johnson, David J Stone, Leo A Celi, and Tom J Pollard · 2017
Cited alongside, same era.
Stein variational policy gradient
Yang Liu, Prajit Ramachandran, Qiang Liu, and Jian Peng · 2017
Cited alongside, same era.
Tanmay Gangwani, Qiang Liu, and Jian Peng · 2018
Later among the works it cites.
Diversity-driven exploration strategy for deep reinforcement learning
Zhang-Wei Hong, Tzu-Yun Shann, Shih-Yang Su, Yi-Hsiang Chang, Tsu-Jui Fu, and Chun-Yi Lee · 2018
Later among the works it cites.
The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care
Matthieu Komorowski, Leo A Celi, Omar Badawi, Anthony C Gordon, and A Aldo Faisal · 2018
Later among the works it cites.
An inference-based policy gradient method for learning options
Matthew Smith, Herke Hoof, and Joelle Pineau · 2018
Later among the works it cites.
Diverse exploration via conjugate policies for policy gradient methods
Andrew Cohen, Xingye Qiao, Lei Yu, Elliot Way, and Xiangrong Tong · 2019
Closest in time.
pybullet-gym
Benjamin Ellenberge · 2019
Closest in time.