Fetching the paper…
Reading the bibliography…
While learning in an unknown Markov Decision Process (MDP), an agent should trade off exploration to discover new information about the MDP, and exploitation of the current knowledge to maximize the reward.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Dynamic programming and optimal control. Vol II
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham M. Kakade and John Langford · 2002
Earlier work this paper cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Peter L. Bartlett and Ambuj Tewari · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Approximate policy iteration: A survey and some new methods
Dimitri P Bertsekas · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
L. Bottou, J. Peters, J. Quinonero-Candela, D. Charles, D. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson · 2013
Earlier work this paper cites.
Thompson sampling for learning parameterized markov decision processes
Aditya Gopalan and Shie Mannor · 2015
Earlier work this paper cites.
Bayesian incentive-compatible bandit exploration
Yishay Mansour, Aleksandrs Slivkins, and Vasilis Syrgkanis · 2015
Cited alongside, same era.
Counterfactual risk minimization: Learning from logged bandit feedback
A. Swaminathan and T. Joachims · 2015
Cited alongside, same era.
Condition-based maintenance for complex systems: Coordinating maintenance and logistics planning for the process industries
Minou Catharina Anselma Olde Keizer · 2016
Cited alongside, same era.
Posterior sampling for reinforcement learning without episodes
Ian Osband and Benjamin Van Roy · 2016
Cited alongside, same era.
Safe policy improvement by minimizing robust baseline regret
M. Petrik, M. Ghavamzadeh, and Y. Chow · 2016
Cited alongside, same era.
Conservative bandits
Learning unknown markov decision processes: A thompson sampling approach
Yi Ouyang, Mukul Gagrani, Ashutosh Nayyar, and Rahul Jain · 2017
Later among the works it cites.
A lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar A. Duéñez-Guzmán, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sébastien Bubeck, and Michael I. Jordan · 2018
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I. Jordan · 2019
Later among the works it cites.
Conservative exploration using interleaving
Sumeet Katariya, Branislav Kveton, Zheng Wen, and Vamsi K. Potluru · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yifan Wu, Roshan Shariff, Tor Lattimore, and Csaba Szepesvári · 2016
Cited alongside, same era.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela P. Schoellig, and Andreas Krause · 2017
Cited alongside, same era.
Conservative contextual linear bandits
Abbas Kazerouni, Mohammad Ghavamzadeh, Yasin Abbasi, and Benjamin Van Roy · 2017
Cited alongside, same era.
Near optimal exploration-exploitation in non-communicating markov decision processes
Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric
Cited in the paper.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Ronald Ortner
Cited in the paper.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Remi Tachet des Combes · 2019
Later among the works it cites.
Regret minimization in infinite-horizon finite markov decision processes
Alessandro Lazaric, Matteo Pirotta, and Ronan Fruit · 2019
Later among the works it cites.
Safe policy improvement with baseline bootstrapping in factored environments
Thiago D. Simão and Matthijs T. J. Spaan · 2019
Later among the works it cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Lin F. Yang and Mengdi Wang · 2019
Later among the works it cites.
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.