Fetching the paper…
Reading the bibliography…
This work tackles the problem of robust zero-shot planning in non-stationary stochastic environments.
State estimation for systems with sojourn-time-dependent Markov model switching
L. Campo, P. Mookerjee, and Y. Bar-Shalom · 1991
Earlier work this paper cites.
Game theory
D. Fudenberg and J. Tirole · 1991
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Hidden-mode Markov decision processes
S. P. Choi, D.-y. Yeung, and N. L. Zhang · 1999
Earlier work this paper cites.
Hidden-mode Markov decision processes for nonstationary sequential decision making
S. P. Choi, D.-Y. Yeung, and N. L. Zhang · 2000
Earlier work this paper cites.
Solving hidden-mode Markov decision problems
S. P. Choi, N. L. Zhang, and D.-Y. Yeung · 2001
Earlier work this paper cites.
Reinforcement learning in dynamic environments using instantiated information
M. A. Wiering · 2001
Earlier work this paper cites.
Multiple model-based reinforcement learning
K. Doya, K. Samejima, K.-i. Katagiri, and M. Kawato · 2002
Earlier work this paper cites.
ε \varepsilon -mdps: Learning in varying environments
I. Szita, B. Takács, and A. Lörincz · 2002
Earlier work this paper cites.
Robust dynamic programming
G. N. Iyengar · 2005
Earlier work this paper cites.
Learning in non-stationary partially observable Markov decision processes
R. Jaulmes, J. Pineau, and D. Precup · 2005
Earlier work this paper cites.
Dealing with non-stationary environments using context detection
B. C. Da Silva, E. W. Basso, A. L. Bazzan, and P. M. Engel · 2006
Earlier work this paper cites.
Bandit based Monte-Carlo planning
L. Kocsis and C. Szepesvári · 2006
Cited alongside, same era.
Value function based reinforcement learning in changing Markovian environments
B. C. Csáji and L. Monostori · 2008
Cited alongside, same era.
Multi-armed bandits in metric spaces
R. Kleinberg, A. Slivkins, and E. Upfal · 2008
Cited alongside, same era.
Optimal transport: old and new , volume 338
C. Villani · 2008
Cited alongside, same era.
Online Markov Decision Processes
E. Even-Dar, S. M. Kakade, and Y. Mansour · 2009
Cited alongside, same era.
Open loop optimistic planning
S. Bubeck and R. Munos · 2010
Cited alongside, same era.
On the locality of action domination in sequential decision making
Online learning in Markov decision processes with changing cost sequences
T. Dick, A. Gyorgy, and C. Szepesvari · 2014
Later among the works it cites.
Sequential decision-making under non-stationary environments via sequential change-point detection
E. Hadoux, A. Beynier, and P. Weng · 2014
Later among the works it cites.
From bandits to monte-carlo tree search: The optimistic principle applied to optimization and planning
R. Munos · 2014
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Later among the works it cites.
Learning to soar: Resource-constrained exploration in reinforcement learning
J. J. Chung, N. R. Lawrance, and S. Sukkarieh · 2015
Later among the works it cites.
Markovian sequential decision-making in non-stationary environments: application to argumentative debates
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Rachelson and M. G. Lagoudakis · 2010
Cited alongside, same era.
A survey of Monte Carlo tree search methods
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton · 2012
Cited alongside, same era.
Online learning in Markov decision processes with adversarially chosen transition probability distributions
Y. Abbasi, P. L. Bartlett, V. Kanade, Y. Seldin, and C. Szepesvári · 2013
Cited alongside, same era.
Trial-based heuristic tree search for finite horizon MDPs
T. Keller and M. Helmert · 2013
Cited alongside, same era.
Reinforcement learning in robust markov decision processes
S. H. Lim, H. Xu, and S. Mannor · 2013
Cited alongside, same era.
PAC Optimal Exploration in Continuous Space Markov Decision Processes
J. Pazis and R. Parr · 2013
Cited alongside, same era.
E. Hadoux · 2015
Later among the works it cites.
Policy gradient in lipschitz Markov Decision Processes
M. Pirotta, M. Restelli, and L. Bascetta · 2015
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Later among the works it cites.
Empirical evaluation of a Q-Learning Algorithm for Model-free Autonomous Soaring
E. Lecarpentier, S. Rapp, M. Melo, and E. Rachelson · 2017
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
K. Asadi, D. Misra, and M. L. Littman · 2018
Later among the works it cites.
Distributional reinforcement learning with quantile regression
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos · 2018
Later among the works it cites.
Open loop execution of tree-search algorithms
E. Lecarpentier, G. Infantes, C. Lesire, and E. Rachelson · 2018
Later among the works it cites.