Fetching the paper…
Reading the bibliography…
Stochastic shortest path (SSP) is a well-known problem in planning and control, in which an agent has to reach a goal state in minimum total expected cost.
An analysis of stochastic shortest path problems
Dimitri P Bertsekas and John N Tsitsiklis · 1991
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L Bartlett and Ambuj Tewari · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
The online loop-free stochastic shortest-path problem
Gergely Neu, András György, and Csaba Szepesvári · 2010
Cited alongside, same era.
The adversarial stochastic shortest path problem with unknown transition probabilities
Gergely Neu, Andras Gyorgy, and Csaba Szepesvári · 2012
Cited alongside, same era.
Stochastic shortest path problems under weak conditions
Dimitri P Bertsekas and Huizhen Yu · 2013
Cited alongside, same era.
Online learning in episodic markovian decision processes by relative entropy policy search
Alexander Zimin and Gergely Neu · 2013
Cited alongside, same era.
Online stochastic shortest path with bandit feedback and unknown transition function
Aviv Rosenberg and Yishay Mansour
Cited in the paper.
Online convex optimization in adversarial markov decision processes
Aviv Rosenberg and Yishay Mansour
Cited in the paper.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Later among the works it cites.
No-regret exploration in goal-oriented reinforcement learning, 2019
Jean Tarbouriech, Evrard Garcelon, Michal Valko, Matteo Pirotta, and Alessandro Lazaric · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…