Fetching the paper…
Reading the bibliography…
We make significant progress toward the stochastic shortest path problem with adversarial costs and unknown transition.
An analysis of stochastic shortest path problems
Dimitri P Bertsekas and John N Tsitsiklis · 1991
Earlier work this paper cites.
A provably efficient sample collection strategy for reinforcement learning
Jean Tarbouriech, Matteo Pirotta, Michal Valko, and Alessandro Lazaric · 2007
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Alina Beygelzimer, John Langford, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
The adversarial stochastic shortest path problem with unknown transition probabilities
Gergely Neu, Andras Gyorgy, and Csaba Szepesvári · 2012
Earlier work this paper cites.
Stochastic shortest path problems under weak conditions
Dimitri P Bertsekas and Huizhen Yu · 2013
Earlier work this paper cites.
Online learning in episodic markovian decision processes by relative entropy policy search
Alexander Zimin and Gergely Neu · 2013
Earlier work this paper cites.
Corralling a band of bandit algorithms
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, and Robert E Schapire · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sébastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Online convex optimization in adversarial Markov decision processes
Aviv Rosenberg and Yishay Mansour · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Minimax regret for stochastic shortest path with adversarial costs and known transition
Liyu Chen, Haipeng Luo, and Chen-Yu Wei · 2020
Later among the works it cites.
Near-optimal regret bounds for stochastic shortest path
Alon Cohen, Haim Kaplan, Yishay Mansour, and Aviv Rosenberg · 2020
Later among the works it cites.
Learning adversarial Markov decision processes with bandit feedback and unknown transition
Chi Jin, Tiancheng Jin, Haipeng Luo, Suvrit Sra, and Tiancheng Yu · 2020
Later among the works it cites.
Bias no more: high-probability data-dependent regret bounds for adversarial bandits and mdps
Chung-Wei Lee, Haipeng Luo, Chen-Yu Wei, and Mengxiao Zhang · 2020
Later among the works it cites.
Stochastic shortest path with adversarially changing costs
Aviv Rosenberg and Yishay Mansour · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimistic policy optimization with bandit feedback
Lior Shani, Yonathan Efroni, Aviv Rosenberg, and Shie Mannor
Cited in the paper.
Optimistic policy optimization with bandit feedback
Lior Shani, Yonathan Efroni, Aviv Rosenberg, and Shie Mannor
Cited in the paper.
No-regret exploration in goal-oriented reinforcement learning
Jean Tarbouriech, Evrard Garcelon, Michal Valko, Matteo Pirotta, and Alessandro Lazaric
Cited in the paper.
Impossible tuning made possible: A new expert algorithm and its applications
Liyu Chen, Haipeng Luo, and Chen-Yu Wei · 2021
Closest in time.