Fetching the paper…
Reading the bibliography…
We initiate the study of dynamic regret minimization for goal-oriented reinforcement learning modeled by a non-stationary stochastic shortest path problem with changing cost and transition functions.
Refined lower bounds for adversarial bandits
Sébastien Gerchinovitz and Tor Lattimore · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Pratik Gajane, Ronald Ortner, and Peter Auer · 2018
Earlier work this paper cites.
Adaptively tracking the best bandit arm with an unknown number of distribution changes
Peter Auer, Pratik Gajane, and Ronald Ortner · 2019
Earlier work this paper cites.
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
Earlier work this paper cites.
Reinforcement learning for non-stationary Markov decision processes: The blessing of (more) optimism
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2020
Earlier work this paper cites.
Near-optimal regret bounds for stochastic shortest path
Alon Cohen, Haim Kaplan, Yishay Mansour, and Aviv Rosenberg · 2020
Earlier work this paper cites.
Dynamic regret of policy optimization in non-stationary environments
Yingjie Fei, Zhuoran Yang, Zhaoran Wang, and Qiaomin Xie · 2020
Earlier work this paper cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Earlier work this paper cites.
Variational regret bounds for reinforcement learning
Ronald Ortner, Pratik Gajane, and Peter Auer · 2020
Earlier work this paper cites.
Algorithms for non-stationary generalized linear bandits
Yoan Russac, Olivier Cappé, and Aurélien Garivier · 2020
Earlier work this paper cites.
Optimistic policy optimization with bandit feedback
Lior Shani, Yonathan Efroni, Aviv Rosenberg, and Shie Mannor · 2020
Cited alongside, same era.
No-regret exploration in goal-oriented reinforcement learning
Jean Tarbouriech, Evrard Garcelon, Michal Valko, Matteo Pirotta, and Alessandro Lazaric · 2020
Cited alongside, same era.
Efficient learning in non-stationary linear Markov decision processes
Ahmed Touati and Pascal Vincent · 2020
Cited alongside, same era.
Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon
Zihan Zhang, Xiangyang Ji, and Simon S Du · 2020
Cited alongside, same era.
Nonstationary reinforcement learning with linear function approximation
Huozhi Zhou, Jinglin Chen, Lav R Varshney, and Ashish Jagmohan · 2020
Cited alongside, same era.
Near-optimal model-free reinforcement learning in non-stationary episodic mdps
Weichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi, and Tamer Basar · 2021
Later among the works it cites.
Learning stochastic shortest path with linear function approximation
Yifei Min, Jiafan He, Tianhao Wang, and Quanquan Gu · 2021
Later among the works it cites.
Stochastic shortest path with adversarially changing costs
Aviv Rosenberg and Yishay Mansour · 2021
Later among the works it cites.
Tracking most severe arm changes in bandits
Joe Suk and Samory Kpotufe · 2021
Later among the works it cites.
Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret
Jean Tarbouriech, Runlong Zhou, Simon S Du, Matteo Pirotta, Michal Valko, and Alessandro Lazaric · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Finding the stochastic shortest path with low regret: The adversarial cost and unknown transition case
Liyu Chen and Haipeng Luo · 2021
Cited alongside, same era.
Minimax regret for stochastic shortest path
Alon Cohen, Yonathan Efroni, Yishay Mansour, and Aviv Rosenberg · 2021
Cited alongside, same era.
A kernel-based approach to non-stationary reinforcement learning in metric spaces
Omar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann, and Michal Valko · 2021
Cited alongside, same era.
Regret bounds for generalized linear bandits under parameter drift
Louis Faury, Yoan Russac, Marc Abeille, and Clément Calauzènes · 2021
Cited alongside, same era.
Online learning for stochastic shortest path model via posterior sampling
Mehdi Jafarnia-Jahromi, Liyu Chen, Rahul Jain, and Haipeng Luo · 2021
Cited alongside, same era.
Corruption-robust exploration in episodic reinforcement learning
Thodoris Lykouris, Max Simchowitz, Alex Slivkins, and Wen Sun · 2021
Cited alongside, same era.
Implicit finite-horizon approximation and efficient optimal algorithms for stochastic shortest path
Liyu Chen, Mehdi Jafarnia-Jahromi, Rahul Jain, and Haipeng Luo
Cited in the paper.
Regret bounds for stochastic shortest path problems with linear function approximation
Daniel Vial, Advait Parulekar, Sanjay Shakkottai, and R Srikant · 2021
Later among the works it cites.
Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach
Chen-Yu Wei and Haipeng Luo · 2021
Later among the works it cites.
A new look at dynamic regret for non-stationary stochastic bandits
Yasin Abbasi-Yadkori, Andras Gyorgy, and Nevena Lazic · 2022
Closest in time.
Yuhao Ding and Javad Lavaei · 2022
Closest in time.
A model selection approach for corruption robust reinforcement learning
Chen-Yu Wei, Christoph Dann, and Julian Zimmert · 2022
Closest in time.