Fetching the paper…
Reading the bibliography…
We propose a black-box reduction that turns a certain reinforcement learning algorithm with optimal regret in a (near-)stationary environment into another algorithm with optimal dynamic regret in a non-stationary environment, importantly without any prior knowledge on the degree of non-stationarity.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Adaptive algorithms for online decision problems
Elad Hazan and Comandur Seshadhri · 2007
Earlier work this paper cites.
Online markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
Parametric bandits: the generalized linear case
Sarah Filippi, Olivier Cappé, Aurélien Garivier, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
The online loop-free stochastic shortest-path problem
Gergely Neu, András György, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Deterministic mdps with adversarial rewards and bandit feedback
Raman Arora, Ofer Dekel, and Ambuj Tewari · 2012
Earlier work this paper cites.
The adversarial stochastic shortest path problem with unknown transition probabilities
Gergely Neu, Andras Gyorgy, and Csaba Szepesvári · 2012
Earlier work this paper cites.
Better rates for any adversarial deterministic mdp
Ofer Dekel and Elad Hazan · 2013
Earlier work this paper cites.
Online markov decision processes under bandit feedback
Gergely Neu, András György, Csaba Szepesvári, and András Antos · 2013
Earlier work this paper cites.
Online learning in episodic markovian decision processes by relative entropy policy search
Alexander Zimin and Gergely Neu · 2013
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Earlier work this paper cites.
Online learning in markov decision processes with changing cost sequences
Travis Dick, Andras Gyorgy, and Csaba Szepesvari · 2014
Earlier work this paper cites.
Strongly adaptive online learning
Amit Daniely, Alon Gonen, and Shai Shalev-Shwartz · 2015
Earlier work this paper cites.
Achieving all with no parameters: Adanormalhedge
Haipeng Luo and Robert E Schapire · 2015
Earlier work this paper cites.
Tracking the best expert in non-stationary stochastic environments
Chen-Yu Wei, Yi-Te Hong, and Chi-Jen Lu · 2016
Earlier work this paper cites.
Improved strongly adaptive online learning using coin betting
Kwang-Sung Jun, Francesco Orabona, Stephen Wright, and Rebecca Willett · 2017
Earlier work this paper cites.
Hedging the drift: Learning to optimize under non-stationarity
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2018
Cited alongside, same era.
Pratik Gajane, Ronald Ortner, and Peter Auer · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Efficient contextual bandits in non-stationary worlds
Haipeng Luo, Chen-Yu Wei, Alekh Agarwal, and John Langford · 2018
Cited alongside, same era.
Adaptively tracking the best bandit arm with an unknown number of distribution changes
Peter Auer, Pratik Gajane, and Ronald Ortner · 2019
Cited alongside, same era.
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Learning adversarial markov decision processes with delayed feedback
Tal Lancewicki, Aviv Rosenberg, and Yishay Mansour · 2020
Later among the works it cites.
Bias no more: high-probability data-dependent regret bounds for adversarial bandits and mdps
Chung-Wei Lee, Haipeng Luo, Chen-Yu Wei, and Mengxiao Zhang · 2020
Later among the works it cites.
Variational regret bounds for reinforcement learning
Ronald Ortner, Pratik Gajane, and Peter Auer · 2020
Later among the works it cites.
Regret bound balancing and elimination for model selection in bandits and rl
Aldo Pacchiano, Christoph Dann, Claudio Gentile, and Peter Bartlett · 2020
Later among the works it cites.
Stochastic shortest path with adversarially changing costs
Aviv Rosenberg and Yishay Mansour · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yifang Chen, Chung-Wei Lee, Haipeng Luo, and Chen-Yu Wei · 2019
Cited alongside, same era.
Learning to optimize under non-stationarity
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Cited alongside, same era.
Online learning for markov decision processes in nonstationary environments: A dynamic regret analysis
Yingying Li and Na Li · 2019
Cited alongside, same era.
Online convex optimization in adversarial markov decision processes
Aviv Rosenberg and Yishay Mansour · 2019
Cited alongside, same era.
Weighted linear bandits for non-stationary environments
Yoan Russac, Claire Vernade, and Olivier Cappé · 2019
Cited alongside, same era.
Regret balancing for bandit and rl model selection
Yasin Abbasi-Yadkori, Aldo Pacchiano, and My Phan · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Later among the works it cites.
Algorithms for non-stationary generalized linear bandits
Yoan Russac, Olivier Cappé, and Aurélien Garivier · 2020
Later among the works it cites.
Optimistic policy optimization with bandit feedback
Lior Shani, Yonathan Efroni, Aviv Rosenberg, and Shie Mannor · 2020
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
David Simchi-Levi and Yunzong Xu · 2020
Later among the works it cites.
Efficient learning in non-stationary linear markov decision processes
Ahmed Touati and Pascal Vincent · 2020
Later among the works it cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Ruosong Wang, Simon S Du, Lin F Yang, and Sham M Kakade · 2020
Later among the works it cites.
A simple approach for non-stationary linear bandits
Peng Zhao, Lijun Zhang, Yuan Jiang, and Zhi-Hua Zhou · 2020
Later among the works it cites.
Nonstationary reinforcement learning with linear function approximation
Huozhi Zhou, Jinglin Chen, Lav R Varshney, and Ashish Jagmohan · 2020
Later among the works it cites.
Minimax regret for stochastic shortest path with adversarial costs and known transition
Liyu Chen, Haipeng Luo, and Chen-Yu Wei · 2021
Closest in time.
A kernel-based approach to non-stationary reinforcement learning in metric spaces
Omar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann, and Michal Valko · 2021
Closest in time.
Regret bounds for generalized linear bandits under parameter drift
Louis Faury, Yoan Russac, Marc Abeille, and Clément Calauzènes · 2021
Closest in time.
Corruption robust exploration in episodic reinforcement learning
Thodoris Lykouris, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2021
Closest in time.
Is model-free learning nearly optimal for non-stationary rl?
Weichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi, and Tamer Başar · 2021
Closest in time.
Non-stationary linear bandits revisited
Peng Zhao and Lijun Zhang · 2021
Closest in time.