Fetching the paper…
Reading the bibliography…
We consider model-free reinforcement learning (RL) in non-stationary Markov decision processes.
Hedging the drift: Learning to optimize under non-stationarity
Cheung, Wang Chi, David Simchi-Levi, Ruihao Zhu. 2019a · 1903
Earlier work this paper cites.
Corruption robust exploration in episodic reinforcement learning
Lykouris, Thodoris, Max Simchowitz, Aleksandrs Slivkins, Wen Sun. 2019 · 1911
Earlier work this paper cites.
Learning adversarial MDPs with bandit feedback and unknown transition
Jin, Chi, Tiancheng Jin, Haipeng Luo, Suvrit Sra, Tiancheng Yu. 2019 · 1912
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, David A. 1975 · 1975
Earlier work this paper cites.
Closing the gap: A learning algorithm for the lost-sales inventory system with lead times
Zhang, Huanan, Xiuli Chao, Cong Shi. 2019 · 1980
Earlier work this paper cites.
Learning from delayed rewards
Watkins, Christopher John Cornish Hellaby. 1989 · 1989
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, Michael L. 1994 · 1994
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, Peter, Nicolo Cesa-Bianchi, Yoav Freund, Robert E Schapire. 2002 · 2002
Earlier work this paper cites.
Temple: Learning template of transitions for sample efficient multi-task RL
Sun, Yanchao, Xiangyu Yin, Furong Huang. 2020 · 2002
Earlier work this paper cites.
Learning and optimization with seasonal patterns
Chen, Ningyuan, Chun Wang, Longlin Wang. 2020b · 2005
Earlier work this paper cites.
A survey of reinforcement learning algorithms for dynamically varying environments
Padakandla, Sindhu. 2020 · 2005
Earlier work this paper cites.
Reinforcement learning for non-stationary Markov decision processes: The blessing of (more) optimism
Cheung, Wang Chi, David Simchi-Levi, Ruihao Zhu. 2020b · 2006
Earlier work this paper cites.
Linear last-iterate convergence for matrix games and stochastic games
Lee, Chung-Wei, Haipeng Luo, Chen-Yu Wei, Mengxiao Zhang. 2020 · 2006
Earlier work this paper cites.
PC-PG: Policy cover directed exploration for provable policy gradient learning
Agarwal, Alekh, Mikael Henaff, Sham Kakade, Wen Sun. 2020 · 2007
Earlier work this paper cites.
Sequential transfer in reinforcement learning with a generative model
Tirinzoni, Andrea, Riccardo Poiani, Marcello Restelli. 2020 · 2007
Earlier work this paper cites.
A nonparametric asymptotic analysis of inventory planning with censored demand
Huh, Woonghee Tim, Paat Rusmevichientong. 2009 · 2009
Earlier work this paper cites.
Intrinsic robustness of the price of anarchy
Roughgarden, Tim. 2009 · 2009
Earlier work this paper cites.
Online learning in Markov decision processes with arbitrarily changing rewards and transitions
Yu, Jia Yuan, Shie Mannor. 2009 · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, Thomas, Ronald Ortner, Peter Auer. 2010 · 2010
Earlier work this paper cites.
Online Markov decision processes under bandit feedback
Neu, Gergely, Andras Antos, András György, Csaba Szepesvári. 2010 · 2010
Earlier work this paper cites.
Online algorithms for the multi-armed bandit problem with Markovian rewards
Tekin, Cem, Mingyan Liu. 2010 · 2010
Earlier work this paper cites.
Efficient learning in non-stationary linear Markov decision processes
Touati, Ahmed, Pascal Vincent. 2020 · 2010
Earlier work this paper cites.
Nonstationary reinforcement learning with linear function approximation
Zhou, Huozhi, Jinglin Chen, Lav R Varshney, Ashish Jagmohan. 2020a · 2010
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
Garivier, Aurélien, Eric Moulines. 2011 · 2011
Earlier work this paper cites.
Informing sequential clinical decision-making through reinforcement learning: An empirical study
Shortreed, Susan M, Eric Laber, Daniel J Lizotte, T Scott Stroup, Joelle Pineau, Susan A Murphy. 2011 · 2011
Cited alongside, same era.
Deterministic MDPs with adversarial rewards and bandit feedback
Arora, Raman, Ofer Dekel, Ambuj Tewari. 2012 · 2012
Cited alongside, same era.
Sample complexity of multi-task reinforcement learning
Brunskill, Emma, Lihong Li. 2013 · 2013
Cited alongside, same era.
Online learning in Markov decision processes with adversarially chosen transition probability distributions
Yadkori, Yasin Abbasi, Peter L Bartlett, Varun Kanade, Yevgeny Seldin, Csaba Szepesvári. 2013 · 2013
Cited alongside, same era.
Stochastic multi-armed-bandit problem with non-stationary rewards
Besbes, Omar, Yonatan Gur, Assaf Zeevi. 2014 · 2014
Cited alongside, same era.
Adaptively tracking the best bandit arm with an unknown number of distribution changes
Auer, Peter, Pratik Gajane, Ronald Ortner. 2019 · 2019
Later among the works it cites.
Provably efficient Q-learning with low switching cost
Bai, Yu, Tengyang Xie, Nan Jiang, Yu-Xiang Wang. 2019 · 2019
Later among the works it cites.
Learning in repeated auctions with budgets: Regret minimization and equilibrium
Balseiro, Santiago R., Yonatan Gur. 2019 · 2019
Later among the works it cites.
Optimal exploration–exploitation in a multi-armed bandit problem with non-stationary rewards
Besbes, Omar, Yonatan Gur, Assaf Zeevi. 2019 · 2019
Later among the works it cites.
A new algorithm for non-stationary contextual bandits: Efficient, optimal and parameter-free
Chen, Yifang, Chung-Wei Lee, Haipeng Luo, Chen-Yu Wei. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dick, Travis, Andras Gyorgy, Csaba Szepesvari. 2014 · 2014
Cited alongside, same era.
Fast convergence of regularized learning in games
Syrgkanis, Vasilis, Alekh Agarwal, Haipeng Luo, Robert E Schapire. 2015 · 2015
Cited alongside, same era.
Dynamic capacity management with general upgrading
Yu, Yueshan, Xin Chen, Fuqiang Zhang. 2015 · 2015
Cited alongside, same era.
Decentralized Q-learning for stochastic teams and games
Arslan, Gürdal, Serdar Yüksel. 2016 · 2016
Cited alongside, same era.
Simple pricing schemes for consumers with evolving values
Chawla, Shuchi, Nikhil R Devanur, Anna R Karlin, Balasubramanian Sivan. 2016 · 2016
Cited alongside, same era.
Multi-armed bandits: Competing with optimal sequences
Karnin, Zohar S, Oren Anava. 2016 · 2016
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Osband, Ian, Benjamin Van Roy. 2016 · 2016
Cited alongside, same era.
Lu, Junwei, Chaoqi Yang, Xiaofeng Gao, Liubin Wang, Changcheng Li, Guihai Chen. 2019 · 2019
Later among the works it cites.
Variational regret bounds for reinforcement learning
Ortner, Ronald, Pratik Gajane, Peter Auer. 2019 · 2019
Later among the works it cites.
Learning to collaborate in Markov decision processes
Radanovic, Goran, Rati Devidze, David Parkes, Adish Singla. 2019 · 2019
Later among the works it cites.
Independent policy gradient methods for competitive reinforcement learning
Daskalakis, Constantinos, Dylan J Foster, Noah Golowich. 2020 · 2020
Closest in time.
Dynamic regret of policy optimization in non-stationary environments
Fei, Yingjie, Zhuoran Yang, Zhaoran Wang, Qiaomin Xie. 2020 · 2020
Closest in time.
POLY-HOOT: Monte-Carlo planning in continuous space MDPs with non-asymptotic analysis
Mao, Weichao, Kaiqing Zhang, Qiaomin Xie, Tamer Başar. 2020 · 2020
Closest in time.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Misra, Dipendra, Mikael Henaff, Akshay Krishnamurthy, John Langford. 2020 · 2020
Closest in time.
Reinforcement learning with perturbed rewards
Wang, Jingkang, Yang Liu, Bo Li. 2020 · 2020
Closest in time.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Zihan, Yuan Zhou, Xiangyang Ji. 2020 · 2020
Closest in time.
A simple approach for non-stationary linear bandits
Zhao, Peng, Lijun Zhang, Yuan Jiang, Zhi-Hua Zhou. 2020 · 2020
Closest in time.
Near-optimal model-free reinforcement learning in non-stationary episodic MDPs
Anonymous. 2021 · 2021
Closest in time.
Meta dynamic pricing: Transfer learning across experiments
Bastani, Hamsa, David Simchi-Levi, Ruihao Zhu. 2021 · 2021
Closest in time.
To interfere or not to interfere: Information revelation and price-setting incentives in a multiagent learning environment
Birge, John R., Hongfan Chen, N. Bora Keskin, Amy Ward. 2021 · 2021
Closest in time.
A kernel-based approach to non-stationary reinforcement learning in metric spaces
Domingues, Omar Darwiche, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann, Michal Valko. 2021 · 2021
Closest in time.
A provably efficient algorithm for linear Markov decision process with low switching cost
Gao, Minbo, Tianle Xie, Simon S Du, Lin F Yang. 2021 · 2021
Closest in time.
Online learning in unknown Markov games
Tian, Yi, Yuanhao Wang, Tiancheng Yu, Suvrit Sra. 2021 · 2021
Closest in time.
Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach
Wei, Chen-Yu, Haipeng Luo. 2021 · 2021
Closest in time.
Marrying stochastic gradient descent with bandits: Learning algorithms for inventory systems with fixed costs
Yuan, Hao, Qi Luo, Cong Shi. 2021 · 2021
Closest in time.
Zhou, Xiang, Yi Xiong, Ningyuan Chen, Xuefeng Gao. 2020b · 2021
Closest in time.