Fetching the paper…
Reading the bibliography…
We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under temporal drifts, ie, both the reward and state transition distributions are allowed to evolve over time, as long as their respective total variations, quantified by suitable metrics, do not exceed certain variation budgets.
Hedging the drift: Learning to optimize under non-stationarity
Cheung, Wang Chi, David Simchi-Levi, Ruihao Zhu. 2019a · 1903
Earlier work this paper cites.
Model-free reinforcement learning in infinite-horizon average-reward markov decision processes
Wei, Chen-Yu, Mehdi Jafarnia-Jahromi, Haipeng Luo, Hiteshi Sharma, Rahul Jain. 2019 · 1910
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Hoeffding, Wassily. 1963 · 1963
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, Martin L. 1994 · 1994
Earlier work this paper cites.
Optimal adaptive policies for markov decision processes
Burnetas, Apostolos N., Michael N. Katehakis. 1997 · 1997
Earlier work this paper cites.
Zhou, Xiang, Ningyuan Chen, Xuefeng Gao, Yi Xiong. 2020 · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, P., N. Cesa-Bianchi, Y. Freund, R. Schapire. 2002a · 2002
Earlier work this paper cites.
Experts in a markov decision process
Even-Dar, Eyal, Sham M Kakade, , Yishay Mansour. 2005 · 2005
Earlier work this paper cites.
Optimistic linear programming gives logarithmic regret for irreducible mdps
Tewari, Ambuj, Peter L. Bartlett. 2008 · 2008
Earlier work this paper cites.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Bartlett, Peter L., Ambuj Tewari. 2009 · 2009
Earlier work this paper cites.
Introduction to algorithms
Cormen, Thomas H., Charles E. Leiserson, Ronald L. Rivest, Clifford Stein. 2009 · 2009
Earlier work this paper cites.
A nonparametric asymptotic analysis of inventory planning with censored demand
Huh, Woonghee Tim, Paat Rusmevichientong. 2009 · 2009
Earlier work this paper cites.
Online learning in markov decision processes with arbitrarily changing rewards and transitions
Yu, Jia Yuan, Shie Mannor. 2009 · 2009
Earlier work this paper cites.
Markov decision processes with arbitrary reward processes
Yu, Jia Yuan, Shie Mannor, Nahum Shimkin. 2009 · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, Thomas, Ronald Ortner, Peter Auer. 2010 · 2010
Earlier work this paper cites.
Online markov decision processes under bandit feedback
Neu, Gergely, Andras Antos, András György, Csaba Szepesvári. 2010 · 2010
Earlier work this paper cites.
Informing sequential clinical decision-making through reinforcement learning: an empirical study
Shortreed, Susan, Eric Laber, Daniel Lizotte, Scott Stroup, Joelle Pineau Susan Murphy. 2010 · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Yasin, David Pál, Csaba. Szepesvári. 2011 · 2011
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
Garivier, Aurélien, Eric Moulines. 2011 · 2011
Earlier work this paper cites.
The arrival of real-time bidding
Google. 2011 · 2011
Earlier work this paper cites.
Deterministic mdps with adversarial rewards and bandit feedback
Arora, Raman, Ofer Dekel, Ambuj Tewari. 2012 · 2012
Cited alongside, same era.
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems
Bubeck, S., N. Cesa-Bianchi. 2012 · 2012
Cited alongside, same era.
The adversarial stochastic shortest path problem with unknown transition probabilities
Neu, Gergely, Andras Gyorgy, Csaba Szepesvari. 2012 · 2012
Cited alongside, same era.
Online learning in markov decision processes with adversarially chosen transition probability distributions
Abbasi-Yadkori, Yasin, Peter L Bartlett, Varun Kanade, Yevgeny Seldin, Csaba Szepesvári. 2013 · 2013
Cited alongside, same era.
Stochastic multi-armed bandit with non-stationary rewards
Besbes, Omar, Yonatan Gur, Assaf Zeevi. 2014 · 2014
Cited alongside, same era.
Online learning in markov decision processes with changing cost sequences
Bandit Algorithms
Lattimore, T., C. Szepesvári. 2018 · 2018
Later among the works it cites.
Efficient contextual bandits in non-stationary worlds
Luo, H., C. Wei, A. Agarwal, J. Langford. 2018 · 2018
Later among the works it cites.
Improvements and generalizations of stochastic knapsack and markovian bandits approximation algorithms
Ma, Will. 2018 · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, Richard S., Andrew G. Barto. 2018 · 2018
Later among the works it cites.
Spectral state compression of markov processes
Zhang, Anru, Mengdi Wang. 2018 · 2018
Later among the works it cites.
Closing the gap: A learning algorithm for the lost-sales inventory system with lead times
Zhang, Huanan, Xiuli Chao, Cong Shi. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dick, Travis, András György, Csaba Szepesvári. 2014 · 2014
Cited alongside, same era.
Tight regret bounds for stochastic combinatorial semi-bandits
Kveton, Branislav, Zheng Wen, Azin Ashkan, Csaba Szepesvári. 2015 · 2015
Cited alongside, same era.
Wireless communications games in fixed and random environments
Zhou, Zhengyuan, Nicholas Bambos. 2015 · 2015
Cited alongside, same era.
Threshold bandits, with and without censored feedback
Abernethy, Jacob, Kareem Amin, Ruihao Zhu. 2016 · 2016
Cited alongside, same era.
Multi-armed bandits: Competing with optimal sequences
Karnin, Z., O. Anava. 2016 · 2016
Cited alongside, same era.
Chasing demand: Learning and earning in a changing environments
Keskin, N., A. Zeevi. 2016 · 2016
Cited alongside, same era.
Robust multiarmed bandit problems
Kim, Michael Jong, Andrew E.B. Lim. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Learning in structured mdps with convex cost functions: Improved regret bounds for inventory management
Agrawal, Shipra, Randy Jia. 2019 · 2019
Closest in time.
Learning in repeated auctions with budgets: Regret minimization and equilibrium
Balseiro, Santiago, Yonatan Gur. 2019 · 2019
Closest in time.
Large scale markov decision processes with changing rewards
Cardoso, Adrian Rivera, He Wang, Huan Xu. 2019 · 2019
Closest in time.
Improved analysis of ucrl2b
Fruit, Ronan, Matteo Pirotta, Alessandro Lazaric. 2019 · 2019
Closest in time.
Learning mean-field games
Guo, Xin, Anran Hu, Renyuan Xu, Junzi Zhang. 2019 · 2019
Closest in time.
Learning adversarial markov decision processes with bandit feedback and unknown transition
Jin, Chi, Tiancheng Jin, Haipeng Luo, Suvrit Sra, Tiancheng Yu. 2019 · 2019
Closest in time.
Online markov decision processes with time-varying transition probabilities and rewards
Li, Yingying, Aoxiao Zhong, Guannan Qu, Na Li. 2019 · 2019
Closest in time.
Variational regret bounds for reinforcement learning
Ortner, Ronald, Pratik Gajane, Peter Auer. 2019 · 2019
Closest in time.
Deep reinforcement learning with applications in transportation
Qin, Zhiwei (Tony), Jian Tang, Jieping Ye. 2019 · 2019
Closest in time.
Online convex optimization in adversarial Markov decision processes
Rosenberg, Aviv, Yishay Mansour. 2019 · 2019
Closest in time.
Randomized linear programming solves the markov decision problem in nearly-linear (sometimes sublinear) running time
Wang, Mengdi. 2019 · 2019
Closest in time.
Marrying stochastic gradient descent with bandits: Learning algorithms for inventory systems with fixed costs
Yuan, Hao, Qi Luo, Cong Shi. 2019 · 2019
Closest in time.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zhang, Zihan, Xiangyang Ji. 2019 · 2019
Closest in time.
Reinforcement with fading memories
Xu, Kuang, Se-Young Yun. 2020 · 2020
Closest in time.