Fetching the paper…
Reading the bibliography…
We consider reinforcement learning (RL) in episodic Markov decision processes (MDPs) with linear function approximation under drifting environment.
Estimation des densités: risque minimax
J Bretagnolle and C Huber · 1979
Earlier work this paper cites.
Gaussian feedback capacity
Thomas M Cover and Sandeep Pombra · 1989
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Complexity analysis of real-time reinforcement learning
Sven Koenig and Reid G Simmons · 1993
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Q \mathchar 29009 -learning with linear function approximation
Francisco S Melo and M Isabel Ribeiro · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Markov decision processes with arbitrary reward processes
Jia Yuan Yu, Shie Mannor, and Nahum Shimkin · 2009
Earlier work this paper cites.
Reinforcement learning design for cancer clinical trials
Yufan Zhao, Michael R Kosorok, and Donglin Zeng · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
Aurélien Garivier and Eric Moulines · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Earlier work this paper cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2016
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Earlier work this paper cites.
Corralling a band of bandit algorithms
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, and Robert E Schapire · 2017
Earlier work this paper cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Earlier work this paper cites.
Real-time bidding by reinforcement learning in display advertising
Han Cai, Kan Ren, Weinan Zhang, Kleanthis Malialis, Jun Wang, Yong Yu, and Defeng Guo · 2017
Earlier work this paper cites.
Discrepancy-based algorithms for non-stationary rested bandits
Corinna Cortes, Giulia DeSalvo, Vitaly Kuznetsov, Mehryar Mohri, and Scott Yang · 2017
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Learning in structured MDPs with convex cost functions: Improved regret bounds for inventory management
Shipra Agrawal and Randy Jia · 2019
Cited alongside, same era.
Solving rubik’s cube with a robot hand
Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang · 2019
Provably efficient reinforcement learning with general value function approximation
Ruosong Wang, Ruslan Salakhutdinov, and Lin F Yang · 2020
Closest in time.
A simple approach for non-stationary linear bandits
Peng Zhao, Lijun Zhang, Yuan Jiang, and Zhi-Hua Zhou · 2020
Closest in time.
A near-optimal change-detection based algorithm for piecewise-stationary combinatorial semi-bandits
Huozhi Zhou, Lingda Wang, Lav R Varshney, and Ee-Peng Lim · 2020
Closest in time.
Improved worst-case regret bounds for randomized least-squares value iteration
Priyank Agrawal, Jinglin Chen, and Nan Jiang · 2021
Closest in time.
A provably efficient model-free posterior sampling method for episodic reinforcement learning
Christoph Dann, Mehryar Mohri, Tong Zhang, and Julian Zimmert · 2021
Closest in time.
Bilinear classes: A structural framework for provable generalization in RL
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The generalized likelihood ratio test meets klucb: an improved algorithm for piece-wise non-stationary bandits
Lilian Besson and Emilie Kaufmann · 2019
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Learning to optimize under non-stationarity
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2019
Cited alongside, same era.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J Russo, and Zheng Wen · 2019
Cited alongside, same era.
Weighted linear bandits for non-stationary environments
Yoan Russac, Claire Vernade, and Olivier Cappé · 2019
Cited alongside, same era.
Worst-case regret bounds for exploration via randomized value functions
Daniel Russo · 2019
Cited alongside, same era.
Nearly optimal algorithms for piecewise-stationary cascading bandits
Lingda Wang, Huozhi Zhou, Bingcong Li, Lav R Varshney, and Zhizhen Zhao · 2019
Cited alongside, same era.
Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Closest in time.
A provably efficient algorithm for linear markov decision process with low switching cost
Minbo Gao, Tianle Xie, Simon S Du, and Lin F Yang · 2021
Closest in time.
Towards deployment-efficient reinforcement learning: Lower bound and optimality
Jiawei Huang, Jinglin Chen, Li Zhao, Tao Qin, Nan Jiang, and Tie-Yan Liu · 2021
Closest in time.
Randomized exploration in reinforcement learning with general value function approximation
Haque Ishfaq, Qiwen Cui, Viet Nguyen, Alex Ayoub, Zhuoran Yang, Zhaoran Wang, Doina Precup, and Lin Yang · 2021
Closest in time.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Closest in time.
Online model selection for reinforcement learning with function approximation
Jonathan Lee, Aldo Pacchiano, Vidya Muthukumar, Weihao Kong, and Emma Brunskill · 2021
Closest in time.
Corruption robust exploration in episodic reinforcement learning
Thodoris Lykouris, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2021
Closest in time.
Near-optimal regret bounds for model-free rl in non-stationary episodic mdps
Weichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi, and Tamer Başar · 2021
Closest in time.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Closest in time.
Online learning in MDPs with linear function approximation and bandit feedback
Gergely Neu and Julia Olkhovskaya · 2021
Closest in time.
Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach
Chen-Yu Wei and Haipeng Luo · 2021
Closest in time.
Learning infinite-horizon average-reward MDPs with linear function approximation
Chen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, and Rahul Jain · 2021
Closest in time.
Randomized exploration is near-optimal for tabular mdp
Zhihan Xiong, Ruoqi Shen, and Simon S Du · 2021
Closest in time.
Optimistic policy optimization is provably efficient in non-stationary mdps
Han Zhong, Zhuoran Yang, and Zhaoran Wang Csaba Szepesvári · 2021
Closest in time.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2021
Closest in time.
On the statistical efficiency of reward-free exploration in non-linear rl
Jinglin Chen, Aditya Modi, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2022
Closest in time.
Near-optimal goal-oriented reinforcement learning in non-stationary environments
Liyu Chen and Haipeng Luo · 2022
Closest in time.
Yuhao Ding and Javad Lavaei · 2022
Closest in time.
Asymptotic instance-optimal algorithms for interactive decision making
Kefan Dong and Tengyu Ma · 2022
Closest in time.
Nearly minimax optimal reinforcement learning with linear function approximation
Pihe Hu, Yu Chen, and Longbo Huang · 2022
Closest in time.
Feel-good thompson sampling for contextual bandits and reinforcement learning
Tong Zhang · 2022
Closest in time.