Fetching the paper…
Reading the bibliography…
We initiate the study of multi-stage episodic reinforcement learning under adversarial corruptions in both the rewards and the transition probabilities of the underlying system extending recent results for the special case of stochastic bandits.
Finite-time regret bounds for the multi-armed bandit problems
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fisher · 2002
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
Eyal Even-Dar, Shie Mannor, and Yishay Mansour · 2006
Earlier work this paper cites.
Exploration–exploitation tradeoff using variance estimates in multi-armed bandits
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvári · 2009
Earlier work this paper cites.
Online markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Online markov decision processes under bandit feedback
Gergely Neu, Andras Antos, András György, and Csaba Szepesvári · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
The best of both worlds: stochastic and adversarial bandits
Sébastien Bubeck and Aleksandrs Slivkins · 2012
Earlier work this paper cites.
The adversarial stochastic shortest path problem with unknown transition probabilities
Gergely Neu, Andras Gyorgy, and Csaba Szepesvári · 2012
Earlier work this paper cites.
Online learning in markov decision processes with adversarially chosen transition probability distributions
Yasin Abbasi-Yadkori, Peter L Bartlett, Varun Kanade, Yevgeny Seldin, and Csaba Szepesvari · 2013
Earlier work this paper cites.
Online learning in episodic markovian decision processes by relative entropy policy search
Alexander Zimin and Gergely Neu · 2013
Earlier work this paper cites.
One practical algorithm for both stochastic and adversarial bandits
Yevgeny Seldin and Aleksandrs Slivkins · 2014
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Earlier work this paper cites.
An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits
Peter Auer and Chao-Kai Chiang · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Earlier work this paper cites.
Fairness in reinforcement learning
Shahin Jabbari, Matthew Joseph, Michael Kearns, Jamie Morgenstern, and Aaron Roth · 2017
Earlier work this paper cites.
An improved parametrization and analysis of the EXP3++ algorithm for stochastic and adversarial bandits
Yevgeny Seldin and Gábor Lugosi · 2017
Cited alongside, same era.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Stochastic bandits robust to adversarial corruptions
Thodoris Lykouris, Vahab Mirrokni, and Renato Paes Leme · 2018
Cited alongside, same era.
More adaptive algorithms for adversarial bandits
Chen-Yu Wei and Haipeng Luo · 2018
Cited alongside, same era.
Robust dynamic assortment optimization in the presence of outlier customers
Xi Chen, Akshay Krishnamurthy, and Yining Wang · 2019
Cited alongside, same era.
Provably efficient q-learning with function approximation via distribution shift error checking oracle
Simon S. Du, Yuping Luo, Ruosong Wang, and Hanrui Zhang · 2019
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Closest in time.
A unifying view of optimism in episodic reinforcement learning
Gergely Neu and Ciara Pike-Burke · 2020
Closest in time.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Closest in time.
Stochastic dueling bandits with adversarial corruption
Arpit Agarwal, Shivani Agarwal, and Prathamesh Patil · 2021
Closest in time.
Improved corruption robust algorithms for episodic reinforcement learning
Yifang Chen, Simon S. Du, and Kevin Jamieson · 2021
Closest in time.
A kernel-based approach to non-stationary reinforcement learning in metric spaces
Omar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann, and Michal Valko · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Better algorithms for stochastic bandits with adversarial corruptions
Anupam Gupta, Tomer Koren, and Kunal Talwar · 2019
Cited alongside, same era.
Stochastic linear optimization with adversarial corruption
Yingkai Li, Edmund Y Lou, and Liren Shan · 2019
Cited alongside, same era.
Online convex optimization in adversarial markov decision processes
Aviv Rosenberg and Yishay Mansour · 2019
Cited alongside, same era.
Worst-case regret bounds for exploration via randomized value functions
Daniel Russo · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin Jamieson · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Closest in time.
Logarithmic regret for reinforcement learning with linear function approximation
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2021
Closest in time.
The best of both worlds: Stochastic and adversarial episodic mdps with unknown transition
Tiancheng Jin, Longbo Huang, and Haipeng Luo · 2021
Closest in time.
Online learning in mdps with linear function approximation and bandit feedback
Gergely Neu and Julia Olkhovskaya · 2021
Closest in time.
Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach
Chen-Yu Wei and Haipeng Luo · 2021
Closest in time.
Optimism in reinforcement learning with generalized linear function approximation
Yining Wang, Ruosong Wang, Simon Shaolei Du, and Akshay Krishnamurthy · 2021
Closest in time.
Robust policy gradient against strong data corruption
Xuezhou Zhang, Yiding Chen, Xiaojin Zhu, and Wen Sun · 2021
Closest in time.
Tsallis-inf: An optimal algorithm for stochastic and adversarial bandits
Julian Zimmert and Yevgeny Seldin · 2021
Closest in time.
A model selection approach for corruption robust reinforcement learning
Chen-Yu Wei, Christoph Dann, and Julian Zimmert · 2022
Closest in time.
Nonstationary reinforcement learning: The blessing of (more) optimism
Wang Chi Cheung, David Simchi-Levi, and Ruihao Zhu · 2023
Closest in time.
Robust dynamic pricing with demand learning in the presence of outlier customers
Xi Chen and Yining Wang · 2023
Closest in time.
Learning product rankings robust to fake users
Negin Golrezaei, Vahideh H. Manshadi, Jon Schneider, and Shreyas Sekar · 2023
Closest in time.
Contextual search in the presence of adversarial corruptions
Akshay Krishnamurthy, Thodoris Lykouris, Chara Podimata, and Robert E. Schapire · 2023
Closest in time.