Fetching the paper…
Reading the bibliography…
We present the E-UC$^3$RL algorithm for regret minimization in Stochastic Contextual Markov Decision Processes (CMDPs).
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, J. and Zhang, T · 2007
Earlier work this paper cites.
Contextual bandit learning with predictable rewards
Agarwal, A., Dudík, M., Kale, S., Langford, J., and Schapire, R · 2012
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Russo, D. and Van Roy, B · 2013
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Agarwal, A., Hsu, D., Kale, S., Langford, J., Li, L., and Schapire, R · 2014
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Osband, I. and Van Roy, B · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Online learning in episodic markovian decision processes by relative entropy policy search
Zimin, A. and Neu, G · 2014
Earlier work this paper cites.
Contextual markov decision processes
Hallak, A., Di Castro, D., and Mannor, S · 2015
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Earlier work this paper cites.
Efficient reinforcement learning in deterministic systems with value function generalization
Wen, Z. and Van Roy, B · 2017
Earlier work this paper cites.
Practical contextual bandits with regression oracles
Foster, D., Agarwal, A., Dudik, M., Luo, H., and Schapire, R · 2018
Earlier work this paper cites.
Markov decision processes with continuous side information
Modi, A., Jiang, N., Singh, S., and Tewari, A · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Introduction to multi-armed bandits
Slivkins, A · 2019
Cited alongside, same era.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Sun, W., Jiang, N., Krishnamurthy, A., Agarwal, A., and Langford, J · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A., Jia, Z., Szepesvari, C., Wang, M., and Yang, L · 2020
Cited alongside, same era.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Foster, D. and Rakhlin, A · 2020
Cited alongside, same era.
Bandit Algorithms
Lattimore, T. and Szepesvári, C · 2020
Cited alongside, same era.
The statistical complexity of interactive decision making
Foster, D. J., Kakade, S. M., Qian, J., and Rakhlin, A · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Jin, C., Liu, Q., and Miryoosefi, S · 2021
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
Simchi-Levi, D. and Xu, Y · 2021
Later among the works it cites.
A general framework for sample-efficient function approximation in reinforcement learning
Chen, Z., Li, C. J., Yuan, A., Gu, Q., and Jordan, M. I · 2022
Closest in time.
Learning efficiently function approximation for contextual MDP
Levy, O. and Mansour, Y · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
No-regret exploration in contextual reinforcement learning
Modi, A. and Tewari, A · 2020
Cited alongside, same era.
Near-optimal regret bounds for stochastic shortest path
Rosenberg, A., Cohen, A., Mansour, Y., and Kaplan, H · 2020
Cited alongside, same era.
Optimistic policy optimization with bandit feedback
Shani, L., Efroni, Y., Rosenberg, A., and Mannor, S · 2020
Cited alongside, same era.
Upper counterfactual confidence bounds: a new optimism principle for contextual bandits
Xu, Y. and Zeevi, A · 2020
Cited alongside, same era.
A provably efficient model-free posterior sampling method for episodic reinforcement learning
Dann, C., Mohri, M., Zhang, T., and Zimmert, J · 2021
Cited alongside, same era.
Efficient first-order contextual bandits: Prediction, allocation, and triangular discrimination
Foster, D. J. and Krishnamurthy, A · 2021
Cited alongside, same era.
Closest in time.
When is partially observable reinforcement learning not scary?
Liu, Q., Chung, A., Szepesvári, C., and Jin, C · 2022
Closest in time.
The role of coverage in online reinforcement learning
Xie, T., Foster, D. J., Bai, Y., Jiang, N., and Kakade, S. M · 2022
Closest in time.
Feel-good thompson sampling for contextual bandits and reinforcement learning
Zhang, T · 2022
Closest in time.
Optimism in face of a context: Regret guarantees for stochastic contextual mdp
Levy, O. and Mansour, Y · 2023
Closest in time.
Efficient rate optimal regret for adversarial contextual mdps using online function approximation
Levy, O., Cohen, A., Cassel, A. B., and Mansour, Y · 2023
Closest in time.
Reinforcement Learning: Foundations
Mannor, S., Mansour, Y., and Tamar, A · 2023
Closest in time.
Uniform-PAC guarantees for model-based RL with bounded eluder dimension
Wu, Y., He, J., and Gu, Q · 2023
Closest in time.