Fetching the paper…
Reading the bibliography…
We study learning contextual MDPs using a function approximation for both the rewards and the dynamics.
Neural network learning: Theoretical foundations , volume 9
Martin Anthony, Peter L Bartlett, Peter L Bartlett, et al · 1999
Earlier work this paper cites.
Contextual bandit learning with predictable rewards
Alekh Agarwal, Miroslav Dudík, Satyen Kale, John Langford, and Robert Schapire · 2012
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Practical contextual bandits with regression oracles
Dylan Foster, Alekh Agarwal, Miroslav Dudík, Haipeng Luo, and Robert Schapire · 2018
Earlier work this paper cites.
Markov decision processes with continuous side information
Aditya Modi, Nan Jiang, Satinder Singh, and Ambuj Tewari · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 2019
Cited alongside, same era.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Dylan Foster and Alexander Rakhlin · 2020
Cited alongside, same era.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Cited alongside, same era.
Bandit Algorithms
T. Lattimore and C. Szepesvári · 2020
Cited alongside, same era.
Upper counterfactual confidence bounds: a new optimism principle for contextual bandits
Yunbei Xu and Assaf Zeevi · 2020
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
On reward-free RL with kernel and neural function approximations: Single-agent MDP and markov game
Shuang Qiu, Jieping Ye, Zhaoran Wang, and Zhuoran Yang · 2021
Later among the works it cites.
Bypassing the monster: A faster and simpler optimal algorithm for contextual bandits under realizability
David Simchi-Levi and Yunzong Xu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Modi and Ambuj Tewari · 2020
Cited alongside, same era.
Later among the works it cites.
Near optimal reward-free reinforcement learning
Zihan Zhang, Simon Du, and Xiangyang Ji · 2021
Later among the works it cites.
On the statistical efficiency of reward-free exploration in non-linear rl
Jinglin Chen, Aditya Modi, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2022
Closest in time.