Fetching the paper…
Reading the bibliography…
Eluder dimension and information gain are two widely used methods of complexity measures in bandit and reinforcement learning.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: no regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 2010
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Dan Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the Eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
An information-theoretic analysis of thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Earlier work this paper cites.
On kernelized multi-armed bandits
Sayak Ray Chowdhury and Aditya Gopalan · 2017
Cited alongside, same era.
Agnostic Q-learning with function approximation in deterministic systems: Tight bounds on approximation error and sample complexity
Simon S Du, Jason D Lee, Gaurav Mahajan, and Ruosong Wang · 2020
Cited alongside, same era.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Dylan Foster and Alexander Rakhlin · 2020
Cited alongside, same era.
Dylan J Foster, Alexander Rakhlin, David Simchi-Levi, and Yunzong Xu · 2020
Cited alongside, same era.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Ruosong Wang, Russ R Salakhutdinov, and Lin Yang · 2020
Cited alongside, same era.
On function approximation in reinforcement learning: Optimism in the face of large state spaces
Zhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang, and Michael I Jordan · 2020
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in RL
Simon S Du, Sham M Kakade, Jason D Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Closest in time.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms, 2021
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Closest in time.
Eluder dimension and generalized rank
Gene Li, Pritish Kamath, Dylan J Foster, and Nathan Srebro · 2021
Closest in time.
On information gain and regret bounds in gaussian process bandits
Sattar Vakili, Kia Khezeli, and Victor Picheny · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Closest in time.