Fetching the paper…
Reading the bibliography…
Information-directed sampling (IDS) has revealed its potential as a data-efficient algorithm for reinforcement learning (RL).
Elements of Information Theory
T.M. Cover and J.A. Thomas · 1991
Earlier work this paper cites.
Asymptotically optimal information-directed sampling
Johannes Kirschner, Tor Lattimore, Claire Vernade, and Csaba Szepesvári · 2011
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
Learning to optimize via information-directed sampling
Daniel Russo and Benjamin Van Roy · 2014
Earlier work this paper cites.
An information-theoretic analysis of thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
An information-theoretic analysis for thompson sampling with many actions
Shi Dong and Benjamin Van Roy · 2018
Earlier work this paper cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Earlier work this paper cites.
Information directed sampling and bandits with heteroscedastic noise
Johannes Kirschner and Andreas Krause · 2018
Earlier work this paper cites.
Information directed sampling for stochastic bandits with graph feedback
Fang Liu, Swapna Buccapatnam, and Ness Shroff · 2018
Earlier work this paper cites.
Information-directed exploration for deep reinforcement learning
Nikolay Nikolov, Johannes Kirschner, Felix Berkenkamp, and Andreas Krause · 2018
Earlier work this paper cites.
Learning to optimize via information-directed sampling
Daniel Russo and Benjamin Van Roy · 2018
Earlier work this paper cites.
An information-theoretic approach to minimax regret in partial monitoring
Tor Lattimore and Csaba Szepesvári · 2019
Cited alongside, same era.
Information-theoretic confidence bounds for reinforcement learning
Xiuyuan Lu and Benjamin Van Roy · 2019
Cited alongside, same era.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J Russo, Zheng Wen, et al · 2019
Cited alongside, same era.
Worst-case regret bounds for exploration via randomized value functions
Daniel Russo · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Bilinear classes: A structural framework for provable generalization in rl
Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Information directed sampling for sparse linear bandits
Botao Hao, Tor Lattimore, and Wei Deng · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Later among the works it cites.
Asymptotically optimal information-directed sampling
Johannes Kirschner, Tor Lattimore, Claire Vernade, and Csaba Szepesvári · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
First-order bayesian regret analysis of thompson sampling
Sébastien Bubeck and Mark Sellke · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Information-Directed Sampling for Reinforcement Learning
Xiuyuan Lu · 2020
Cited alongside, same era.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Ruosong Wang, Russ R Salakhutdinov, and Lin Yang · 2020
Cited alongside, same era.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 2020
Cited alongside, same era.
Tor Lattimore and Andras Gyorgy · 2021
Later among the works it cites.
Reinforcement learning, bit by bit
Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, and Zheng Wen · 2021
Later among the works it cites.
Ucb momentum q-learning: Correcting the bias without forgetting
Pierre Ménard, Omar Darwiche Domingues, Xuedong Shang, and Michal Valko · 2021
Later among the works it cites.
Capacity of noisy permutation channels
Jennifer Tang and Yury Polyanskiy · 2021
Later among the works it cites.
Feel-good thompson sampling for contextual bandits and reinforcement learning
Tong Zhang · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2021
Later among the works it cites.
Contextual information-directed sampling
Botao Hao, Tor Lattimore, and Chao Qing · 2022
Closest in time.
Satisficing in time-sensitive bandit learning
Daniel Russo and Benjamin Van Roy · 2022
Closest in time.