Fetching the paper…
Reading the bibliography…
Intrinsic rewards play a central role in handling the exploration-exploitation trade-off when designing sequential decision-making algorithms, in both foundational theory and state-of-the-art deep reinforcement learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Exponential polynomials
Eric Temple Bell · 1934
Earlier work this paper cites.
Associative reinforcement learning using linear probabilistic concepts
Naoki Abe and Philip M Long · 1999
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
The value of unlabeled data for classification problems
Tong Zhang and F Oles · 2000
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas P Hayes, and Sham M Kakade · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Linearly parameterized bandits
Paat Rusmevichientong and John N Tsitsiklis · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Efficient bayes-adaptive reinforcement learning using sample-based search
Arthur Guez, David Silver, and Peter Dayan · 2012
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Shipra Agrawal and Navin Goyal · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Batch-mode active learning via error bound minimization
Quanquan Gu, Tong Zhang, and Jiawei Han · 2014
Earlier work this paper cites.
Convergence rates of active learning for maximum likelihood estimation
Kamalika Chaudhuri, Sham Kakade, Praneeth Netrapalli, and Sujay Sanghavi · 2015
Earlier work this paper cites.
Fast gradient descent for drifting least squares regression, with application to bandits
Nathaniel Korda, LA Prashanth, and Rémi Munos · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2018
Later among the works it cites.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, and Sham M Kakade · 2019
Later among the works it cites.
Model-predictive policy learning with uncertainty regularization for driving in dense traffic
Mikael Henaff, Alfredo Canziani, and Yann LeCun · 2019
Later among the works it cites.
Perturbed-history exploration in stochastic linear bandits
Branislav Kveton, Csaba Szepesvari, Mohammad Ghavamzadeh, and Craig Boutilier · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
Deepak Pathak, Dhiraj Gandhi, and Abhinav Gupta · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc G Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Deep bayesian active learning with image data
Yarin Gal, Riashat Islam, and Zoubin Ghahramani · 2017
Cited alongside, same era.
Scalable generalized linear bandits: Online computation and hashing
Kwang-Sung Jun, Aniruddha Bhargava, Robert Nowak, and Rebecca Willett · 2017
Cited alongside, same era.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Cited alongside, same era.
Count-based exploration in feature space for reinforcement learning
Jarryd Martin, Suraj Narayanan Sasikumar, Tom Everitt, and Marcus Hutter · 2017
Cited alongside, same era.
Aleksandrs Slivkins · 2019
Later among the works it cites.
Deep neural linear bandits: Overcoming catastrophic forgetting through likelihood matching
Tom Zahavy and Shie Mannor · 2019
Later among the works it cites.
Deep batch active learning by diverse, uncertain gradient lower bounds
Jordan T Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal · 2020
Later among the works it cites.
The elliptical potential lemma revisited
Alexandra Carpentier, Claire Vernade, and Yasin Abbasi-Yadkori · 2020
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Weitong Zhang, Dongruo Zhou, Lihong Li, and Quanquan Gu · 2020
Later among the works it cites.
Neural contextual bandits with ucb-based exploration
Dongruo Zhou, Lihong Li, and Quanquan Gu · 2020
Later among the works it cites.
Gone fishing: Neural active learning with fisher embeddings
Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Sham Kakade · 2021
Closest in time.
Principled exploration via optimistic bootstrapping and backward induction
Chenjia Bai, Lingxiao Wang, Zhaoran Wang, Lei Han, Jianye Hao, Animesh Garg, and Peng Liu · 2021
Closest in time.
An efficient algorithm for generalized linear bandit: Online stochastic gradient descent and thompson sampling
Qin Ding, Cho-Jui Hsieh, and James Sharpnack · 2021
Closest in time.
Randomized exploration for reinforcement learning with general value function approximation
Haque Ishfaq, Qiwen Cui, Viet Nguyen, Alex Ayoub, Zhuoran Yang, Zhaoran Wang, Doina Precup, and Lin F Yang · 2021
Closest in time.
Online limited memory neural-linear bandits with likelihood matching
Ofir Nabati, Tom Zahavy, and Shie Mannor · 2021
Closest in time.