Fetching the paper…
Reading the bibliography…
Bayesian methods for machine learning have been widely investigated, yielding principled methods for incorporating prior information into inference algorithms.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. Thompson · 1933
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 1935
Earlier work this paper cites.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Dual control theory, parts i and ii
A. Feldbaum · 1961
Earlier work this paper cites.
Optimal control of Markov decision processes with incomplete state estimation
K. Astrom · 1965
Earlier work this paper cites.
The optimal control of partially observable Markov processes over a finite horizon
R. Smallwood and E. Sondik · 1973
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
J. Gittins · 1979
Earlier work this paper cites.
Neuron-like elements that can solve difficult learning control problems
A. Barto, R. Sutton, and C. Anderson · 1983
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
R. Sutton · 1984
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. Lai and H. Robbins · 1985
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. Sutton · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
C. Watkins · 1989
Earlier work this paper cites.
Likelihood ratio gradient estimation for stochastic systems
P. Glynn · 1990
Earlier work this paper cites.
Bayes-Hermite quadrature
A. O’Hagan · 1991
Earlier work this paper cites.
Statistical Signal Processing
L. Scharf · 1991
Earlier work this paper cites.
DYNA, an integrated architecture for learning, planning, and reacting
R. Sutton · 1991
Earlier work this paper cites.
Bayesian adaptive control of time varying systems
R. Ravikanth, S. Meyn, and L. Brown · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. Williams · 1992
Earlier work this paper cites.
Discrete-time Bayesian adaptive control problems with complete information
O. Zane · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less real time
A. Moore and C. Atkeson · 1993
Earlier work this paper cites.
Markov Decision Processes
M. Puterman · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems
G. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
TD-Gammon, a self-teaching backgammon program, achieves master-level play
G. Tesauro · 1994
Earlier work this paper cites.
A short proof of the gittins index theorem
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
Optimal adaptive control of uncertain stochastic discrete linear systems
I. Rusnak · 1995
Earlier work this paper cites.
Adaptive dual control methods: An overview
B. Wittenmark · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
D. Bertsekas and J. Tsitsiklis · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. Bradtke and A. Barto · 1996
Earlier work this paper cites.
A comparison of direct and model-based reinforcement learning
C. G. Atkeson and J. C. Santamaria · 1997
Earlier work this paper cites.
Multitask learning
R. Caruana · 1997
Earlier work this paper cites.
Knightcap: A chess program that learns by combining TD( λ \lambda ) with game-tree search
J. Baxter, A. Tridgell, and L. Weaver · 1998
Earlier work this paper cites.
Elevator group control using multiple reinforcement learning agents
R. Crites and A. Barto · 1998
Earlier work this paper cites.
Bayesian Q-learning
R. Dearden, N. Friedman, and S. J. Russell · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. Kaelbling, M. Littman, and A. Cassandra · 1998
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 1998
Earlier work this paper cites.
Simulated-Based Methods for Markov Decision Processes
P. Marbach · 1998
Earlier work this paper cites.
Learning agents for uncertain environments (extended abstract)
S. Russell · 1998
Earlier work this paper cites.
An Introduction to Reinforcement Learning
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
Least-squares temporal difference learning
J. Boyan · 1999
Earlier work this paper cites.
Model based Bayesian exploration
R. Dearden, N. Friedman, and D. Andre · 1999
Earlier work this paper cites.
Efficient Bayesian parameter estimation in large discrete domains
N. Friedman and Y. Singer · 1999
Earlier work this paper cites.
Exploiting generative models in discriminative classifiers
T. Jaakkola and D. Haussler · 1999
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large Markov decision processes
M. Kearns, Y. Mansour, and A. Ng · 1999
Earlier work this paper cites.
Some PAC-Bayesian theorems
D. McAllester · 1999
Earlier work this paper cites.
A model of inductive bias learning
J. Baxter · 2000
Earlier work this paper cites.
Survey of adaptive dual control methods
N. M. Filatov and H. Unbehauen · 2000
Earlier work this paper cites.
Actor-Critic algorithms
V. Konda and J. Tsitsiklis · 2000
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Ng and S. Russell · 2000
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
M. Strens · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
J. Baxter and P. Bartlett · 2001
Earlier work this paper cites.
Experiments with infinite-horizon policy-gradient estimation
J. Baxter, P. Bartlett, and L. Weaver · 2001
Earlier work this paper cites.
Monte-Carlo algorithms for the improvement of finite-state stochastic controllers: Application to Bayes-adaptive Markov decision processes
M. Duff · 2001
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Cited alongside, same era.
Optimal Learning: Computational Procedures for Bayes-Adaptive Markov Decision Processes
M. Duff · 2002
Cited alongside, same era.
Sparse online greedy support vector regression
Y. Engel, S. Mannor, and R. Meir · 2002
Cited alongside, same era.
Artificial Intelligence: A Modern Approach (2nd Edition)
S. Russell and P. Norvig · 2002
Cited alongside, same era.
Learning with Kernels
B. Schölkopf and A. Smola · 2002
Cited alongside, same era.
Bayesian reinforcement learning in continuous POMDPs with Gaussian processes
P. Dallaire, C. Besse, S. Ross, and B. Chaib-draa · 2009
Later among the works it cites.
The infinite partially observable Markov decision process
F. Doshi-Velez · 2009
Later among the works it cites.
Near-Bayesian exploration in polynomial time
J. Kolter and A. Ng · 2009
Later among the works it cites.
A unifying framework for computational reinforcement learning theory
L. Li · 2009
Later among the works it cites.
Smarter sampling in model-based Bayesian reinforcement learning
P. Castro and D. Precup · 2010
Later among the works it cites.
Cooperative games with overlapping coalitions
G. Chalkiadakis, E. Elkinda, E. Markakis, M. Polukarov, and N. Jennings · 2010
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
R. Brafman and M. Tennenholtz · 2003
Cited alongside, same era.
Bayes meets Bellman: The Gaussian process approach to temporal difference learning
Y. Engel, S. Mannor, and R. Meir · 2003
Cited alongside, same era.
Adaptive control of nonlinear stochastic systems by particle filtering
A. Greenfield and A. Brockwell · 2003
Cited alongside, same era.
Least-squares policy iteration
M. Lagoudakis and R. Parr · 2003
Cited alongside, same era.
The cross entropy method for fast policy search
S. Mannor, R. Rubinstein, and Y. Gat · 2003
Cited alongside, same era.
Point-based value iteration: an anytime algorithm for POMDPs
J. Pineau, G. Gordon, and S. Thrun · 2003
Cited alongside, same era.
Percentile optimization for Markov decision processes with parameter uncertainty
E. Delage and S. Mannor · 2010
Later among the works it cites.
Nonparametric Bayesian policy priors for reinforcement learning
F. Doshi-Velez, D. Wingate, N. Roy, and J. Tenenbaum · 2010
Later among the works it cites.
Large-Scale Inference: Empirical Bayes Methods fro Estimation, Testing, and Prediction
B. Efron · 2010
Later among the works it cites.
PAC-Bayesian model selection for reinforcement learning
M. M. Fard and J. Pineau · 2010
Later among the works it cites.
Web-scale Bayesian click-through rate prediction for sponsored search advertising in Microsoft’s Bing search engine
T. Graepel, J.Q. Candela, T. Borchert, and R. Herbrich · 2010
Later among the works it cites.
Bayesian multi-task reinforcement learning
A. Lazaric and M. Ghavamzadeh · 2010
Later among the works it cites.
Linearly parameterized bandits
P. Rusmevichientong and J. N. Tsitsiklis · 2010
Later among the works it cites.
A modern Bayesian look at the multi-armed bandit
S. Scott · 2010
Later among the works it cites.
Variance-based rewards for approximate Bayesian reinforcement learning
V. Sorg, S. Sing, and R. Lewis · 2010
Later among the works it cites.
Integrating sample-based planning and model-based reinforcement learning
T. Walsh, S. Goschin, and M. Littman · 2010
Later among the works it cites.
Approaching Bayes-optimality using Monte-Carlo tree search
J. Asmuth and M. Littman · 2011
Later among the works it cites.
Apprenticeship learning about multiple intentions
M. Babes, V. Marivate, K. Subramanian, and M. Littman · 2011
Later among the works it cites.
An empirical evaluation of Thompson sampling
O. Chapelle and L. Li · 2011
Later among the works it cites.
Map inference for Bayesian inverse reinforcement learning
J. Choi and K. Kim · 2011
Later among the works it cites.
Bayesian multi-task inverse reinforcement learning
C. Dimitrakakis and C. Rothkopf · 2011
Later among the works it cites.
Reinforcement learning with limited reinforcement: Using Bayes risk for active learning in POMDPs
F. Doshi-Velez, J. Pineau, and N. Roy · 2011
Later among the works it cites.
PAC-Bayesian policy evaluation for reinforcement learning
M. M. Fard, J. Pineau, and C. Szepesvari · 2011
Later among the works it cites.
Computing a classic index for finite-horizon bandits
J. Niño-Mora · 2011
Later among the works it cites.
Bayesian Reinforcement Learning for POMDP-based Dialogue Systems
S Png · 2011
Later among the works it cites.
Bayesian reinforcement learning for POMDP-based dialogue systems
S. Png and J. Pineau · 2011
Later among the works it cites.
Approximate Dynamic Programming: Solving the curses of dimensionality (2nd Edition)
W. B. Powell · 2011
Later among the works it cites.
A Bayesian approach for learning and planning in partially observable Markov decision processes
S. Ross, J. Pineau, B. Chaib-draa, and P. Kreitmann · 2011
Later among the works it cites.
Hessian matrix distribution for Bayesian policy gradient reinforcement learning
N. Vien, H. Yu, and T. Chung · 2011
Later among the works it cites.
Analysis of Thompson sampling for the multi-armed bandit problem
S. Agrawal and N. Goyal · 2012
Later among the works it cites.
Near-optimal BRL using optimistic local transitions
M. Araya-Lopez, V. Thomas, and O. Buffet · 2012
Later among the works it cites.
Robust adaptive Markov decision processes: Planning with model uncertainty
L. Bertuccelli, A. Wu, and J. How · 2012
Later among the works it cites.
Tractable objectives for robust policy optimization
K. Chen and M. Bowling · 2012
Later among the works it cites.
Nonparametric Bayesian inverse reinforcement learning for multiple reward functions
J. Choi and K. Kim · 2012
Later among the works it cites.
The grand challenge of computer Go: Monte Carlo tree search and extensions
S. Gelly, L. Kocsis, M. Schoenauer, M. Sebag, D. Silver, C. Szepesvari, and O. Teytaud · 2012
Later among the works it cites.
Efficient Bayes-adaptive reinforcement learning using sample-based search
A. Guez, D. Silver, and P. Dayan · 2012
Later among the works it cites.
Model-based Bayesian Reinforcement Learning with Generalized Priors
J. Asmuth · 2013
Later among the works it cites.
Coordination in multiagent reinforcement learning: A Bayesian approach
G. Chalkiadakis and C. Boutilier · 2013
Later among the works it cites.
Bayesian policy gradient and actor-critic algorithms
M. Ghavamzadeh, Y. Engel, and M. Valko · 2013
Later among the works it cites.
A greedy approximation of Bayesian reinforcement learning with probably optimistic transition model
K. Kawaguchi and M. Araya-Lopez · 2013
Later among the works it cites.
Reinforcement learning in robotics: A survey
J. Kober, D. Bagnell, and J. Peters · 2013
Later among the works it cites.
(More) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Later among the works it cites.
Optimistic planning for belief-augmented Markov Decision Processes
R. Munos R. Fonteneau, L. Busoniu · 2013
Later among the works it cites.
Automatic ad format selection via contextual bandits
L. Tang, R. Rosales, A. Singh, and D. Agarwal · 2013
Later among the works it cites.
Thompson sampling for complex online problems
A. Gopalan, S. Mannor, and Y. Mansour · 2014
Later among the works it cites.
Stochastic regret minimization via Thompson sampling
S. Guha and K. Munagala · 2014
Later among the works it cites.
https://github.com/mcastron/BBRL/, 2015
Bbrl a c++ open-source library used to compare bayesian reinforcement learning algorithms · 2015
Later among the works it cites.
Bayesian optimal control of smoothly parameterized systems
Y. Abbasi-Yadkori and C. Szepesvari · 2015
Later among the works it cites.
Benchmarking for bayesian reinforcement learning
M. Castronovo, D. Ernst, and R. Fonteneau A. Couetoux · 2015
Later among the works it cites.
Thompson sampling for learning parameterized markov decision processes
A. Gopalan and S. Mannor · 2015
Later among the works it cites.
On the prior sensitivity of Thompson sampling
C. Liu and L. Li · 2015
Later among the works it cites.