Fetching the paper…
Reading the bibliography…
In reinforcement learning the Q-values summarize the expected future rewards that the agent will attain.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
Dynamic programming
R. Bellman · 1957
Earlier work this paper cites.
Information theory and statistical mechanics
E. T. Jaynes · 1957
Earlier work this paper cites.
A general theory of measurement applications to utility
J. Pfanzag · 1959
Earlier work this paper cites.
Value of information lotteries
R. A. Howard · 1967
Earlier work this paper cites.
Decision Analysis: Introductory Lectures on Choices under Uncertainty
H. Raiffa · 1968
Earlier work this paper cites.
Rules for ordering uncertain prospects
J. Hadar and W. R. Russell · 1969
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
R. J. Williams and J. Peng · 1991
Earlier work this paper cites.
Reinforcement Learning: an Introduction
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
M. Strens · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Dynamic programming and optimal control
D. P. Bertsekas · 2005
Earlier work this paper cites.
Prediction, learning, and games
N. Cesa-Bianchi and G. Lugosi · 2006
Earlier work this paper cites.
Invariant utility functions and certain equivalent transformations
A. E. Abbas · 2007
Earlier work this paper cites.
Theory of games and economic behavior (commemorative edition)
J. Von Neumann and O. Morgenstern · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
A. L. Strehl and M. L. Littman · 2008
Earlier work this paper cites.
Near-Bayesian exploration in polynomial time
J. Z. Kolter and A. Y. Ng · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Variance-based rewards for approximate Bayesian reinforcement learning
J. Sorg, S. Singh, and R. L. Lewis · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Cited alongside, same era.
Near-optimal BRL using optimistic local transitions
M. Araya, O. Buffet, and V. Thomas · 2012
Cited alongside, same era.
Dynamic policy programming
M. G. Azar, V. Gómez, and H. J. Kappen · 2012
Cited alongside, same era.
Elements of information theory
T. M. Cover and J. A. Thomas · 2012
Cited alongside, same era.
ECOS: An SOCP solver for embedded systems
A. Domahidi, E. Chu, and S. Boyd · 2013
Cited alongside, same era.
(More) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. Van Roy · 2013
Cited alongside, same era.
Boltzmann exploration done right
N. Cesa-Bianchi, C. Gentile, G. Neu, and G. Lugosi · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans · 2017
Later among the works it cites.
SCS: Splitting conic solver, version 2.0.2
B. O’Donoghue, E. Chu, N. Parikh, and S. Boyd · 2017
Later among the works it cites.
Combining policy gradient and Q-learning
B. O’Donoghue, R. Munos, K. Kavukcuoglu, and V. Mnih · 2017
Later among the works it cites.
Deep exploration via randomized value functions
I. Osband, D. Russo, Z. Wen, and B. Van Roy · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Osband, B. Van Roy, and Z. Wen · 2014
Cited alongside, same era.
Markov decision processes: Discrete stochastic dynamic programming
M. L. Puterman · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
D. Russo and B. Van Roy · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
C. Dann and E. Brunskill · 2015
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
R. Fox, A. Pakman, and N. Tishby · 2015
Cited alongside, same era.
Bayesian reinforcement learning: A survey
M. Ghavamzadeh, S. Mannor, J. Pineau, and A. Tamar · 2015
Cited alongside, same era.
Later among the works it cites.
Gaussian-Dirichlet posterior dominance in sequential learning
I. Osband and B. Van Roy · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning
I. Osband and B. Van Roy · 2017
Later among the works it cites.
A short variational proof of equivalence between policy gradients and soft Q learning
P. H. Richemond and B. Maginnis · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. M. N. Heess, and M. Riedmiller · 2018
Closest in time.
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Closest in time.
Reinforcement learning and control as probabilistic inference: Tutorial and review
S. Levine · 2018
Closest in time.
The uncertainty bellman equation and exploration
B. O’Donoghue, I. Osband, R. Munos, and V. Mnih · 2018
Closest in time.
A tutorial on thompson sampling
D. J. Russo, B. Van Roy, A. Kazerouni, I. Osband, and Z. Wen · 2018
Closest in time.
Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction
G. Li, Y. Wei, Y. Chi, Y. Gu, and Y. Chen · 2020
Closest in time.
Stochastic matrix games with bandit feedback
B. O’Donoghue, T. Lattimore, and I. Osband · 2020
Closest in time.
Making sense of reinforcement learning and probabilistic inference
B. O’Donoghue, I. Osband, and C. Ionescu · 2020
Closest in time.
Almost optimal model-free reinforcement learningvia reference-advantage decomposition
Z. Zhang, Y. Zhou, and X. Ji · 2020
Closest in time.
Operator splitting for a homogeneous embedding of the linear complementarity problem
B. O’Donoghue · 2021
Closest in time.