Fetching the paper…
Reading the bibliography…
Leveraging an equivalence property in the state-space of a Markov Decision Process (MDP) has been investigated in several studies.
Probability inequalities for sums of bounded random variables
W. Hoeffding · 1963
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and H. Robbins · 1985
Earlier work this paper cites.
Optimal adaptive policies for Markov decision processes
A. N. Burnetas and M. N. Katehakis · 1997
Earlier work this paper cites.
Model reduction techniques for computing approximately optimal solutions for Markov decision processes
T. Dean, R. Givan, and S. Leach · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Finite time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Equivalence notions and model minimization in Markov decision processes
R. Givan, T. Dean, and M. Greig · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. Kakade · 2003
Earlier work this paper cites.
Inequalities for the L1 deviation of the empirical distribution
T. Weissman, E. Ordentlich, G. Seroussi, S. Verdu, and M. J. Weinberger · 2003
Earlier work this paper cites.
Metrics for finite Markov decision processes
N. Ferns, P. Panangaden, and D. Precup · 2004
Earlier work this paper cites.
Approximate homomorphisms: A framework for non-exact minimization in Markov decision processes
B. Ravindran and A. G. Barto · 2004
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
L. Li, T. J. Walsh, and M. L. Littman · 2006
Earlier work this paper cites.
Efficient reinforcement learning with relocatable action models
B. R. Leffler, M. L. Littman, and T. Edmunds · 2007
Cited alongside, same era.
Self-normalized processes: Limit theory and Statistical Applications
V. H. Peña, T. L. Lai, and Q.-M. Shao · 2008
Cited alongside, same era.
An analysis of model-based interval estimation for Markov decision processes
A. L. Strehl and M. L. Littman · 2008
Cited alongside, same era.
The adaptive k-meteorologists problem and its application to structure learning and feature selection in reinforcement learning
C. Diuk, L. Li, and B. R. Leffler · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and Cs. Szepesvári · 2011
Selecting near-optimal approximate state representations in reinforcement learning
R. Ortner, O.-A. Maillard, and D. Ryabko · 2014
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Later among the works it cites.
ASAP-UCT: Abstraction of state-action pairs in UCT
A. Anand, A. Grover, Mausam, and P. Singla · 2015
Later among the works it cites.
Near optimal behavior via approximate state abstraction
D. Abel, D. Hershkowitz, and M. L. Littman · 2016
Later among the works it cites.
Consistent algorithms for clustering time series
A. Khaleghi, D. Ryabko, J. Mary, and P. Preux · 2016
Later among the works it cites.
Efficient Bayesian clustering for reinforcement learning
T. Mandel, Y.-E. Liu, E. Brunskill, and Z. Popovic · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bisimulation metrics for continuous Markov decision processes
N. Ferns, P. Panangaden, and D. Precup · 2011
Cited alongside, same era.
Sample complexity of multi-task reinforcement learning
E. Brunskill and L. Li · 2013
Cited alongside, same era.
Model selection in Markovian processes
A. Hallak, D. Di-Castro, and S. Mannor · 2013
Cited alongside, same era.
Adaptive aggregation for reinforcement learning in average reward Markov decision processes
R. Ortner · 2013
Cited alongside, same era.
How hard is my MDP? “the distribution-norm to the rescue”
O.-A. Maillard, T. A. Mann, and S. Mannor · 2014
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
C. Dann, T. Lattimore, and E. Brunskill · 2017
Later among the works it cites.
Minimax regret bounds for reinforcement learning
M. Gheshlaghi Azar, I. Osband, and R. Munos · 2017
Later among the works it cites.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
R. Fruit, M. Pirotta, A. Lazaric, and R. Ortner · 2018
Later among the works it cites.
Variance-aware regret bounds for undiscounted reinforcement learning in MDPs
M. S. Talebi and O.-A. Maillard · 2018
Later among the works it cites.
Mathematics of statistical sequential decision making
O.-A. Maillard · 2019
Closest in time.