Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) in Markov decision processes (MDPs) with large state spaces is a challenging problem.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Memoryless policies: Theoretical limitations and practical results
Michael L. Littman · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
D. Bertsekas and J. Tsitsiklis · 1996
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J. Walsh, and Michael L. Littman · 2006
Earlier work this paper cites.
Point-based value iteration for continuous pomdps
Josep M Porta, Nikos Vlassis, Matthijs TJ Spaan, and Pascal Poupart · 2006
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Freedman’s inequality for matrix martingales
Joel A Tropp et al · 2011
Earlier work this paper cites.
A method of moments for mixture models and hidden Markov models
Animashree Anandkumar, Daniel Hsu, and Sham M Kakade · 2012
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Cited alongside, same era.
On the computational complexity of stochastic controller optimization in pomdps
Nikos Vlassis, Michael L Littman, and David Barber · 2012
Cited alongside, same era.
A gang of bandits
Nicolo Cesa-Bianchi, Claudio Gentile, and Giovanni Zappella · 2013
Cited alongside, same era.
Sequential transfer in multi-armed bandit with finite set of models
Mohammad Gheshlaghi-Azar, Alessandro Lazaric, and Emma Brunskill · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Adaptive aggregation for reinforcement learning in average reward Markov decision processes
Ronald Ortner · 2013
Online clustering of bandits
Claudio Gentile, Shuai Li, and Giovanni Zappella · 2014
Later among the works it cites.
Uniform chernoff and dvoretzky-kiefer-wolfowitz-type inequalities for Markov chains and related processes
Aryeh Kontorovich, Roi Weiss, et al · 2014
Later among the works it cites.
Latent bandits
Odalric-Ambrym Maillard and Shie Mannor · 2014
Later among the works it cites.
Fast and guaranteed tensor decomposition via sketching
Yining Wang, Hsiao-Yu Tung, Alex J Smola, and Anima Anandkumar · 2015
Later among the works it cites.
Low-rank bandits with latent mixtures
Aditya Gopalan, Odalric-Ambrym Maillard, and Mohammadi Zaki · 2016
Closest in time.
A PAC rl algorithm for episodic POMDPs
Zhaohan Daniel Guo, Shayan Doroudi, and Emma Brunskill · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey of point-based pomdp solvers
Guy Shani, Joelle Pineau, and Robert Kaplow · 2013
Cited alongside, same era.
Nonparametric estimation of multi-view latent variable models
Le Song, Animashree Anandkumar, Bo Dai, and Bo Xie · 2013
Cited alongside, same era.
Tensor decompositions for learning latent variable models
Animashree Anandkumar, Rong Ge, Daniel Hsu, Sham M Kakade, and Matus Telgarsky · 2014
Cited alongside, same era.
Partial monitoring—classification, regret bounds, and algorithms
Gábor Bartók, Dean P Foster, Dávid Pál, Alexander Rakhlin, and Csaba Szepesvári · 2014
Cited alongside, same era.
Open problem: Approximate planning of POMDPs in the class of memoryless policies
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar
Cited in the paper.
Reinforcement learning of POMDPs using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar
Cited in the paper.
Contextual decision processes with low bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2016
Closest in time.
PAC reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Closest in time.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Closest in time.
Ubev-a more practical algorithm for episodic rl with near-optimal PAC and regret guarantees
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Closest in time.