Fetching the paper…
Reading the bibliography…
In this work, we consider the regret minimization problem for reinforcement learning in latent Markov Decision Processes (LMDP).
The optimal control of partially observable markov processes over a finite horizon
R. D. Smallwood and E. J. Sondik · 1973
Earlier work this paper cites.
The complexity of Markov decision processes
C. H. Papadimitriou and J. N. Tsitsiklis · 1987
Earlier work this paper cites.
Matrix perturbation theory
G. W. Stewart · 1990
Earlier work this paper cites.
Memoryless policies: Theoretical limitations and practical results
M. L. Littman · 1994
Earlier work this paper cites.
Reinforcement learning algorithm for partially observable markov decision problems
T. Jaakkola, S. P. Singh, and M. I. Jordan · 1995
Earlier work this paper cites.
Learning policies for partially observable environments: Scaling up
M. L. Littman, A. R. Cassandra, and L. P. Kaelbling · 1995
Earlier work this paper cites.
PEGASUS: a policy search method for large MDPs and POMDPs
A. Y. Ng and M. Jordan · 2000
Earlier work this paper cites.
Predictive representations of state
M. L. Littman and R. S. Sutton · 2002
Earlier work this paper cites.
Predictive state representations: a new theory for modeling dynamical systems
S. Singh, M. R. James, and M. R. Rudary · 2004
Earlier work this paper cites.
Heuristic search value iteration for POMDPs
T. Smith and R. Simmons · 2004
Earlier work this paper cites.
Perseus: Randomized point-based value iteration for POMDPs
M. T. Spaan and N. Vlassis · 2005
Earlier work this paper cites.
Anytime point-based approximations for large POMDPs
J. Pineau, G. Gordon, and S. Thrun · 2006
Earlier work this paper cites.
k-means++ the advantages of careful seeding
D. Arthur and S. Vassilvitskii · 2007
Earlier work this paper cites.
On-line expectation–maximization algorithm for latent data models
O. Cappé and E. Moulines · 2009
Earlier work this paper cites.
Sensitivity analysis of POMDP value functions
S. Ross, M. Izadi, M. Mercer, and D. Buckeridge · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
M. E. Taylor and P. Stone · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2010
Cited alongside, same era.
An online spectral learning algorithm for partially observable nonlinear dynamical systems
B. Boots and G. J. Gordon · 2011
Cited alongside, same era.
Closing the learning-planning loop with predictive state representations
B. Boots, S. M. Siddiqi, and G. J. Gordon · 2011
Cited alongside, same era.
Finding optimal memoryless policies of POMDPs under the expected average reward criterion
Y. Li, B. Yin, and H. Xi · 2011
Cited alongside, same era.
A method of moments for mixture models and hidden markov models
Improving predictive state representations via gradient descent
N. Jiang, A. Kulesza, and S. Singh · 2016
Later among the works it cites.
PAC reinforcement learning with rich observations
A. Krishnamurthy, A. Agarwal, and J. Langford · 2016
Later among the works it cites.
PAC continuous state online multitask reinforcement learning with identification
Y. Liu, Z. Guo, and E. Brunskill · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Later among the works it cites.
On context-dependent clustering of bandits
C. Gentile, S. Li, P. Kar, A. Karatzoglou, G. Zappella, and E. Etrue · 2017
Later among the works it cites.
Contextual decision processes with low bellman rank are PAC-learnable
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Anandkumar, D. Hsu, and S. M. Kakade · 2012
Cited alongside, same era.
Momdps: a solution for modelling adaptive management problems
I. Chadès, J. Carwardine, T. Martin, S. Nicol, R. Sabbadin, and O. Buffet · 2012
Cited alongside, same era.
A spectral algorithm for learning hidden markov models
D. Hsu, S. M. Kakade, and T. Zhang · 2012
Cited alongside, same era.
Sample complexity of multi-task reinforcement learning
E. Brunskill and L. Li · 2013
Cited alongside, same era.
Tensor decompositions for learning latent variable models
A. Anandkumar, R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky · 2014
Cited alongside, same era.
Online clustering of bandits
C. Gentile, S. Li, and G. Zappella · 2014
Cited alongside, same era.
Latent bandits
O.-A. Maillard and S. Mannor · 2014
Cited alongside, same era.
N. Jiang, A. Krishnamurthy, A. Agarwal, J. Langford, and R. E. Schapire · 2017
Later among the works it cites.
On oracle-efficient PAC RL with rich observations
C. Dann, N. Jiang, A. Krishnamurthy, A. Agarwal, J. Langford, and R. E. Schapire · 2018
Later among the works it cites.
Markov decision processes with continuous side information
A. Modi, N. Jiang, S. Singh, and A. Tewari · 2018
Later among the works it cites.
Multi-model markov decision processes
L. N. Steimle, D. L. Kaufman, and B. T. Denton · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Computation of weighted sums of rewards for concurrent MDPs
P. Buchholz and D. Scheftelowitsch · 2019
Later among the works it cites.
Provably efficient rl with rich observations via latent state decoding
S. Du, A. Krishnamurthy, N. Jiang, A. Agarwal, M. Dudik, and J. Langford · 2019
Later among the works it cites.
Sample-efficient reinforcement learning of undercomplete POMDPs
C. Jin, S. M. Kakade, A. Krishnamurthy, and Q. Liu · 2020
Later among the works it cites.
The EM algorithm gives sample-optimality for learning mixtures of well-separated gaussians
J. Kwon and C. Caramanis · 2020
Later among the works it cites.
EM converges for a mixture of many linear regressions
J. Kwon and C. Caramanis · 2020
Later among the works it cites.
On the minimax optimality of the EM algorithm for learning two-component mixed linear regression
J. Kwon, N. Ho, and C. Caramanis · 2020
Later among the works it cites.