Fetching the paper…
Reading the bibliography…
Partial observability is a common challenge in many reinforcement learning applications, which requires an agent to maintain memory, infer latent states, and integrate this past information into exploration.
On the definition of a family of automata
M. P. Schützenberger · 1961
Earlier work this paper cites.
Realizations by stochastic finite automata
J. W. Carlyle and A. Paz · 1971
Earlier work this paper cites.
The complexity of markov decision processes
C. H. Papadimitriou and J. N. Tsitsiklis · 1987
Earlier work this paper cites.
Acting under uncertainty: Discrete bayesian models for mobile-robot navigation
A. R. Cassandra, L. P. Kaelbling, and J. A. Kurien · 1996
Earlier work this paper cites.
Discrete-time, discrete-valued observable operator models: a tutorial
H. Jaeger · 1998
Earlier work this paper cites.
Planning treatment of ischemic heart disease with partially observable markov decision processes
M. Hauskrecht and H. Fraser · 2000
Earlier work this paper cites.
Observable operator models for discrete stochastic time series
H. Jaeger · 2000
Earlier work this paper cites.
Complexity of finite-horizon markov decision process problems
M. Mundhenk, J. Goldsmith, C. Lusena, and E. Allender · 2000
Earlier work this paper cites.
Predictive representations of state
M. L. Littman and R. S. Sutton · 2002
Earlier work this paper cites.
Reinforcement learning in pomdps without resets
E. Even-Dar, S. M. Kakade, and Y. Mansour · 2005
Earlier work this paper cites.
Learning nonsingular phylogenies and hidden markov models
E. Mossel and S. Roch · 2005
Cited alongside, same era.
Model-based bayesian reinforcement learning in partially observable domains
P. Poupart and N. Vlassis · 2008
Cited alongside, same era.
Bayes-adaptive pomdps
S. Ross, B. Chaib-draa, and J. Pineau · 2008
Cited alongside, same era.
Quasi-deterministic partially observable markov decision processes
C. Besse and B. Chaib-Draa · 2009
Cited alongside, same era.
Hilbert space embeddings of hidden markov models
L. Song, B. Boots, S. Siddiqi, G. J. Gordon, and A. Smola · 2010
Cited alongside, same era.
Closing the learning-planning loop with predictive state representations
B. Boots, S. M. Siddiqi, and G. J. Gordon · 2011
Cited alongside, same era.
On the computational complexity of stochastic controller optimization in pomdps
N. Vlassis, M. L. Littman, and D. Barber · 2012
Later among the works it cites.
Tensor decompositions for learning latent variable models
A. Anandkumar, R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky · 2014
Later among the works it cites.
Reinforcement learning of pomdps using spectral methods
K. Azizzadenesheli, A. Lazaric, and A. Anandkumar · 2016
Later among the works it cites.
A pac rl algorithm for episodic pomdps
Z. D. Guo, S. Doroudi, and E. Brunskill · 2016
Later among the works it cites.
Pac reinforcement learning with rich observations
A. Krishnamurthy, A. Agarwal, and J. Langford · 2016
Later among the works it cites.
Learning overcomplete hmms
V. Sharan, S. M. Kakade, P. S. Liang, and G. Valiant · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Faster teaching by pomdp planning
A. N. Rafferty, E. Brunskill, T. L. Griffiths, and P. Shafto · 2011
Cited alongside, same era.
A method of moments for mixture models and hidden markov models
A. Anandkumar, D. Hsu, and S. M. Kakade · 2012
Cited alongside, same era.
Deterministic pomdps revisited
B. Bonet · 2012
Cited alongside, same era.
A spectral algorithm for learning hidden markov models
D. Hsu, S. M. Kakade, and T. Zhang · 2012
Cited alongside, same era.
Iterative planning for deterministic qdec-pomdps
S. Bazinin and G. Shani · 2018
Later among the works it cites.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
N. Brown and T. Sandholm · 2018
Later among the works it cites.
Policy gradient in partially observable environments: Approximation and convergence
A. Kamyar, Y. Yue, and A. Anandkumar · 2018
Later among the works it cites.
A short note on concentration inequalities for random vectors with subgaussian norm
C. Jin, P. Netrapalli, R. Ge, S. M. Kakade, and M. I. Jordan · 2019
Later among the works it cites.