Fetching the paper…
Reading the bibliography…
Reinforcement learning in partially observed Markov decision processes (POMDPs) faces two challenges.
A short note on concentration inequalities for random vectors with subgaussian norm
Jin, C · 1902
Earlier work this paper cites.
The optimal control of partially observable Markov processes
Sondik, E. J · 1971
Earlier work this paper cites.
The complexity of Markov decision processes
Papadimitriou, C. H · 1987
Earlier work this paper cites.
Empirical Processes in M-estimation
Geer, S. A · 2000
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Measuring statistical dependence with Hilbert-Schmidt norms
Gretton, A · 2005
Earlier work this paper cites.
Sample-efficient reinforcement learning of undercomplete POMDPs
Jin, C · 2006
Earlier work this paper cites.
Zhang, T · 2006
Earlier work this paper cites.
A Hilbert space embedding for distributions
Smola, A · 2007
Earlier work this paper cites.
Learning for control from multiple demonstrations
Coates, A · 2008
Earlier work this paper cites.
Dynamical variational autoencoders: A comprehensive review
Girin, L · 2008
Earlier work this paper cites.
A method of moments for mixture models and hidden markov models
Anandkumar, A · 2012
Earlier work this paper cites.
On the computational complexity of stochastic controller optimization in POMDPs
Vlassis, N · 2012
Cited alongside, same era.
Playing Atari with deep reinforcement learning
Mnih, V · 2013
Cited alongside, same era.
A survey of point-based POMDP solvers
Shani, G · 2013
Cited alongside, same era.
Deep recurrent Q-learning for partially observable MDPs
Hausknecht, M · 2015
Cited alongside, same era.
Supervised learning for dynamical system learning
Hefny, A · 2015
Cited alongside, same era.
Recurrent reinforcement learning: A hybrid approach
Li, X · 2015
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D · 2017
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J · 2019
Later among the works it cites.
Flambe: Structural complexity and representation learning of low rank MDPs
Agarwal, A · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Later among the works it cites.
Model-free representation learning and exploration in low-rank MDPs
Modi, A · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Mnih, V · 2015
Cited alongside, same era.
Reinforcement learning of POMDPs using spectral methods
Azizzadenesheli, K · 2016
Cited alongside, same era.
A PAC RL algorithm for episodic POMDPs
Guo, Z. D · 2016
Cited alongside, same era.
Learning to navigate in complex environments
Mirowski, P · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D · 2016
Cited alongside, same era.
Learning to filter with predictive state inference machines
Sun, W · 2016
Cited alongside, same era.
Representation learning for online and offline RL in low-rank MDPs
Uehara, M · 2021
Later among the works it cites.
Sample-efficient reinforcement learning for POMDPs with linear function approximations
Cai, Q · 2022
Closest in time.
Learning to control partially observed systems with finite memory
Cayci, S · 2022
Closest in time.
Provable reinforcement learning with a short-term memory
Efroni, Y · 2022
Closest in time.
When is partially observable reinforcement learning not scary?
Liu, Q · 2022
Closest in time.