Fetching the paper…
Reading the bibliography…
We study reinforcement learning for partially observed Markov decision processes (POMDPs) with infinite observation and state spaces, which remains less investigated theoretically.
Dota 2 with large scale deep reinforcement learning
Berner, C · 1912
Earlier work this paper cites.
An approach to inequalities for the distributions of infinite-dimensional martingales
Pinelis, I · 1992
Earlier work this paper cites.
Optimum bounds for the distributions of martingales in Banach spaces
Pinelis, I · 1994
Earlier work this paper cites.
Acting under uncertainty: Discrete bayesian models for mobile-robot navigation
Cassandra, A. R · 1996
Earlier work this paper cites.
Learning with kernels
Smola, A. J · 1998
Earlier work this paper cites.
Planning treatment of ischemic heart disease with partially observable markov decision processes
Hauskrecht, M · 2000
Earlier work this paper cites.
Observable operator models for discrete stochastic time series
Jaeger, H · 2000
Earlier work this paper cites.
Sample-efficient reinforcement learning of undercomplete POMDPs
Jin, C · 2006
Earlier work this paper cites.
Information theoretic regret bounds for online nonlinear control
Kakade, S · 2006
Earlier work this paper cites.
PC-PG: Policy cover directed exploration for provable policy gradient learning
Agarwal, A · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Auer, P · 2008
Earlier work this paper cites.
Causal inference in statistics: An overview
Pearl, J · 2009
Earlier work this paper cites.
Faster teaching by POMDP planning
Rafferty, A. N · 2011
Earlier work this paper cites.
A spectral algorithm for learning hidden Markov models
Hsu, D · 2012
Cited alongside, same era.
On the computational complexity of stochastic controller optimization in POMDPs
Vlassis, N · 2012
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V · 2015
Cited alongside, same era.
Reinforcement learning of POMDPs using spectral methods
Azizzadenesheli, K · 2016
Cited alongside, same era.
A PAC RL algorithm for episodic POMDPs
Guo, Z. D · 2016
Cited alongside, same era.
Generalization and exploration via randomized value functions
Osband, I · 2016
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Later among the works it cites.
A selective review of negative control methods in epidemiology
Shi, X · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Yang, L · 2020
Later among the works it cites.
Bennett, A · 2021
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in RL
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Silver, D · 2016
Cited alongside, same era.
Markov decision processes with unobserved confounders: A causal approach
Zhang, J · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Cited alongside, same era.
Mastering the game of Go without human knowledge
Silver, D · 2017
Cited alongside, same era.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Brown, N · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Jin, C · 2018
Cited alongside, same era.
Du, S. S · 2021
Later among the works it cites.
Causal inference under unmeasured confounding with negative controls: A minimax learning approach
Kallus, N · 2021
Later among the works it cites.
Model-free learning for two-player zero-sum partially observable Markov games with perfect recall
Kozuno, T · 2021
Later among the works it cites.
RL for latent MDPs: Regret guarantees and a lower bound
Kwon, J · 2021
Later among the works it cites.
A spectral approach to off-policy evaluation for POMDPs
Nair, Y · 2021
Later among the works it cites.
Sublinear regret for learning POMDPs
Xiong, Y · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture Markov decision processes
Zhou, D · 2021
Later among the works it cites.
Planning in observable POMDPs in quasipolynomial time
Golowich, N · 2022
Closest in time.