Fetching the paper…
Reading the bibliography…
In this paper, for POMDPs, we provide the convergence of a Q learning algorithm for control policies using a finite history of past observations and control actions, and, consequentially, we establish near optimality of such limit Q functions under explicit filter stability conditions.
On the notion of recurrence in discrete stochastic processes
M. Kac · 1947
Earlier work this paper cites.
Central limit theorem for nonstationary Markov chains. i
R.L. Dobrushin · 1956
Earlier work this paper cites.
Probability Measures on Metric Spaces
K.R. Parthasarathy · 1967
Earlier work this paper cites.
Incomplete information in Markovian decision models
D. Rhenius · 1974
Earlier work this paper cites.
Reduction of a controlled Markov model with incomplete data to a problem with complete information in the case of Borel state and control spaces
A.A. Yushkevich · 1976
Earlier work this paper cites.
Controlled Markov processes with arbitrary numerical criteria
E. A. Feinberg · 1982
Earlier work this paper cites.
Adaptive Markov Control Processes
O. Hernández-Lerma · 1989
Earlier work this paper cites.
A survey of algorithmic methods for partially observed Markov decision processes
W.S. Lovejoy · 1991
Earlier work this paper cites.
A survey of solution techniques for the partially observed Markov decision process
C.C. White · 1991
Earlier work this paper cites.
Memory approaches to reinforcement learning in non-Markovian domains
Long-Ji Lin and Tom M Mitchell · 1992
Earlier work this paper cites.
Learning without state-estimation in partially observable markovian decision processes
Satinder P. Singh, Tommi Jaakkola, and Michael I. Jordan · 1994
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
Tommi T. Jaakkola, M. I. Jordan, and S. P. Singh · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
Finite-memory suboptimal design for partially observed markov decision processes
C. C. White-III and W. T. Scherer · 1994
Earlier work this paper cites.
Reinforcement learning algorithm for partially observable markov decision problems
Tommi Jaakkola, Satinder P. Singh, and Michael I. Jordan · 1995
Earlier work this paper cites.
Discrete-Time Markov Control Processes: Basic Optimality Criteria
O. Hernandez-Lerma and J. B. Lasserre · 1996
Cited alongside, same era.
Reinforcement learning with selective perception and hidden state
Andrew McCallum · 1997
Cited alongside, same era.
Convergence of probability measures
P. Billingsley · 1999
Cited alongside, same era.
An improved grid-based approximation algorithm for POMDPs
R. Zhou and E.A. Hansen · 2001
Cited alongside, same era.
Learning rates for q-learning
Eyal Even-Dar, Yishay Mansour, and Peter Bartlett · 2003
Cited alongside, same era.
Perseus: Randomized point-based value iteration for pomdps
N. Vlassis and M. T. J. Spaan · 2005
Cited alongside, same era.
Partially observable total-cost Markov decision process with weakly continuous transition probabilities
E.A. Feinberg, P.O. Kasyanov, and M.Z. Zgurovsky · 2016
Later among the works it cites.
Partially observed Markov decision processes: from filtering to controlled sensing
V. Krishnamurthy · 2016
Later among the works it cites.
On the asymptotic optimality of finite approximations to markov decision processes with borel spaces
N. Saldi, S. Yüksel, and T. Linder · 2017
Later among the works it cites.
Converse results on filter stability criteria and stochastic non-linear observability
C. McDonald and S. Yüksel · 2018
Later among the works it cites.
Weak feller property of non-linear filters
A. D. Kara, N. Saldi, and S. Yüksel · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Pineau, G. Gordon, and S. Thrun · 2006
Cited alongside, same era.
Point-based value iteration for continuous pomdps
J. M. Porta, N. Vlassis, M. T. J. Spaan, and P. Poupart · 2006
Cited alongside, same era.
Optimal transport: old and new
C. Villani · 2008
Cited alongside, same era.
On near optimality of the set of finite-state controllers for average cost pomdp
Huizhen Yu and Dimitri P Bertsekas · 2008
Cited alongside, same era.
A density projection approach to dimension reduction for continuous-state POMDPs
E. Zhou, M. C. Fu, and S. I. Marcus · 2008
Cited alongside, same era.
Solving continuous-state POMDPs via density projection
E. Zhou, M. C. Fu, and S. I. Marcus · 2010
Cited alongside, same era.
Observability and filter stability for partially observed markov processes
C. McDonald and S. Yüksel · 2019
Later among the works it cites.
Approximate information state for partially observed systems
J. Subramanian and A. Mahajan · 2019
Later among the works it cites.
Martin Wainwright · 2019
Later among the works it cites.
Near optimality of finite memory feedback policies in partially observed markov decision processes
Ali Devran Kara and Serdar Yuksel · 2020
Later among the works it cites.
Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Later among the works it cites.
Exponential filter stability via Dobrushin’s coefficient
C. McDonald and S. Yüksel · 2020
Later among the works it cites.
Finite model approximations for partially observed markov decision processes with discounted cost
N. Saldi, S. Yüksel, and T. Linder · 2020
Later among the works it cites.
Planning in observable POMDPs in quasipolynomial time
N. Golowich, A. Moitra, and D. Rohatgi · 2022
Closest in time.
C. McDonald and S. Yüksel · 2022
Closest in time.