Fetching the paper…
Reading the bibliography…
It is well known that for any finite state Markov decision process (MDP) there is a memoryless deterministic policy that maximizes the expected reward.
Dynamic programming
Richard Bellman · 1957
Earlier work this paper cites.
Dynamic Programming and Markov Processes
Ronald A. Howard · 1960
Earlier work this paper cites.
The optimal control of partially observable Markov processes over the infinite horizon: Discounted costs
Edward J. Sondik · 1978
Earlier work this paper cites.
Introduction to Stochastic Dynamic Programming: Probability and Mathematical
Sheldon M. Ross · 1983
Earlier work this paper cites.
Reinforcement learning with perceptual aliasing: The perceptual distinctions approach
Lonnie Chrisman · 1992
Earlier work this paper cites.
Learning without state-estimation in partially observable Markovian decision processes
Satinder P. Singh, Tommi Jaakkola, and Michael I. Jordan · 1994
Earlier work this paper cites.
Reinforcement learning algorithm for partially observable Markov decision problems
Tommi Jaakkola, Satinder P. Singh, and Michael I. Jordan · 1995
Cited alongside, same era.
Learning policies for partially observable environments: Scaling up
Michael L. Littman, Anthony R. Cassandra, and Leslie Pack Kaelbling · 1995
Cited alongside, same era.
Approximating optimal policies for partially observable stochastic domains
Ronald Parr and Stuart Russell · 1995
Cited alongside, same era.
Reinforcement Learning with Selective Perception and Hidden State
Andrew Kachites McCallum · 1996
Cited alongside, same era.
Finite-memory Control of Partially Observable Systems
Eric Anton Hansen · 1998
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Cited alongside, same era.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L. Bartlett · 2001
Later among the works it cites.
How the Body Shapes the Way We Think: A New View of Intelligence
Rolf Pfeifer and Josh C. Bongard · 2006
Later among the works it cites.
Neighborliness of marginal polytopes
Thomas Kahle · 2010
Later among the works it cites.
Exponential Families with Incompatible Statistics and Their Entropy Distance
Stephan Weis · 2010
Later among the works it cites.
Selection criteria for neuromanifolds of stochastic dynamics
Nihat Ay, Guido Montúfar, and Johannes Rauh · 2013
Later among the works it cites.
Geometry and expressive power of conditional restricted Boltzmann machines
Guido Montúfar, Nihat Ay, and Keyan Ghazi-Zahedi · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Closest in time.