Fetching the paper…
Reading the bibliography…
Partially Observable Markov Decision Processes (POMDPs) are rich environments often used in machine learning.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Cited alongside, same era.
Inverse reinforcement learning in partially observable environments
Jaedeug Choi and Kee-Eung Kim · 2011
Cited alongside, same era.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…