Fetching the paper…
Reading the bibliography…
We propose a theoretical framework for approximate planning and learning in partially observed systems.
Learning causal state representations of partially observable environments, 2019
A. Zhang, Z. C. Lipton, L. Pineda, K. Azizzadenesheli, A. Anandkumar, L. Itti, J. Pineau, and T. Furlanello · 1906
Earlier work this paper cites.
Decision Processes , chapter Towards an Economic Theory of Organization and Information
J. Marschak · 1954
Earlier work this paper cites.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Linear automaton transformations
A. Nerode · 1958
Earlier work this paper cites.
Team decision problems
R. Radner · 1962
Earlier work this paper cites.
Optimal control of Markov processes with incomplete state information
K. J. Aström · 1965
Earlier work this paper cites.
Theory of Self-Adaptive Control Systems , chapter Admissible Adaptive Control, pages 14–18
H. Kwakernaak · 1965
Earlier work this paper cites.
Sufficient statistics in the optimal control of stochastic systems
C. Striebel · 1965
Earlier work this paper cites.
Optimum maintenance with incomplete information
J. E. Eckles · 1968
Earlier work this paper cites.
A counterexample in stochastic optimum control
H. S. Witsenhausen · 1968
Earlier work this paper cites.
Introduction to Stochastic Control Theory
K. J. Aström · 1970
Earlier work this paper cites.
Information pattern for linear discrete-time models with stochastic coefficients
T. Bohlin · 1970
Earlier work this paper cites.
Separation of estimation and control for discrete time systems
H. S. Witsenhausen · 1971
Earlier work this paper cites.
An example of interaction between information and control: The transparency of a game
J.-M. Bismut · 1972
Earlier work this paper cites.
Information states for linear stochastic systems
M. Davis and P. Varaiya · 1972
Earlier work this paper cites.
Economic theory of teams
J. Marschak and R. Radner · 1972
Earlier work this paper cites.
The optimal control of partially observable Markov processes over a finite horizon
R. D. Smallwood and E. J. Sondik · 1973
Earlier work this paper cites.
Solution of some nonclassical lqg stochastic decision problems
N. Sandell and M. Athans · 1974
Earlier work this paper cites.
Convergence of discretization procedures in dynamic programming
D. Bertsekas · 1975
Earlier work this paper cites.
Dynamic programming approach to decentralized stochastic control problems
T. Yoshikawa · 1975
Earlier work this paper cites.
Some remarks on the concept of state
H. S. Witsenhausen · 1976
Earlier work this paper cites.
Survey of decentralized control methods for large scale systems
N. Sandell, P. Varaiya, M. Athans, and M. Safonov · 1978
Earlier work this paper cites.
Approximations of dynamic programs, I
W. Whitt · 1978
Earlier work this paper cites.
Optimal causal coding-decoding problems
J. C. Walrand and P. Varaiya · 1983
Earlier work this paper cites.
Toward a quantitative theory of self-generated complexity
P. Grassberger · 1986
Earlier work this paper cites.
Stochastic Systems: Estimation, Identification and Adaptive Control
P. R. Kumar and P. Varaiya · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, R. J. Williams, et al · 1986
Earlier work this paper cites.
Decentralized optimal control of Markov chains with a common past information set
M. Aicardi, F. Davoli, and R. Minciardi · 1987
Earlier work this paper cites.
Algorithms for Partially Observable Markov Decision Processes
H.-T. Cheng · 1988
Earlier work this paper cites.
Complexity and forecasting in dynamical systems
P. Grassberger · 1988
Earlier work this paper cites.
Inferring statistical complexity
J. P. Crutchfield and K. Young · 1989
Earlier work this paper cites.
Probablity Metrics and the Stability of Stochastic Models
S. T. Rachev · 1991
Earlier work this paper cites.
Closed-loop control with delayed information
E. Altman and P. Nain · 1992
Earlier work this paper cites.
Overcoming incomplete perception with utile distinction memory
R. A. McCallum · 1993
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
A. R. Cassandra, L. P. Kaelbling, and M. L. Littman · 1994
Earlier work this paper cites.
Memoryless policies: Theoretical limitations and practical results
M. L. Littman · 1994
Earlier work this paper cites.
Reinforcement learning algorithm for partially observable markov decision problems
T. Jaakkola, S. P. Singh, and M. I. Jordan · 1995
Earlier work this paper cites.
Reinforcement learning of non-markov decision processes
S. D. Whitehead and L.-J. Lin · 1995
Earlier work this paper cites.
Neuro-dynamic Programming
D. Bertsekas and J. Tsitsiklis · 1996
Earlier work this paper cites.
Planning in stochastic domains: Problem characteristics and approximation
N. Zhang and W. Liu · 1996
Earlier work this paper cites.
Stochastic approximation with two time scales
V. S. Borkar · 1997
Earlier work this paper cites.
Incremental pruning: A simple, fast, exact method for partially observable Markov decision processes
A. Cassandra, M. L. Littman, and N. L. Zhang · 1997
Earlier work this paper cites.
An improved policy iteratioll algorithm for partially observable MDPs
E. A. Hansen · 1997
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
How does the value function of a Markov decision process depend on the transition probabilities?
A. Müller · 1997
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
A. Müller · 1997
Earlier work this paper cites.
A separation theorem for periodic sharing information patterns in decentralized control
J. M. Ooi, S. M. Verbout, J. T. Ludwig, and G. W. Wornell · 1997
Earlier work this paper cites.
Solving POMDPs by searching in policy space
E. A. Hansen · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Using eligibility traces to find the best memoryless policy in partially observable markov decision processes
J. Loch and S. P. Singh · 1998
Earlier work this paper cites.
Markov decision processes with noise-corrupted and delayed state observations
J. L. Bander and C. C. White · 1999
Earlier work this paper cites.
Learning finite-state controllers for partially observable environments
N. Meuleau, L. Peshkin, K.-E. Kim, and L. P. Kaelbling · 1999
Cited alongside, same era.
The complexity of optimal queuing network control
C. H. Papadimitriou and J. N. Tsitsiklis · 1999
Cited alongside, same era.
Experimental results on learning stochastic memoryless policies for partially observable markov decision processes
J. K. Williams and S. P. Singh · 1999
Cited alongside, same era.
Observable operator models for discrete stochastic time series
H. Jaeger · 2000
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Cited alongside, same era.
Infinite-horizon policy-gradient estimation
J. Baxter and P. L. Bartlett · 2001
On the locality of action domination in sequential decision making
E. Rachelson and M. G. Lagoudakis · 2010
Later among the works it cites.
Recurrent policy gradients
D. Wierstra, A. Förster, J. Peters, and J. Schmidhuber · 2010
Later among the works it cites.
Closing the learning-planning loop with predictive state representations
B. Boots, S. M. Siddiqi, and G. J. Gordon · 2011
Later among the works it cites.
Information theory: coding theorems for discrete memoryless systems
I. Csiszar and J. Körner · 2011
Later among the works it cites.
Bisimulation metrics for continuous Markov decision processes
N. Ferns, P. Panangaden, and D. Precup · 2011
Later among the works it cites.
Finding optimal memoryless policies of POMDPs under the expected average reward criterion
Y. Li, B. Yin, and H. Xi · 2011
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Computational Mechanics: Pattern and prediction, structure and simplicity
C. R. Shalizi and J. P. Crutchfield · 2001
Cited alongside, same era.
Reinforcement learning with long short-term memory
B. Bakker · 2002
Cited alongside, same era.
The complexity of decentralized control of markov decision processes
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein · 2002
Cited alongside, same era.
Predictive representations of state
M. L. Littman, R. S. Sutton, and S. P. Singh · 2002
Cited alongside, same era.
Equivalence notions and model minimization in Markov decision processes
R. Givan, T. Dean, and M. Greig · 2003
Cited alongside, same era.
A planning algorithm for predictive state representations
M. T. Izadi and D. Precup · 2003
Cited alongside, same era.
Sequential Decision Making in Decentralized Systems
A. Nayyar · 2011
Later among the works it cites.
Optimal control strategies in delayed sharing information structures
A. Nayyar, A. Mahajan, and D. Teneketzis · 2011
Later among the works it cites.
Closing the gap: Improved bounds on optimal POMDP solutions
P. Poupart, K.-E. Kim, and D. Kim · 2011
Later among the works it cites.
An algorithm for quantization of discrete probability distributions
Y. A. Reznik · 2011
Later among the works it cites.
A bayesian approach for learning and planning in partially observable markov decision processes
S. Ross, J. Pineau, B. Chaib-draa, and P. Kreitmann · 2011
Later among the works it cites.
Networked Markov decision processes with delays
S. Adlakha, S. Lall, and A. Goldsmith · 2012
Later among the works it cites.
Discrete-time Markov control processes: basic optimality criteria
O. Hernández-Lerma and J. B. Lasserre · 2012
Later among the works it cites.
Information structures in optimal decentralized control
A. Mahajan, N. C. Martins, M. C. Rotkowitz, and S. Yüksel · 2012
Later among the works it cites.
On the empirical estimation of integral probability metrics
B. K. Sriperumbudur, K. Fukumizu, A. Gretton, B. Schölkopf, and G. R. G. Lanckriet · 2012
Later among the works it cites.
Optimal decentralized control of coupled subsystems with control sharing
A. Mahajan · 2013
Later among the works it cites.
Decentralized stochastic control with partial history sharing: A common information approach
A. Nayyar, A. Mahajan, and D. Teneketzis · 2013
Later among the works it cites.
Equivalence of distance-based and RKHS-based statistics in hypothesis testing
D. Sejdinovic, B. Sriperumbudur, A. Gretton, and K. Fukumizu · 2013
Later among the works it cites.
A survey of point-based POMDP solvers
G. Shani, J. Pineau, and R. Kaplow · 2013
Later among the works it cites.
Team optimal control of coupled subsystems with mean-field sharing
J. Arabneydi and A. Mahajan · 2014
Later among the works it cites.
Efficient learning and planning with compressed predictive states
W. Hamilton, M. M. Fard, and J. Pineau · 2014
Later among the works it cites.
Reinforcement learning in decentralized stochastic control systems with partial history sharing
J. Arabneydi and A. Mahajan · 2015
Later among the works it cites.
Deep recurrent Q-learning for partially observable MDPs
M. Hausknecht and P. Stone · 2015
Later among the works it cites.
Memory-based control with recurrent neural networks, 2015
N. Heess, J. J. Hunt, T. P. Lillicrap, and D. Silver · 2015
Later among the works it cites.
A concise introduction to decentralized POMDPs , volume 1
F. A. Oliehoek and C. Amato · 2015
Later among the works it cites.
Temporal logic motion planning using POMDPs with parity objectives: Case study paper
M. Svoreňová, M. Chmelík, K. Leahy, H. F. Eniser, K. Chatterjee, I. Černá, and C. Belta · 2015
Later among the works it cites.
Near optimal behavior via approximate state abstraction
D. Abel, D. Hershkowitz, and M. Littman · 2016
Later among the works it cites.
Reinforcement learning of POMDPs using spectral methods
K. Azizzadenesheli, A. Lazaric, and A. Anandkumar · 2016
Later among the works it cites.
Optimally solving dec-pomdps as continuous-state mdps
J. S. Dibangoye, C. Amato, O. Buffet, and F. Charpillet · 2016
Later among the works it cites.
Improving predictive state representations via gradient descent
N. Jiang, A. Kulesza, and S. P. Singh · 2016
Later among the works it cites.
Learning for decentralized control of multiagent systems in large, partially-observable stochastic environments
M. Liu, C. Amato, E. P. Anesta, J. D. Griffith, and J. P. How · 2016
Later among the works it cites.
Decentralized stochastic control
A. Mahajan and M. Mannan · 2016
Later among the works it cites.
Probability metrics
V. M. Zolotarev · 2016
Later among the works it cites.
POMDPs.jl: A framework for sequential decision making under uncertainty
M. Egorov, Z. N. Sunberg, E. Balaban, T. A. Wheeler, J. K. Gupta, and M. J. Kochenderfer · 2017
Later among the works it cites.
On improving deep reinforcement learning for POMDPs, 2017
P. Zhu, X. Li, P. Poupart, and G. Miao · 2017
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
K. Asadi, D. Misra, and M. Littman · 2018
Later among the works it cites.
Learning internal state models in partially observable environments;
A. Baisero and C. Amato · 2018
Later among the works it cites.
Exploration by random network distillation
Y. Burda, H. Edwards, A. Storkey, and O. Klimov · 2018
Later among the works it cites.
Sufficient conditions for the value function and optimal strategy to be even and quasi-convex
J. Chakravorty and A. Mahajan · 2018
Later among the works it cites.
D. Ha and J. Schmidhuber · 2018
Later among the works it cites.
Deep variational reinforcement learning for POMDPs
M. Igl, L. Zintgraf, T. A. Le, F. Wood, and S. Whiteson · 2018
Later among the works it cites.
Finite approximations in discrete-time stochastic control
N. Saldi, T. Linder, and S. Yüksel · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
On overfitting and asymptotic bias in batch reinforcement learning with partial observability
V. Francois-Lavet, G. Rabusseau, J. Pineau, D. Ernst, and R. Fonteneau · 2019
Later among the works it cites.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare · 2019
Later among the works it cites.
Bayesian reinforcement learning in factored POMDPs
S. Katt, F. A. Oliehoek, and C. Amato · 2019
Later among the works it cites.
On data-processing and majorization inequalities for f-divergences with applications
I. Sason · 2019
Later among the works it cites.
Approximate information state for partially observed system
J. Subramanian and A. Mahajan · 2019
Later among the works it cites.
Lifelong learning with a changing action set
Y. Chandak, G. Theocharous, C. Nota, and P. S. Thomas · 2020
Closest in time.
Approximate information state for reinforcement learning in partially observed systems
J. Subramanian, A. Sinha, R. Seraj, and A. Mahajan · 2020
Closest in time.