Fetching the paper…
Reading the bibliography…
We propose a new approach to value function approximation which combines linear temporal difference reinforcement learning with subspace identification.
The most predictable criterion
Harold Hotelling · 1935
Earlier work this paper cites.
Optimal control of Markov decision processes with incomplete state estimation
K. J. Aström · 1965
Earlier work this paper cites.
The Optimal Control of Partially Observable Markov Processes
E. J. Sondik · 1971
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
Anthony R. Cassandra, Leslie P. Kaelbling, and Michael R. Littman · 1994
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J. Bradtke and Andrew G. Barto · 1996
Earlier work this paper cites.
Matrix Computations
Gene H. Golub and Charles F. Van Loan · 1996
Earlier work this paper cites.
Subspace Identification for Linear Systems: Theory, Implementation, Applications
P. Van Overschee and B. De Moor · 1996
Earlier work this paper cites.
Optimal stopping of markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
John N. Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Multivariate Reduced-rank Regression: Theory and Applications
Gregory C. Reinsel and Rajabather Palani Velu · 1998
Earlier work this paper cites.
Least-squares temporal difference learning
Justin A. Boyan · 1999
Earlier work this paper cites.
Observable operator models for discrete stochastic time series
Herbert Jaeger · 2000
Earlier work this paper cites.
Causality: models, reasoning, and inference
Judea Pearl · 2000
Cited alongside, same era.
Dynamic data factorization
S. Soatto and A. Chiuso · 2001
Cited alongside, same era.
Value-directed compression of pomdps
Pascal Poupart and Craig Boutilier · 2002
Cited alongside, same era.
Predictive representations of state
Michael Littman, Richard Sutton, and Satinder Singh · 2002
Cited alongside, same era.
Least-squares policy iteration
Michail G. Lagoudakis and Ronald Parr · 2003
Cited alongside, same era.
Predictive state representations: A new theory for modeling dynamical systems
Satinder Singh, Michael James, and Matthew Rudary · 2004
Cited alongside, same era.
Learning low dimensional predictive representations
Learning predictive state representations using non-blind policies
Michael Bowling, Peter McCracken, Michael James, James Neufeld, and Dana Wilkinson · 2006
Later among the works it cites.
Improving approximate value iteration using memories and predictive state representations
Michael R. James, Ton Wessling, and Nikos A. Vlassis · 2006
Later among the works it cites.
A generalized kalman filter for fixed point approximation and efficient temporal-difference learning
David Choi and Benjamin Roy · 2006
Later among the works it cites.
Compact spectral bases for value function approximation using kronecker factorization
Jeff Johns, Sridhar Mahadevan, and Chang Wang · 2007
Later among the works it cites.
A novel orthogonal nmf-based belief compression for pomdps
Xin Li, William K. W. Cheung, Jiming Liu, and Zhili Wu · 2007
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthew Rosencrantz, Geoffrey J. Gordon, and Sebastian Thrun · 2004
Cited alongside, same era.
Representation policy iteration
Sridhar Mahadevan · 2005
Cited alongside, same era.
Samuel meets amarel: automating value function approximation using global state space analysis
Sridhar Mahadevan · 2005
Cited alongside, same era.
Y.: Planning in pomdps using multiplicity automata
Eyal Even-dar · 2005
Cited alongside, same era.
Subspace Methods for System Identification
Tohru Katayama · 2005
Cited alongside, same era.
Ronald Parr, Lihong Li, Gavin Taylor, Christopher Painter-Wakefield, and Michael L. Littman · 2008
Later among the works it cites.
Regularization and feature selection in least-squares temporal difference learning
J. Zico Kolter and Andrew Y. Ng · 2009
Later among the works it cites.
A spectral algorithm for learning hidden Markov models
Daniel Hsu, Sham Kakade, and Tong Zhang · 2009
Later among the works it cites.
Compressing pomdps using locality preserving non-negative matrix factorization, 2010
Georgios Theocharous and Sridhar Mahadevan · 2010
Closest in time.
Closing the learning-planning loop with predictive state representations
Byron Boots, Sajid M. Siddiqi, and Geoffrey J. Gordon · 2010
Closest in time.
Reduced-rank hidden Markov models
Sajid Siddiqi, Byron Boots, and Geoffrey J. Gordon · 2010
Closest in time.