Fetching the paper…
Reading the bibliography…
We propose a new reinforcement learning algorithm for partially observable Markov decision processes (POMDP) based on spectral decomposition methods.
The optimal control of partially observable Markov processes
E. J. Sondik · 1971
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
Arthur P Dempster, Nan M Laird, and Donald B Rubin · 1977
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
A.G. Barto, R.S. Sutton, and C.W. Anderson · 1983
Earlier work this paper cites.
The complexity of markov decision processes
Christos Papadimitriou and John N. Tsitsiklis · 1987
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Memoryless policies: Theoretical limitations and practical results
Michael L. Littman · 1994
Earlier work this paper cites.
Learning without state-estimation in partially observable markovian decision processes
Satinder P Singh, Tommi Jaakkola, and Michael I Jordan · 1994
Earlier work this paper cites.
Reinforcement learning algorithm for partially observable markov decision problems
Tommi Jaakkola, Satinder P. Singh, and Michael I. Jordan · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
D. Bertsekas and J. Tsitsiklis · 1996
Earlier work this paper cites.
Using eligibility traces to find the best memoryless policy in partially observable markov decision processes
John Loch and Satinder P Singh · 1998
Earlier work this paper cites.
On the computability of infinite-horizon partially observable markov decision processes
Omid Madani · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
John K. Williams and Satinder P. Singh · 1998
Earlier work this paper cites.
Planning treatment of ischemic heart disease with partially observable markov decision processes
Milos Hauskrecht and Hamish Fraser · 2000
Earlier work this paper cites.
Pegasus: A policy search method for large mdps and pomdps
Andrew Y. Ng and Michael Jordan · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L. Bartlett · 2001
Earlier work this paper cites.
Predictive representations of state
Michael L. Littman, Richard S. Sutton, and Satinder Singh · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolò Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Reinforcement learning for POMDPs based on action values and stochastic optimization
Theodore J. Perkins · 2002
Cited alongside, same era.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2003
Cited alongside, same era.
Bounded finite state controllers
Pascal Poupart and Craig Boutilier · 2003
Cited alongside, same era.
Policy search by dynamic programming
J. A. Bagnell, Sham M Kakade, Jeff G. Schneider, and Andrew Y. Ng · 2004
Cited alongside, same era.
Efficient planning and tracking in pomdps with large observation spaces
A. Atrash and J. Pineau · 2006
Cited alongside, same era.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Cited alongside, same era.
A method of moments for mixture models and hidden markov models
Animashree Anandkumar, Daniel Hsu, and Sham M Kakade · 2012
Later among the works it cites.
Building adaptive dialogue systems via bayes-adaptive pomdps
Shaowei Png, J. Pineau, and B. Chaib-draa · 2012
Later among the works it cites.
Partially observable markov decision processes
Matthijs T.J. Spaan · 2012
Later among the works it cites.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Later among the works it cites.
Combinatorial multi-armed bandit: General framework and applications
Wei Chen, Yajun Wang, and Yang Yuan · 2013
Later among the works it cites.
Regret bounds for reinforcement learning with policy advice
M. Gheshlaghi-Azar, A. Lazaric, and E. Brunskill · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Planning algorithms
Steven M LaValle · 2006
Cited alongside, same era.
Logarithmic online regret bounds for undiscounted reinforcement learning
P Ortner and R Auer · 2007
Cited alongside, same era.
Bayes-adaptive pomdps
Stephane Ross, Brahim Chaib-draa, and Joelle Pineau · 2007
Cited alongside, same era.
Concentration inequalities for dependent random variables via the martingale method
Leonid Aryeh Kontorovich, Kavita Ramanan, et al · 2008
Cited alongside, same era.
Model-based bayesian reinforcement learning in partially observable domains
P. Poupart and N. Vlassis · 2008
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2009
Cited alongside, same era.
Later among the works it cites.
Sequential transfer in multi-armed bandit with finite set of models
Mohammad Gheshlaghi azar, Alessandro Lazaric, and Emma Brunskill · 2013
Later among the works it cites.
On learning parametric-output hmms
Aryeh Kontorovich, Boaz Nadler, and Roi Weiss · 2013
Later among the works it cites.
Nonparametric estimation of multi-view latent variable models
Le Song, Animashree Anandkumar, Bo Dai, and Bo Xie · 2013
Later among the works it cites.
Tensor decompositions for learning latent variable models
Animashree Anandkumar, Rong Ge, Daniel Hsu, Sham M Kakade, and Matus Telgarsky · 2014
Later among the works it cites.
Resource-efficient stochastic optimization of a locally smooth function under correlated bandit feedback
M. Gheshlaghi-Azar, A. Lazaric, and E. Brunskill · 2014
Later among the works it cites.
Efficient learning and planning with compressed predictive states
William Hamilton, Mahdi Milani Fard, and Joelle Pineau · 2014
Later among the works it cites.
Uniform chernoff and dvoretzky-kiefer-wolfowitz-type inequalities for markov chains and related processes
Aryeh Kontorovich, Roi Weiss, et al · 2014
Later among the works it cites.
Selecting near-optimal approximate state representations in reinforcement learning
Ronald Ortner, Odalric-Ambrym Maillard, and Daniil Ryabko · 2014
Later among the works it cites.
Mixing time estimation in reversible markov chains from a single sample path
Daniel J Hsu, Aryeh Kontorovich, and Csaba Szepesvári · 2015
Later among the works it cites.
Online learning and optimization of markov jump affine models
Sevi Baltaoglu, Lang Tong, and Qing Zhao · 2016
Closest in time.
A pac rl algorithm for episodic pomdps
Zhaohan Daniel Guo, Shayan Doroudi, and Emma Brunskill · 2016
Closest in time.
Contextual-mdps for pac-reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Closest in time.