Fetching the paper…
Reading the bibliography…
Solving Partially Observable Markov Decision Processes (POMDPs) is hard.
Yasin Abbasi-Yadkori, Nevena Lazic, Csaba Szepesvari, and Gellert Weisz · 1908
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Logarithmic regret bound in partially observable linear dynamical systems
Sahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, and Anima Anandkumar · 2003
Earlier work this paper cites.
Model-based online learning of pomdps
Guy Shani, Ronen I Brafman, and Solomon E Shimony · 2005
Earlier work this paper cites.
Explore more and improve regret in linear quadratic regulators
Sahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, and Anima Anandkumar · 2007
Earlier work this paper cites.
Bayes-adaptive pomdps
Stephane Ross, Brahim Chaib-draa, and Joelle Pineau · 2007
Earlier work this paper cites.
Model-based bayesian reinforcement learning in partially observable domains
Pascal Poupart and Nikos Vlassis · 2008
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Peter L Bartlett and Ambuj Tewari · 2009
Earlier work this paper cites.
Learning to explore and exploit in pomdps
Chenghui Cai, Xuejun Liao, and Lawrence Carin · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
The infinite regionalized policy representation
Miao Liu, Xuejun Liao, and Lawrence Carin · 2011
Earlier work this paper cites.
Modelling transition dynamics in mdps with rkhs embeddings
Steffen Grunewalder, Guy Lever, Luca Baldassarre, Massi Pontil, and Arthur Gretton · 2012
Earlier work this paper cites.
Bayesian nonparametric methods for partially-observable reinforcement learning
Finale Doshi-Velez, David Pfau, Frank Wood, and Nicholas Roy · 2013
Earlier work this paper cites.
An optimistic posterior sampling strategy for bayesian reinforcement learning
Raphaël Fonteneau, Nathan Korda, and Rémi Munos · 2013
Earlier work this paper cites.
Online expectation maximization for reinforcement learning in pomdps
Miao Liu, Xuejun Liao, and Lawrence Carin · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Cited alongside, same era.
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Thompson sampling for learning parameterized markov decision processes
Aditya Gopalan and Shie Mannor · 2015
Cited alongside, same era.
Stochastic systems: Estimation, identification, and adaptive control
Panqanamala Ramana Kumar and Pravin Varaiya · 2015
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Later among the works it cites.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zihan Zhang and Xiangyang Ji · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Later among the works it cites.
Root-n-regret for learning in markov decision processes with function approximation and low bellman rank
Kefan Dong, Jian Peng, Yining Wang, and Yuan Zhou · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Experimental results: Reinforcement learning of pomdps using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar · 2017
Cited alongside, same era.
Dynamic programming and optimal control, vol i and ii, 4th edition
Dimitri P Bertsekas · 2017
Cited alongside, same era.
Thompson sampling for stochastic control: The finite parameter case
Michael Jong Kim · 2017
Cited alongside, same era.
Policy gradient in partially observable environments: Approximation and convergence
Kamyar Azizzadenesheli, Yisong Yue, and Animashree Anandkumar · 2018
Cited alongside, same era.
Regret bounds for robust adaptive control of the linear quadratic regulator
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu · 2018
Cited alongside, same era.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Ronald Ortner · 2018
Cited alongside, same era.
Botao Hao, Nevena Lazic, Yasin Abbasi-Yadkori, Pooria Joulani, and Csaba Szepesvari · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
Naive exploration is optimal for online lqr
Max Simchowitz and Dylan Foster · 2020
Later among the works it cites.
Jayakumar Subramanian, Amit Sinha, Raihan Seraj, and Aditya Mahajan · 2020
Later among the works it cites.
Online learning of the kalman filter with logarithmic regret
Anastasios Tsiamis and George Pappas · 2020
Later among the works it cites.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Ruosong Wang, Russ R Salakhutdinov, and Lin Yang · 2020
Later among the works it cites.
Model-free reinforcement learning in infinite-horizon average-reward markov decision processes
Chen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, Hiteshi Sharma, and Rahul Jain · 2020
Later among the works it cites.
Online learning for stochastic shortest path model via posterior sampling
Mehdi Jafarnia-Jahromi, Liyu Chen, Rahul Jain, and Haipeng Luo · 2021
Closest in time.
Tianhao Wang, Dongruo Zhou, and Quanquan Gu · 2021
Closest in time.
Learning infinite-horizon average-reward mdps with linear function approximation
Chen-Yu Wei, Mehdi Jafarnia-Jahromi, Haipeng Luo, and Rahul Jain · 2021
Closest in time.