Fetching the paper…
Reading the bibliography…
Applications of Reinforcement Learning (RL), in which agents learn to make a sequence of decisions despite lacking complete information about the latent states of the controlled system, that is, they act under partial observability of the states, are ubiquitous.
The large-sample distribution of the likelihood ratio for testing composite hypotheses
Samuel S Wilks · 1938
Earlier work this paper cites.
A new family of optimal adaptive controllers for Markov chains
P Kumar and A Becker · 1982
Earlier work this paper cites.
The complexity of Markov decision processes
Christos H Papadimitriou and John N Tsitsiklis · 1987
Earlier work this paper cites.
Adaptive treatment allocation and the multi-armed bandit problem
Tze Leung Lai · 1987
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 1995
Earlier work this paper cites.
Acting under uncertainty: Discrete Bayesian models for mobile-robot navigation
Anthony R Cassandra, Leslie Pack Kaelbling, and James A Kurien · 1996
Earlier work this paper cites.
Planning treatment of ischemic heart disease with partially observable Markov decision processes
Milos Hauskrecht and Hamish Fraser · 2000
Earlier work this paper cites.
Complexity of finite-horizon Markov decision process problems
Martin Mundhenk, Judy Goldsmith, Christopher Lusena, and Eric Allender · 2000
Earlier work this paper cites.
Observable operator models for discrete stochastic time series
Herbert Jaeger · 2000
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Sara A Geer, Sara van de Geer, and D Williams · 2000
Earlier work this paper cites.
On the generalization ability of online learning algorithms
Nicolo Cesa-Bianchi, Alex Conconi, and Claudio Gentile · 2004
Earlier work this paper cites.
From resource allocation to strategy
Joseph L Bower and Clark G Gilbert · 2005
Earlier work this paper cites.
Learning nonsingular phylogenies and hidden Markov models
Elchanan Mossel and Sébastien Roch · 2005
Earlier work this paper cites.
Reinforcement learning in POMDPs without resets
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2005
Earlier work this paper cites.
From ε \varepsilon -entropy to KL-entropy: Analysis of minimum information complexity density estimation
Tong Zhang · 2006
Earlier work this paper cites.
Bayes-adaptive POMDPs
Stephane Ross, Brahim Chaib-draa, and Joelle Pineau · 2007
Earlier work this paper cites.
Model-based Bayesian reinforcement learning in partially observable domains
Pascal Poupart and Nikos Vlassis · 2008
Earlier work this paper cites.
Inventory control with product returns: The impact of imperfect information
Marisa P De Brito and Erwin A Van Der Laan · 2009
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Cited alongside, same era.
Towards fully autonomous driving: Systems and algorithms
Jesse Levinson, Jake Askeland, Jan Becker, Jennifer Dolson, David Held, Soeren Kammel, J Zico Kolter, Dirk Langer, Oliver Pink, Vaughan Pratt, et al · 2011
Cited alongside, same era.
On the computational complexity of stochastic controller optimization in POMDPs
Nikos Vlassis, Michael L Littman, and David Barber · 2012
Cited alongside, same era.
A spectral algorithm for learning hidden Markov models
Daniel Hsu, Sham M Kakade, and Tong Zhang · 2012
Cited alongside, same era.
Eluder dimension and the sample complexity of optimistic exploration
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
Provably efficient RL with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Learning near optimal policies with low inherent Bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Russo and Benjamin Van Roy · 2013
Cited alongside, same era.
Tensor decompositions for learning latent variable models
Animashree Anandkumar, Rong Ge, Daniel Hsu, Sham M Kakade, and Matus Telgarsky · 2014
Cited alongside, same era.
Reinforcement learning of POMDPs using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar · 2016
Cited alongside, same era.
A PAC RL algorithm for episodic POMDPs
Zhaohan Daniel Guo, Shayan Doroudi, and Emma Brunskill · 2016
Cited alongside, same era.
PAC reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Later among the works it cites.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Later among the works it cites.
Sublinear regret for learning POMDPs
Yi Xiong, Ningyuan Chen, Xuefeng Gao, and Xiang Zhou · 2021
Later among the works it cites.
Online learning for unknown partially observable mdps
Mehdi Jafarnia-Jahromi, Rahul Jain, and Ashutosh Nayyar · 2021
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in RL
Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of RL problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Reward biased maximum likelihood estimation for reinforcement learning
Akshay Mete, Rahul Singh, Xi Liu, and PR Kumar · 2021
Later among the works it cites.
Representation learning for online and offline RL in low-rank MDPs
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Later among the works it cites.
Planning in observable POMDPs in quasipolynomial time
Noah Golowich, Ankur Moitra, and Dhruv Rohatgi · 2022
Closest in time.
Provable reinforcement learning with a short-term memory
Yonathan Efroni, Chi Jin, Akshay Krishnamurthy, and Sobhan Miryoosefi · 2022
Closest in time.