Fetching the paper…
Reading the bibliography…
Provably sample-efficient Reinforcement Learning (RL) with rich observations and function approximation has witnessed tremendous recent progress, particularly when the underlying function approximators are linear.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
The equivalence of two extremum problems
Jack Kiefer and Jacob Wolfowitz · 1960
Earlier work this paper cites.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 1964
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R.S. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Nicolo Cesa-Bianchi, Alex Conconi, and Claudio Gentile · 2004
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Vivek S Borkar · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Earlier work this paper cites.
Combinatorial bandits
Nicolo Cesa-Bianchi and Gábor Lugosi · 2012
Earlier work this paper cites.
Theory Of Optimal Experiments
V.V. Fedorov · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Online learning in mdps with side information
Yasin Abbasi-Yadkori and Gergely Neu · 2014
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M.L. Puterman · 2014
Earlier work this paper cites.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Earlier work this paper cites.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2016
Earlier work this paper cites.
Posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel V Todorov, and Sham M Kakade · 2017
Cited alongside, same era.
A tutorial on thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2017
Cited alongside, same era.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni · 2017
Cited alongside, same era.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song · 2018
Cited alongside, same era.
Active learning for nonlinear system identification with guarantees
Horia Mania, Michael I Jordan, and Benjamin Recht · 2020
Later among the works it cites.
Learning the linear quadratic regulator from nonlinear observations
Zakaria Mhammedi, Dylan J Foster, Max Simchowitz, Dipendra Misra, Wen Sun, Akshay Krishnamurthy, Alexander Rakhlin, and John Langford · 2020
Later among the works it cites.
Naive exploration is optimal for online lqr
Max Simchowitz and Dylan Foster · 2020
Later among the works it cites.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Ruosong Wang, Russ R Salakhutdinov, and Lin Yang · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Lin Yang and Mengdi Wang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christoph Dann · 2018
Cited alongside, same era.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
Markov decision processes with continuous side information
Aditya Modi, Nan Jiang, Satinder Singh, and Ambuj Tewari · 2018
Cited alongside, same era.
Online control with adversarial disturbances
Naman Agarwal, Brian Bullins, Elad Hazan, Sham Kakade, and Karan Singh · 2019
Cited alongside, same era.
Provably efficient rl with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Cited alongside, same era.
Eugene Ie, Vihan Jain, Jing Wang, Sanmit Narvekar, Ritesh Agarwal, Rui Wu, Heng-Tze Cheng, Morgane Lustman, Vince Gatto, Paul Covington, et al · 2019
Cited alongside, same era.
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Cited alongside, same era.
The Finite Element Method: Its Basis and Fundamentals
O.C. Zienkiewicz, R.L. Taylor, and J.Z. Zhu · 2020
Later among the works it cites.
Deep radial-basis value functions for continuous control
Kavosh Asadi, Neev Parikh, Ronald E Parr, George D Konidaris, and Michael L Littman · 2021
Later among the works it cites.
A provably efficient model-free posterior sampling method for episodic reinforcement learning
Christoph Dann, Mehryar Mohri, Tong Zhang, and Julian Zimmert · 2021
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Simon S Du, Sham M Kakade, Jason D Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Later among the works it cites.
Provably correct optimization and exploration with non-linear policies
Fei Feng, Wotao Yin, Alekh Agarwal, and Lin Yang · 2021
Later among the works it cites.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Later among the works it cites.
Online sparse reinforcement learning
Botao Hao, Tor Lattimore, Csaba Szepesvári, and Mengdi Wang · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Later among the works it cites.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Later among the works it cites.
Agnostic reinforcement learning with low-rank mdps and rich observations
Ayush Sekhari, Christoph Dann, Mehryar Mohri, Yishay Mansour, and Karthik Sridharan · 2021
Later among the works it cites.
Masatoshi Uehara, Masaaki Imaizumi, Nan Jiang, Nathan Kallus, Wen Sun, and Tengyang Xie · 2021
Later among the works it cites.
Feel-good thompson sampling for contextual bandits and reinforcement learning
Tong Zhang · 2021
Later among the works it cites.
Efficient reinforcement learning in block mdps: A model-free representation learning approach
Xuezhou Zhang, Yuda Song, Masatoshi Uehara, Mengdi Wang, Wen Sun, and Alekh Agarwal · 2022
Closest in time.