Fetching the paper…
Reading the bibliography…
This paper provides a statistical analysis of high-dimensional batch Reinforcement Learning (RL) using sparse linear function approximation.
Yasin Abbasi-Yadkori, Nevena Lazic, Csaba Szepesvari, and Gellert Weisz · 1908
Earlier work this paper cites.
Polynomial approximation – a new computational technique in dynamic programming
I. R. Bellman, R. Kalaba, and B. Kotkin · 1963
Earlier work this paper cites.
Generalized polynomial approximations in Markovian decision processes
Paul J Schweitzer and Abraham Seidmann · 1985
Earlier work this paper cites.
Efficient memory-based learning for robot control
Andrew William Moore · 1990
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri P. Bertsekas · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the Lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S Sutton, and Satinder Singh · 2000
Earlier work this paper cites.
Atomic decomposition by basis pursuit
Scott Shaobing Chen, David L Donoho, and Michael A Saunders · 2001
Earlier work this paper cites.
GradientDICE: Rethinking generalized offline estimation of stationary values
Shangtong Zhang, Bo Liu, and Shimon Whiteson · 2001
Earlier work this paper cites.
GenDICE: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Lihong Li, and Dale Schuurmans · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
Ming Yuan and Yi Lin · 2006
Earlier work this paper cites.
On model selection consistency of lasso
Peng Zhao and Bin Yu · 2006
Earlier work this paper cites.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2007
Earlier work this paper cites.
Sparsity oracle inequalities for the Lasso
F. Bunea, A. Tsybakov, and M. Wegkamp · 2007
Earlier work this paper cites.
Fitted Q-iteration in continuous action-space MDPs
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Regularized fitted Q-iteration: Application to planning
Amir massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Simultaneous analysis of Lasso and Dantzig selector
Peter J Bickel, Ya’acov Ritov, Alexandre B Tsybakov, et al · 2009
Cited alongside, same era.
Regularization and feature selection in least-squares temporal difference learning
J Zico Kolter and Andrew Y Ng · 2009
Cited alongside, same era.
Sharp thresholds for high-dimensional and noisy sparsity recovery using l1-constrained quadratic programming (lasso)
Martin J Wainwright · 2009
Cited alongside, same era.
Algorithms for Reinforcement Learning
Csaba Szepesvári · 2010
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Cited alongside, same era.
Statistics for high-dimensional data: methods, theory and applications
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Later among the works it cites.
Consistent on-line off-policy evaluation
Assaf Hallak and Shie Mannor · 2017
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peter Bühlmann and Sara Van De Geer · 2011
Cited alongside, same era.
ℓ 1 \ell^{1} -penalized projected Bellman residual
Matthieu Geist and Bruno Scherrer · 2011
Cited alongside, same era.
Finite-sample analysis of Lasso-TD
Mohammad Ghavamzadeh, Alessandro Lazaric, Rémi Munos, and Matthew Hoffman · 2011
Cited alongside, same era.
Regularized least squares temporal difference learning with nested
Matthew W Hoffman, Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2011
Cited alongside, same era.
A Dantzig selector approach to temporal difference learning
Matthieu Geist, Bruno Scherrer, Alessandro Lazaric, and Mohammad Ghavamzadeh · 2012
Cited alongside, same era.
Efficient reinforcement learning for high dimensional linear quadratic systems
Morteza Ibrahimi, Adel Javanmard, and Benjamin V Roy · 2012
Cited alongside, same era.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Cited alongside, same era.
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2019
Later among the works it cites.
Batch policy learning under constraints
Hoang Le, Cameron Voloshin, and Yisong Yue · 2019
Later among the works it cites.
DualDICE: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Later among the works it cites.
Minimax weight and Q-function learning for off-policy evaluation
Masatoshi Uehara and Nan Jiang · 2019
Later among the works it cites.
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
Martin J Wainwright · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Later among the works it cites.
Minimax-optimal off-policy evaluation with linear function approximation
Yaqi Duan and Mengdi Wang · 2020
Closest in time.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Nathan Kallus and Masatoshi Uehara · 2020
Closest in time.
A maximum-entropy approach to off-policy evaluation in average-reward mdps
Nevena Lazic, Dong Yin, Mehrdad Farajtabar, Nir Levine, Dilan Gorur, Chris Harris, and Dale Schuurmans · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Closest in time.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Lin F Yang and Mengdi Wang · 2020
Closest in time.
Off-policy evaluation via the regularized lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Closest in time.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Ming Yin and Yu-Xiang Wang · 2020
Closest in time.
Frequentist regret bounds for randomized least-squares value iteration
Andrea Zanette, David Brandfonbrener, Emma Brunskill, Matteo Pirotta, and Alessandro Lazaric · 2020
Closest in time.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2020
Closest in time.