Fetching the paper…
Reading the bibliography…
We study the problem of model selection in batch policy optimization: given a fixed, partial-feedback dataset and $M$ model classes, learn a policy with performance that is competitive with the policy derived from the best model class.
Adaptive model selection using empirical complexities
Gábor Lugosi and Andrew B Nobel · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Data-dependent margin-based generalization bounds for classification
András Antos, Balázs Kégl, Tamás Linder, and Gábor Lugosi · 2002
Earlier work this paper cites.
Model selection and error estimation
Peter L Bartlett, Stéphane Boucheron, and Gábor Lugosi · 2002
Earlier work this paper cites.
Concentration inequalities and model selection
Pascal Massart · 2007
Earlier work this paper cites.
Fast rates for estimation error and oracle inequalities for model selection
Peter L Bartlett · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Model selection in reinforcement learning
Amir-massoud Farahmand and Csaba Szepesvári · 2011
Earlier work this paper cites.
A tail inequality for quadratic forms of subgaussian random vectors
Daniel Hsu, Sham Kakade, Tong Zhang, et al · 2012
Earlier work this paper cites.
Random design analysis of ridge regression
Daniel Hsu, Sham M Kakade, and Tong Zhang · 2012
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Bounded regret in stochastic multi-armed bandits
Sébastien Bubeck, Vianney Perchet, and Philippe Rigollet · 2013
Earlier work this paper cites.
Agnostic notes on regression adjustments to experimental data: Reexamining freedman’s critique
Winston Lin · 2013
Earlier work this paper cites.
Hanson-wright inequality and sub-gaussian concentration
Mark Rudelson and Roman Vershynin · 2013
Earlier work this paper cites.
Causal inference in statistics, social, and biomedical sciences
Guido W Imbens and Donald B Rubin · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Corralling a band of bandit algorithms
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, and Robert E Schapire · 2017
Cited alongside, same era.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Model selection in contextual stochastic bandit problems
Aldo Pacchiano, My Phan, Yasin Abbasi-Yadkori, Anup Rao, Julian Zimmert, Tor Lattimore, and Csaba Szepesvari · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Caglar Gulcehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2020
Later among the works it cites.
Adaptive estimator selection for off-policy evaluation
Yi Su, Pavithra Srinath, and Akshay Krishnamurthy · 2020
Later among the works it cites.
Offline policy selection under uncertainty
Mengjiao Yang, Bo Dai, Ofir Nachum, George Tucker, and Dale Schuurmans · 2020
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lecture notes for statistics 311/electrical engineering 377
John Duchi · 2019
Cited alongside, same era.
Model selection for contextual bandits
Dylan Foster, Akshay Krishnamurthy, and Haipeng Luo · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2019
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Cited alongside, same era.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Cited alongside, same era.
Closest in time.
A workflow for offline model-free robotic reinforcement learning
Aviral Kumar, Anikait Singh, Stephen Tian, Chelsea Finn, and Sergey Levine · 2021
Closest in time.
Online model selection for reinforcement learning with function approximation
Jonathan Lee, Aldo Pacchiano, Vidya Muthukumar, Weihao Kong, and Emma Brunskill · 2021
Closest in time.
Leveraging good representations in linear contextual bandits
Matteo Papini, Andrea Tirinzoni, Marcello Restelli, Alessandro Lazaric, and Matteo Pirotta · 2021
Closest in time.
Model selection for offline reinforcement learning: Practical considerations for healthcare settings
Shengpu Tang and Jenna Wiens · 2021
Closest in time.
Pessimistic model-based offline rl: Pac bounds and posterior sampling under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Closest in time.
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Closest in time.
Batch value-function approximation with only realizability
Tengyang Xie and Nan Jiang · 2021
Closest in time.
On the optimality of batch policy optimization algorithms
Chenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai, Tor Lattimore, Lihong Li, Csaba Szepesvari, and Dale Schuurmans · 2021
Closest in time.
Provably efficient representation learning in low-rank markov decision processes
Weitong Zhang, Jiafan He, Dongruo Zhou, Amy Zhang, and Quanquan Gu · 2021
Closest in time.
Towards hyperparameter-free policy selection for offline reinforcement learning
Siyuan Zhang and Nan Jiang · 2021
Closest in time.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Closest in time.