Bootstrapping with models: Confidence intervals for off-policy evaluation
Josiah P. Hanna, Peter Stone, and Scott Niekum · 2017
Later among the works it cites.
A linearly relaxed approximate linear program for Markov decision processes
Original
Chandrashekar Lakshminarayanan, Shalabh Bhatnagar, and Csaba Szepesvari · 2017
Later among the works it cites.
The empirical likelihood approach to quantifying uncertainty in sample average approximation
Henry Lam and Enlu Zhou · 2017
Later among the works it cites.
Variance-based regularization with convex objectives
Hongseok Namkoong and John C Duchi · 2017
Later among the works it cites.
Monte Carlo confidence sets for identified sets
X. Chen, T. M. Christensen, and E. Tamer · 2018
Later among the works it cites.
Evaluating reinforcement learning algorithms in observational health settings, 2018
Original
Omer Gottesman, Fredrik Johansson, Joshua Meier, Jack Dent, Donghun Lee, Srivatsan Srinivasan, Linying Zhang, Yi Ding, David Wihl, Xuefeng Peng, Jiayu Yao, Isaac Lage, Christopher Mosch, Li wei H. Lehman, Matthieu Komorowski, Matthieu Komorowski, Aldo Faisal, Leo Anthony Celi, David Sontag, and Finale Doshi-Velez · 2018
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Later among the works it cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Original
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2019
Later among the works it cites.
Top- k k off-policy correction for a REINFORCE recommender system
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed H Chi · 2019
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Later among the works it cites.
Empirical likelihood for contextual bandits
Original
Nikos Karampatziakis, John Langford, and Paul Mineiro · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Remi Tachet Des Combes · 2019
Later among the works it cites.
On gradient descent ascent for nonconvex-concave minimax problems
Original
Tianyi Lin, Chi Jin, and Michael I. Jordan · 2019
Later among the works it cites.
Minimax weight and Q-function learning for off-policy evaluation
Original
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Original
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
Optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Later among the works it cites.
The importance of pessimism in fixed-dataset policy optimization
Original
Jacob Buckman, Carles Gelada, and Marc G Bellemare · 2020
Closest in time.
Pessimism about unknown unknowns inspires conservatism
Michael K Cohen and Marcus Hutter · 2020
Closest in time.
Confident Off-Policy Evaluation and Selection through Self-Normalized Importance Weighting
Original
Ilja Kuzborskij, Claire Vernade, András György, Csaba Szepesvári · 2020
Closest in time.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.
Reinforcement learning via Fenchel-Rockafellar duality
Original
Ofir Nachum and Bo Dai · 2020
Closest in time.
Doubly robust bias reduction in infinite horizon off-policy estimation
Ziyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou, and Qiang Liu · 2020
Closest in time.
Mopo: Model-based offline policy optimization
Original
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Closest in time.