Fetching the paper…
Reading the bibliography…
We offer an experimental benchmark and empirical study for off-policy policy evaluation (OPE) in reinforcement learning, which is a key problem in many safety critical applications.
Adjustment of an inverse matrix corresponding to a change in one element of a given matrix
Jack Sherman and Winifred J. Morrison · 1950
Earlier work this paper cites.
A generalization of sampling without replacement from a finite universe
Daniel G Horvitz and Donovan J Thompson · 1952
Earlier work this paper cites.
Monte carlo methods
John Michael Hammersley and David Christopher Handscomb · 1964
Earlier work this paper cites.
Weighted uniform sampling—a monte carlo technique for reducing variance
Michael JD Powell and J Swann · 1966
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
Multi-agent reinforcement learning for traffic light control
MA Wiering · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Susan A Murphy, Mark J van der Laan, James M Robins, and Conduct Problems Prevention Research Group · 2001
Earlier work this paper cites.
Doubly robust estimation in missing data and causal inference models
Heejung Bang and James M. Robins · 2005
Earlier work this paper cites.
An empirical comparison of supervised learning algorithms
Rich Caruana and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data
Joseph DY Kang, Joseph L Schafer, et al · 2007
Earlier work this paper cites.
An empirical evaluation of supervised learning in high dimensions
Rich Caruana, Nikos Karampatziakis, and Ainur Yessenalina · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Olivier Chapelle and Lihong Li · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Earlier work this paper cites.
Off-policy actor-critic
Thomas Degris, Martha White, and Richard S Sutton · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Qui nonero Candela, Denis X. Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.
Off-policy evaluation in Markov decision processes
Cosmin Paduraru · 2013
Earlier work this paper cites.
Policy evaluation with temporal differences: A survey and comparison
Christoph Dann, Gerhard Neumann, Jan Peters, et al · 2014
Cited alongside, same era.
Toward minimax off-policy value estimation
Lihong Li, Rémi Munos, and Csaba Szepesvári · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Q(lambda) with off-policy corrections
Anna Harutyunyan, Marc G. Bellemare, Tom Stepleton, and Rémi Munos · 2016
Batch policy learning under constraints
Hoang M Le, Cameron Voloshin, and Yisong Yue · 2019
Closest in time.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Closest in time.
Challenging common assumptions in the unsupervised learning of disentangled representations
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem · 2019
Closest in time.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Closest in time.
Learning when-to-treat policies
Xinkun Nie, Emma Brunskill, and Stefan Wager · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Remi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni · 2017
Cited alongside, same era.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
Philip S Thomas, Georgios Theocharous, Mohammad Ghavamzadeh, Ishan Durugkar, and Emma Brunskill · 2017
Cited alongside, same era.
Counterfactual off-policy evaluation with gumbel-max structural causal models
Michael Oberst and David Sontag · 2019
Closest in time.
Off-policy evaluation in partially observable environments
Guy Tennenholtz, Shie Mannor, and Uri Shalit · 2019
Closest in time.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Closest in time.
Coindice: Off-policy confidence interval estimation
Bo Dai, Ofir Nachum, Yinlam Chow, Lihong Li, Csaba Szepesvári, and Dale Schuurmans · 2020
Closest in time.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Closest in time.
Minimax value interval for off-policy evaluation and policy optimization
Nan Jiang and Jiawei Huang · 2020
Closest in time.
Off-policy learning in two-stage recommender systems
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang, Minmin Chen, Jiaxi Tang, Lichan Hong, and Ed H Chi · 2020
Closest in time.
A large-scale open dataset for bandit algorithms
Yuta Saito, Shunsuke Aihara, Megumi Matsutani, and Yusuke Narita · 2020
Closest in time.
Minimax weight and q-function learning for off-policy evaluation
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2020
Closest in time.
Off-policy evaluation via the regularized lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Closest in time.
Gradientdice: Rethinking generalized offline estimation of stationary values
Shangtong Zhang, Bo Liu, and Shimon Whiteson · 2020
Closest in time.
The ai economist: Improving equality and productivity with ai-driven tax policies
Stephan Zheng, Alexander Trott, Sunil Srinivasa, Nikhil Naik, Melvin Gruesbeck, David C Parkes, and Richard Socher · 2020
Closest in time.
Benchmarks for deep off-policy evaluation
Justin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker, ziyu wang, Alexander Novikov, Mengjiao Yang, Michael R Zhang, Yutian Chen, Aviral Kumar, Cosmin Paduraru, Sergey Levine, and Thomas Paine · 2021
Closest in time.
Minimax model learning
Cameron Voloshin, Nan Jiang, and Yisong Yue · 2021
Closest in time.
Autoregressive dynamics models for offline policy evaluation and optimization
Michael R Zhang, Tom Le Paine, Ofir Nachum, Cosmin Paduraru, George Tucker, Ziyu Wang, and Mohammad Norouzi · 2021
Closest in time.