Fetching the paper…
Reading the bibliography…
Off-policy evaluation (OPE) in contextual bandits has seen rapid adoption in real-world systems, since it enables offline evaluation of new policies using only historic log data.
A generalization of sampling without replacement from a finite universe
Horvitz, D. G. and Thompson, D. J · 1952
Earlier work this paper cites.
The continuum-armed bandit problem
Agrawal, R · 1995
Earlier work this paper cites.
Optimal pointwise adaptive methods in nonparametric estimation
Lepski, O. V. and Spokoiny, V. G · 1997
Earlier work this paper cites.
Random forests
Breiman, L · 2001
Earlier work this paper cites.
Nearly tight bounds for the continuum-armed bandit problem
Kleinberg, R · 2004
Earlier work this paper cites.
Efficient multiple-click models in web search
Guo, F., Liu, C., and Wang, Y. M · 2009
Earlier work this paper cites.
Using continuous action spaces to solve discrete problems
Van Hasselt, H. and Wiering, M. A · 2009
Earlier work this paper cites.
X-armed bandits
Bubeck, S., Munos, R., Stoltz, G., and Szepesvári, C · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Chu, W., Li, L., Reyzin, L., and Schapire, R · 2011
Earlier work this paper cites.
Generalized value functions for large action sets
Pazis, J. and Parr, R · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Édouard Duchesnay · 2011
Earlier work this paper cites.
Fast reinforcement learning with large action sets using error-correcting output codes for mdp factorization
Dulac-Arnold, G., Denoyer, L., Preux, P., and Gallinari, P · 2012
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S. and Goyal, N · 2013
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
Dudík, M., Erhan, D., Langford, J., and Li, L · 2014
Earlier work this paper cites.
Click models for web search
Chuklin, A., Markov, I., and Rijke, M. d · 2015
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., and Coppin, B · 2015
Earlier work this paper cites.
High confidence policy improvement
Thomas, P., Theocharous, G., and Ghavamzadeh, M · 2015
Earlier work this paper cites.
A neural click model for web search
Borisov, A., Markov, I., De Rijke, M., and Serdyukov, P · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E · 2016
Cited alongside, same era.
Off-policy evaluation for slate recommendation
Swaminathan, A., Krishnamurthy, A., Agarwal, A., Dudik, M., Langford, J., Jose, D., and Zitouni, I · 2017
Cited alongside, same era.
Optimal and adaptive off-policy evaluation in contextual bandits
Wang, Y.-X., Agarwal, A., and Dudık, M · 2017
Cited alongside, same era.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., and Wang, Y.-X · 2017
Cited alongside, same era.
More robust doubly robust off-policy evaluation
Farajtabar, M., Chow, Y., and Ghavamzadeh, M · 2018
Cited alongside, same era.
Policy evaluation and optimization with continuous treatments
Kallus, N. and Zhou, A · 2018
Empirical study of off-policy policy evaluation for reinforcement learning
Voloshin, C., Le, H. M., Jiang, N., and Yue, Y · 2019
Later among the works it cites.
Combining experimental and observational data to estimate treatment effects on long term outcomes
Athey, S., Chetty, R., and Imbens, G · 2020
Later among the works it cites.
Debiasing grid-based product search in e-commerce
Guo, R., Zhao, X., Henderson, A., Hong, L., and Liu, H · 2020
Later among the works it cites.
On the role of surrogates in the efficient estimation of treatment effects with limited outcome data
Kallus, N. and Mao, X · 2020
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N. and Uehara, M · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Offline evaluation of ranking policies with click models
Li, S., Abbasi-Yadkori, Y., Kveton, B., Muthukrishnan, S., Vinay, V., and Wen, Z · 2018
Cited alongside, same era.
Breaking the curse of horizon: infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D · 2018
Cited alongside, same era.
Cross-fitting and fast remainder rates for semiparametric estimation
Newey, W. K. and Robins, J. R · 2018
Cited alongside, same era.
The surrogate index: Combining short-term proxies to estimate long-term treatment effects more rapidly and precisely
Athey, S., Chetty, R., Imbens, G. W., and Kang, H · 2019
Cited alongside, same era.
Learning action representations for reinforcement learning
Chandak, Y., Theocharous, G., Kostas, J., Jordan, S., and Thomas, P · 2019
Cited alongside, same era.
Semi-parametric efficient policy learning with continuous actions
Demirer, M., Syrgkanis, V., Lewis, G., and Chernozhukov, V · 2019
Cited alongside, same era.
Later among the works it cites.
Counterfactual evaluation of slate recommendations with sequential reward interactions
McInerney, J., Brost, B., Chandar, P., Mehrotra, R., and Carterette, B · 2020
Later among the works it cites.
Off-policy bandits with deficient support
Sachdeva, N., Su, Y., and Joachims, T · 2020
Later among the works it cites.
Doubly robust estimator for ranking metrics with post-click conversions
Saito, Y · 2020
Later among the works it cites.
Open bandit dataset and pipeline: Towards realistic and reproducible off-policy evaluation
Saito, Y., Aihara, S., Matsutani, M., and Narita, Y · 2020
Later among the works it cites.
Balanced off-policy evaluation in general action spaces
Sondhi, A., Arbour, D., and Dimmery, D · 2020
Later among the works it cites.
Semiparametric estimation of long-term treatment effects
Chen, J. and Ritzwoller, D. M · 2021
Later among the works it cites.
Optimal off-policy evaluation from multiple logging policies
Kallus, N., Saito, Y., and Uehara, M · 2021
Later among the works it cites.
Learning from extreme bandit feedback
Lopez, R., Dhillon, I. S., and Jordan, M. I · 2021
Later among the works it cites.
Subgaussian and differentiable importance sampling for off-policy evaluation and learning
Metelli, A. M., Russo, A., and Restelli, M · 2021
Later among the works it cites.
Evaluating the robustness of off-policy evaluation
Saito, Y., Udagawa, T., Kiyohara, H., Mogi, K., Narita, Y., and Tateno, K · 2021
Later among the works it cites.
Improved estimator selection for off-policy evaluation
Tucker, G. and Lee, J · 2021
Later among the works it cites.
Control variates for slate off-policy evaluation
Vlassis, N., Chandrashekar, A., Gil, F. A., and Kallus, N · 2021
Later among the works it cites.
Doubly robust off-policy evaluation for ranking policies under the cascade behavior model
Kiyohara, H., Saito, Y., Matsuhiro, T., Narita, Y., Shimizu, N., and Yamamoto, Y · 2022
Closest in time.