Fetching the paper…
Reading the bibliography…
This paper studies offline policy learning, which aims at utilizing observations collected a priori (from either fixed or adaptively evolving behavior policies) to learn an optimal individualized decision rule that achieves the best overall outcomes for a given population.
Algaedice: Policy gradient from arbitrary experience
Nachum, O., Dai, B., Kostrikov, I., Chow, Y., Li, L., and Schuurmans, D. (2019) · 1912
Earlier work this paper cites.
Self-normalized processes: exponential inequalities, moment bounds and iterated logarithm laws
De la Pena, V. H., Klass, M. J., and Lai, T. L. (2004) · 1933
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
Sequential medical trials
Anscombe, F. (1963) · 1963
Earlier work this paper cites.
Cross-validatory choice and assessment of statistical predictions
Stone, M. (1974) · 1974
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D. A. (1975) · 1975
Earlier work this paper cites.
Adaptive treatment assignment methods and clinical trials
Simon, R. (1977) · 1977
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L., Robbins, H., et al. (1985) · 1985
Earlier work this paper cites.
On asymptotically efficient estimation in semiparametric models
Schick, A. (1986) · 1986
Earlier work this paper cites.
On learning sets and functions
Natarajan, B. K. (1989) · 1989
Earlier work this paper cites.
Ideal spatial adaptation by wavelet shrinkage
Donoho, D. L. and Johnstone, J. M. (1994) · 1994
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Hoeffding, W. (1994) · 1994
Earlier work this paper cites.
Estimation of regression coefficients when some regressors are not always observed
Robins, J. M., Rotnitzky, A., and Zhao, L. P. (1994) · 1994
Earlier work this paper cites.
Variable selection via nonconcave penalized likelihood and its oracle properties
Fan, J. and Li, R. (2001) · 2001
Earlier work this paper cites.
Gendice: Generalized offline estimation of stationary values
Zhang, R., Dai, B., Li, L., and Schuurmans, D. (2020) · 2002
Earlier work this paper cites.
Optimal dynamic treatment regimes
Murphy, S. A. (2003) · 2003
Earlier work this paper cites.
Statistical treatment rules for heterogeneous populations
Manski, C. F. (2004) · 2004
Earlier work this paper cites.
Local rademacher complexities
Bartlett, P. L., Bousquet, O., and Mendelson, S. (2005) · 2005
Earlier work this paper cites.
An experimental design for the development of adaptive treatment strategies
Murphy, S. A. (2005) · 2005
Earlier work this paper cites.
Convexity, classification, and risk bounds
Bartlett, P. L., Jordan, M. I., and McAuliffe, J. D. (2006) · 2006
Earlier work this paper cites.
Local rademacher complexities and oracle inequalities in risk minimization
Koltchinskii, V. (2006) · 2006
Earlier work this paper cites.
The adaptive lasso and its oracle properties
Zou, H. (2006) · 2006
Earlier work this paper cites.
The multiphase optimization strategy (most) and the sequential multiple assignment randomized trial (smart): new methods for more potent ehealth interventions
Collins, L. M., Murphy, S. A., and Strecher, V. (2007) · 2007
Earlier work this paper cites.
Data-adaptive selection of the truncation level for inverse-probability-of-treatment-weighted estimators
Bembom, O. and van der Laan, M. J. (2008) · 2008
Earlier work this paper cites.
Introduction to nonparametric estimation
Tsybakov, A. B. (2008) · 2008
Earlier work this paper cites.
The importance of pessimism in fixed-dataset policy optimization
Buckman, J., Gelada, C., and Bellemare, M. G. (2020) · 2009
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Hastie, T., Tibshirani, R., Friedman, J. H., and Friedman, J. H. (2009) · 2009
Earlier work this paper cites.
Asymptotics for statistical treatment rules
Hirano, K. and Porter, J. R. (2009) · 2009
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Maurer, A. and Pontil, M. (2009) · 2009
Earlier work this paper cites.
Self-normalized processes: Limit theory and Statistical Applications
Peña, V. H., Lai, T. L., and Shao, Q.-M. (2009) · 2009
Earlier work this paper cites.
Minimax regret treatment choice with finite samples
Stoye, J. (2009) · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Earlier work this paper cites.
Estimating means of bounded random variables by betting
Waudby-Smith, I. and Ramdas, A. (2020) · 2010
Earlier work this paper cites.
Multiclass learnability and the ERM principle
Daniely, A., Sabato, S., Ben-David, S., and Shalev-Shwartz, S. (2011) · 2011
Cited alongside, same era.
The battle trial: Personalizing therapy for lung cancerthe battle trial: Personalizing therapy for lung cancer
Kim, E. S., Herbst, R. S., Wistuba, I. I., Lee, J. J., Blumenschein, G. R., Tsao, A., Stewart, D. J., Hicks, M. E., Erasmus, J., Gupta, S., et al. (2011) · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Li, L., Chu, W., Langford, J., and Wang, X. (2011) · 2011
Cited alongside, same era.
Towards optimal problem dependent generalization error bounds in statistical learning theory
Xu, Y. and Zeevi, A. (2020) · 2011
Cited alongside, same era.
Decoupling: from dependence to independence
De la Pena, V. and Giné, E. (2012) · 2012
Cited alongside, same era.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., and Wang, Y.-X. (2019) · 2019
Later among the works it cites.
Trustworthy online controlled experiments: A practical guide to a/b testing
Kohavi, R., Tang, D., and Xu, Y. (2020) · 2020
Later among the works it cites.
policytree: Policy learning via doubly robust empirical welfare maximization over trees
Sverdrup, E., Kanodia, A., Zhou, Z., Athey, S., and Wager, S. (2020) · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Uehara, M., Huang, J., and Jiang, N. (2020) · 2020
Later among the works it cites.
Policy learning with observational data
Athey, S. and Wager, S. (2021) · 2021
Later among the works it cites.
Risk minimization from adaptively collected data: Guarantees for supervised and policy learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimax regret treatment choice with covariates or with limited validity of experiments
Stoye, J. (2012) · 2012
Cited alongside, same era.
Estimating optimal treatment regimes from a classification perspective
Zhang, B., Tsiatis, A. A., Davidian, M., Zhang, M., and Laber, E. (2012) · 2012
Cited alongside, same era.
Estimating individualized treatment rules using outcome weighted learning
Zhao, Y., Zeng, D., Rush, A. J., and Kosorok, M. R. (2012) · 2012
Cited alongside, same era.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S. and Goyal, N. (2013) · 2013
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D. X., Chickering, D. M., Portugaly, E., Ray, D., Simard, P., and Snelson, E. (2013) · 2013
Cited alongside, same era.
Openml: networked science in machine learning
Vanschoren, J., van Rijn, J. N., Bischl, B., and Torgo, L. (2013) · 2013
Cited alongside, same era.
Dynamic treatment regimes
Chakraborty, B. and Murphy, S. A. (2014) · 2014
Cited alongside, same era.
Bibaut, A., Kallus, N., Dimakopoulou, M., Chambaz, A., and van der Laan, M. (2021) · 2021
Later among the works it cites.
Confidence intervals for policy evaluation in adaptive experiments
Hadad, V., Hirshberg, D. A., Zhan, R., Wager, S., and Athey, S. (2021) · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z. (2021) · 2021
Later among the works it cites.
Confident off-policy evaluation and selection through self-normalized importance weighting
Kuzborskij, I., Vernade, C., Gyorgy, A., and Szepesvári, C. (2021) · 2021
Later among the works it cites.
Optidice: Offline policy optimization via stationary distribution correction estimation
Lee, J., Jeon, W., Lee, B., Pineau, J., and Kim, K.-E. (2021) · 2021
Later among the works it cites.
Adaptive experimental design: Prospects and applications in political science
Offer-Westort, M., Coppock, A., and Green, D. P. (2021) · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2021) · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and Sun, W. (2021) · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Yin, M. and Wang, Y.-X. (2021) · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Zanette, A., Wainwright, M. J., and Brunskill, E. (2021) · 2021
Later among the works it cites.
Policy learning with adaptively collected data
Zhan, R., Ren, Z., Athey, S., and Zhou, Z. (2021) · 2021
Later among the works it cites.
Offline reinforcement learning under value and density-ratio realizability: the power of gaps
Chen, J. and Jiang, N. (2022) · 2022
Closest in time.
Conformalized survival analysis with adaptive cutoffs
Gui, Y., Hore, R., Ren, Z., and Barber, R. F. (2022) · 2022
Closest in time.
Optimal conservative offline rl with general function approximation via augmented lagrangian
Rashidinejad, P., Zhu, H., Yang, K., Russell, S., and Jiao, J. (2022) · 2022
Closest in time.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L., Li, G., Wei, Y., Chen, Y., and Chi, Y. (2022) · 2022
Closest in time.
Anytime-valid off-policy inference for contextual bandits
Waudby-Smith, I., Wu, L., Ramdas, A., Karampatziakis, N., and Mineiro, P. (2022) · 2022
Closest in time.
The efficacy of pessimism in asynchronous q-learning
Yan, Y., Li, G., Chen, Y., and Fan, J. (2022) · 2022
Closest in time.
Offline reinforcement learning with realizability and single-policy concentrability
Zhan, W., Huang, B., Huang, A., Jiang, N., and Lee, J. (2022) · 2022
Closest in time.
Offline multi-action policy learning: Generalization and optimization
Zhou, Z., Athey, S., and Wager, S. (2022) · 2022
Closest in time.
Causal effect estimation after propensity score trimming with continuous treatments
Branson, Z., Kennedy, E. H., Balakrishnan, S., and Wasserman, L. (2023) · 2023
Closest in time.
Steel: Singularity-aware reinforcement learning
Chen, X., Qi, Z., and Wan, R. (2023) · 2023
Closest in time.
Upper bounds on the natarajan dimensions of some function classes
Jin, Y. (2023) · 2023
Closest in time.
Sensitivity analysis of individual treatment effects: A robust conformal inference approach
Jin, Y., Ren, Z., and Candès, E. J. (2023) · 2023
Closest in time.
Off-policy evaluation beyond overlap: partial identification through smoothness
Khan, S., Saveski, M., and Ugander, J. (2023) · 2023
Closest in time.
Average treatment effect on the treated, under lack of positivity
Liu, Y., Li, H., Zhou, Y., and Matsouaka, R. (2023) · 2023
Closest in time.
Mou, W., Ding, P., Wainwright, M. J., and Bartlett, P. L. (2023) · 2023
Closest in time.
Viper: Provably efficient algorithm for offline rl with neural function approximation
Nguyen-Tang, T. and Arora, R. (2023) · 2023
Closest in time.
Positivity-free policy learning with observational data
Zhao, P., Chambaz, A., Josse, J., and Yang, S. (2023) · 2023
Closest in time.