Fetching the paper…
Reading the bibliography…
Algorithmic decisions and recommendations are used in many high-stakes decision-making settings such as criminal justice, medicine, and public policy.
On the application of probability theory to agricultural experiments. essay on principles. section 9
Neyman, J. (1990 [1923]) · 1923
Earlier work this paper cites.
Tacc fragging procedures
PACAF, H. Q. (1969) · 1969
Earlier work this paper cites.
Estimating causal effects of treatments in randomized and nonrandomized studies
Rubin, D. B. (1974) · 1974
Earlier work this paper cites.
Randomization analysis of experimental data: The fisher randomization test comment
Rubin, D. B. (1980) · 1980
Earlier work this paper cites.
The central role of the propensity score in observational studies for causal effects
Rosenbaum, P. R. and Rubin, D. B. (1983) · 1983
Earlier work this paper cites.
On the conductance of order Markov chains
Karzanov, A. and Khachiyan, L. (1990) · 1990
Earlier work this paper cites.
Percentile performance criteria for limiting average markov decision processes
Filar, J. A., Krass, D., and Ross, K. W. (1995) · 1995
Earlier work this paper cites.
Chance-constrained model predictive control
Schwarm, A. T. and Nikolaou, M. (1999) · 1999
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Friedman, J. H. (2001) · 2001
Earlier work this paper cites.
A model to predict survival in patients with end-stage liver disease
Kamath, P. S., Wiesner, R. H., Malinchoc, M., Kremers, W., Therneau, T. M., Kosberg, C. L., D’Amico, G., Dickson, E. R., and Kim, W. R. (2001) · 2001
Earlier work this paper cites.
Td algorithm for the variance of return and mean-variance reinforcement learning
Sato, M., Kimura, H., and Kobayashi, S. (2001) · 2001
Earlier work this paper cites.
Pu, H. and Zhang, B. (2020) · 2002
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Geibel, P. and Wysotzki, F. (2005) · 2005
Earlier work this paper cites.
Gaussian processes for machine learning
Rasmussen, C. E. and Williams, C. K. I. (2006) · 2006
Earlier work this paper cites.
Gaussian processes for machine learning
Rasmussen, C. E. and Williams, C. K. I. (2006) · 2006
Earlier work this paper cites.
Percentile optimization in uncertain markov decision processes with application to efficient exploration
Delage, E. and Mannor, S. (2007) · 2007
Earlier work this paper cites.
Minimax-regret treatment choice with missing outcome data
Manski, C. F. (2007) · 2007
Earlier work this paper cites.
The offset tree for learning with partial labels
Beygelzimer, A. and Langford, J. (2009) · 2009
Earlier work this paper cites.
The importance of pessimism in fixed-dataset policy optimization
Buckman, J., Gelada, C., and Bellemare, M. G. (2020) · 2009
Earlier work this paper cites.
Bart: Bayesian additive regression trees
Chipman, H. A., George, E. I., and McCulloch, R. E. (2010) · 2010
Earlier work this paper cites.
Percentile optimization for markov decision processes with parameter uncertainty
Delage, E. and Mannor, S. (2010) · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Earlier work this paper cites.
Bart: Bayesian additive regression trees
Chipman, H. A., George, E. I., and McCulloch, R. E. (2010) · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L. (2011) · 2011
Earlier work this paper cites.
Bayesian nonparametric modeling for causal inference
Hill, J. L. (2011) · 2011
Earlier work this paper cites.
Performance guarantees for individualized treatment rules
Qian, M. and Murphy, S. A. (2011) · 2011
Earlier work this paper cites.
The problem of metrics: Assessing progress and effectiveness in the Vietnam War
Daddis, G. A. (2012) · 2012
Earlier work this paper cites.
Estimating optimal treatment regimes from a classification perspective
Zhang, B., Tsiatis, A. A., Davidian, M., Zhang, M., and Laber, E. (2012) · 2012
Cited alongside, same era.
Estimating individualized treatment rules using outcome weighted learning
Zhao, Y., Zeng, D., Rush, A. J., and Kosorok, M. R. (2012) · 2012
Cited alongside, same era.
Automatic ad format selection via contextual bandits
Tang, L., Rosales, R., Singh, A., and Agarwal, D. (2013) · 2013
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F. (2015) · 2015
Cited alongside, same era.
Batch learning from logged bandit feedback through counterfactual risk minimization
Swaminathan, A. and Joachims, T. (2015) · 2015
Cited alongside, same era.
Mean-variance and value at risk in multi-armed bandit problems
Vakili, S. and Zhao, Q. (2015) · 2015
Randomized control trial evaluation of the implementation of the psa-dmf system in dane county, wi
Greiner, J., Halen, R., Stubenberg, M., and Griffin, C. L. (2020) · 2020
Later among the works it cites.
Bayesian regression tree models for causal inference: Regularization, confounding, and heterogeneous effects (with discussion)
Hahn, P. R., Murray, J. S., and Carvalho, C. M. (2020) · 2020
Later among the works it cites.
Bayesian regression tree models for causal inference: Regularization, confounding, and heterogeneous effects (with discussion)
Hahn, P. R., Murray, J. S., and Carvalho, C. M. (2020) · 2020
Later among the works it cites.
Policy learning with observational data
Athey, S. and Wager, S. (2021) · 2021
Later among the works it cites.
Safe policy learning through extrapolation: Application to pre-trial risk assessment
Ben-Michael, E., Greiner, D. J., Imai, K., and Jiang, Z. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stochastic control with uncertain parameters via chance constrained control
Vitus, M. P., Zhou, Z., and Tomlin, C. J. (2015) · 2015
Cited alongside, same era.
Stochastic linear model predictive control with chance constraints – a review
Farina, M., Giulioni, L., and Scattolini, R. (2016) · 2016
Cited alongside, same era.
Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy
Luedtke, A. R. and Van Der Laan, M. J. (2016) · 2016
Cited alongside, same era.
A nonparametric bayesian analysis of heterogenous treatment effects in digital experimentation
Taddy, M., Gardner, M., Chen, L., and Draper, D. (2016) · 2016
Cited alongside, same era.
A nonparametric bayesian analysis of heterogenous treatment effects in digital experimentation
Taddy, M., Gardner, M., Chen, L., and Draper, D. (2016) · 2016
Cited alongside, same era.
Stan: A probabilistic programming language
Carpenter, B., Gelman, A., Hoffman, M. D., Lee, D., Goodrich, B., Betancourt, M., Brubaker, M., Guo, J., Li, P., and Riddell, A. (2017) · 2017
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z. (2021) · 2021
Later among the works it cites.
Minimax-optimal policy learning under unobserved confounding
Kallus, N. and Zhou, A. (2021) · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2021) · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and Sun, W. (2021) · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A. (2021) · 2021
Later among the works it cites.
Optimal decision rules under partial identification
Yata, K. (2021) · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Yin, M. and Wang, Y.-X. (2021) · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Zanette, A., Wainwright, M. J., and Brunskill, E. (2021) · 2021
Later among the works it cites.
Safe policy learning through extrapolation: Application to pre-trial risk assessment
Ben-Michael, E., Greiner, D. J., Imai, K., and Jiang, Z. (2021) · 2021
Later among the works it cites.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Bai, C., Wang, L., Yang, Z., Deng, Z., Garg, A., Liu, P., and Wang, Z. (2022) · 2022
Later among the works it cites.
Offline reinforcement learning under value and density-ratio realizability: the power of gaps
Chen, J. and Jiang, N. (2022) · 2022
Later among the works it cites.
Policy learning” without”overlap: Pessimism and generalized empirical bernstein’s inequality
Jin, Y., Ren, Z., Yang, Z., and Wang, Z. (2022) · 2022
Later among the works it cites.
What’s the harm? sharp bounds on the fraction negatively affected by treatment
Kallus, N. (2022) · 2022
Later among the works it cites.
Safe chance constrained reinforcement learning for batch process control
Mowbray, M., Petsagkourakis, P., del Rio-Chanona, E. A., and Zhang, D. (2022) · 2022
Later among the works it cites.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L., Li, G., Wei, Y., Chen, Y., and Chi, Y. (2022) · 2022
Later among the works it cites.
On the safety of interpretable machine learning: A maximum deviation approach
Wei, D., Nair, R., Dhurandhar, A., Varshney, K. R., Daly, E. M., and Singh, M. (2022) · 2022
Later among the works it cites.
The efficacy of pessimism in asynchronous q-learning
Yan, Y., Li, G., Chen, Y., and Fan, J. (2022) · 2022
Later among the works it cites.
Safe policy learning under regression discontinuity designs
Zhang, Y., Ben-Michael, E., and Imai, K. (2022) · 2022
Later among the works it cites.
Offline multi-action policy learning: Generalization and optimization
Zhou, Z., Athey, S., and Wager, S. (2022) · 2022
Later among the works it cites.
Experimental evaluation of computer-assisted human decision-making: Application to pretrial risk assessment instrument (with discussion)
Imai, K., Jiang, Z., Greiner, D. J., Halen, R., and Shin, S. (2023) · 2023
Closest in time.
Voting rights, markov chains, and optimization by short bursts
Cannon, S., Goldbloom-Helzner, A., Gupta, V., Matthews, J., and Suwal, B. (2023) · 2023
Closest in time.
Policy learning with counterfactual asymmetric utilities
Ben-Michael, E., Imai, K., and Jiang, Z. (2024) · 2024
Closest in time.