Fetching the paper…
Reading the bibliography…
In real-world scenarios, datasets collected from randomized experiments are often constrained by size, due to limitations in time and budget.
1912
Earlier work this paper cites.
Banach, S. (1922), “Sur les opérations dans les ensembles abstraits et leur application aux équations intégrales,” Fundamenta mathematicae
1922
Earlier work this paper cites.
Bernstein, S. N. (1946), Theory of Probability (in Russian)
1946
Earlier work this paper cites.
Kullback, S. and Leibler, R. A. (1951), “On information and sufficiency,” The annals of mathematical statistics
1951
Earlier work this paper cites.
Nadler Jr, S. B. (1969), “Multi-valued contraction mappings.”
1969
Earlier work this paper cites.
Dedecker, J. and Louhichi, S. (2002), “Maximal inequalities and empirical central limit theorems,” in Empirical process techniques for dependent data
2002
Earlier work this paper cites.
Kakade, S. and Langford, J. (2002), “Approximately optimal approximate reinforcement learning,” in Proceedings of the Nineteenth International Conference on Machine Learning
2002
Earlier work this paper cites.
Murphy, S. A. (2003), “Optimal dynamic treatment regimes,” Journal of the Royal Statistical Society: Series B
2003
Earlier work this paper cites.
Robins, J. M. (2004), “Optimal structural nested models for optimal sequential decisions,” in Proceedings of the Second Seattle Symposium in Biostatistics
2004
Earlier work this paper cites.
Bradley, R. C. (2005), “Basic properties of strong mixing conditions. A survey and some open questions,”
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
Riedmiller, M. (2005), “Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method,” in Machine Learning: ECML 2005: 16th European Conference on Machine Learning, Porto, Portugal, October 3-7, 2005. Proceedings 16
2005
Earlier work this paper cites.
Antos, A., Szepesvári, C., and Munos, R. (2008), “Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path,” Machine Learning
2008
Earlier work this paper cites.
Munos, R. and Szepesvári, C. (2008), “Finite-Time Bounds for Fitted Value Iteration.” Journal of Machine Learning Research
2008
Earlier work this paper cites.
Siciliano, B., Khatib, O., and Kröger, T. (2008), Springer handbook of robotics
2008
Earlier work this paper cites.
Pearl, J. (2009), Causality
2009
Earlier work this paper cites.
Qian, M. and Murphy, S. A. (2011), “Performance guarantees for individualized treatment rules,” Annals of statistics
2011
Earlier work this paper cites.
Todorov, E., Erez, T., and Tassa, Y. (2012), “MuJoCo: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems
2012
Earlier work this paper cites.
Chakraborty, B. and Moodie, E. E. M. (2013), Statistical methods for dynamic treatment regimes
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
Zhang, B., Tsiatis, A. A., Laber, E. B., and Davidian, M. (2013), “Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions,” Biometrika
2013
Earlier work this paper cites.
Puterman, M. L. (2014), Markov decision processes: discrete stochastic dynamic programming
2014
Earlier work this paper cites.
Imbens, G. W. and Rubin, D. B. (2015), Causal inference in statistics, social, and biomedical sciences
2015
Earlier work this paper cites.
LeCun, Y., Bengio, Y., and Hinton, G. (2015), “Deep learning,” nature
2015
Earlier work this paper cites.
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015), “Human-level control through deep reinforcement learning,” nature
2015
Earlier work this paper cites.
Zhao, Y.-Q., Zeng, D., Laber, E. B., and Kosorok, M. R. (2015), “New statistical learning methods for estimating optimal dynamic treatment regimes,” Journal of the American Statistical Association
2015
Earlier work this paper cites.
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016), “Mastering the game of Go with deep neural networks and tree search,” nature
2016
Cited alongside, same era.
Zhang, J. and Bareinboim, E. (2016), “Markov decision processes with unobserved confounders: A causal approach,” Tech. rep., Technical report, Technical Report R-23, Purdue AI Lab
2016
Cited alongside, same era.
Ertefaie, A. and Strawderman, R. L. (2018), “Constructing dynamic treatment regimes over indefinite time horizons,” Biometrika
2018
Cited alongside, same era.
Kallus, N. and Zhou, A. (2018), “Confounding-robust policy improvement,” Advances in neural information processing systems
2018
Cited alongside, same era.
Kleinberg, J., Lakkaraju, H., Leskovec, J., Ludwig, J., and Mullainathan, S. (2018), “Human decisions and machine predictions,” The quarterly journal of economics
Chang, J., Uehara, M., Sreenivas, D., Kidambi, R., and Sun, W. (2021), “Mitigating covariate shift in imitation learning via offline data with partial coverage,” Advances in Neural Information Processing Systems
2021
Later among the works it cites.
2021
Later among the works it cites.
Jin, Y., Yang, Z., and Wang, Z. (2021), “Is pessimism provably efficient for offline rl?” in International Conference on Machine Learning
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Shi, C., Fan, A., Song, R., and Lu, W. (2018), “High-dimensional A-learning for optimal dynamic treatment regimes,” Annals of statistics
2018
Cited alongside, same era.
Sutton, R. S. and Barto, A. G. (2018), Reinforcement learning: An introduction
2018
Cited alongside, same era.
Wang, L., Zhou, Y., Song, R., and Sherwood, B. (2018), “Quantile-optimal treatment regimes,” Journal of the American Statistical Association
2018
Cited alongside, same era.
Xu, Z., Li, Z., Guan, Q., Zhang, D., Li, Q., Nan, J., Liu, C., Bian, W., and Ye, J. (2018), “Large-scale order dispatch in on-demand ride-hailing platforms: A learning and planning approach,” in Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining
2018
Cited alongside, same era.
Zhang, Y., Laber, E. B., Davidian, M., and Tsiatis, A. A. (2018), “Interpretable dynamic treatment regimes,” Journal of the American Statistical Association
2018
Cited alongside, same era.
Chen, J. and Jiang, N. (2019), “Information-theoretic considerations in batch reinforcement learning,” in International Conference on Machine Learning
2019
Cited alongside, same era.
2021
Later among the works it cites.
Mo, W., Qi, Z., and Liu, Y. (2021), “Learning optimal distributionally robust individualized treatment rules,” Journal of the American Statistical Association
2021
Later among the works it cites.
Nie, X. and Wager, S. (2021), “Quasi-oracle estimation of heterogeneous treatment effects,” Biometrika
2021
Later among the works it cites.
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2021), “Bridging offline reinforcement learning and imitation learning: A tale of pessimism,” Advances in Neural Information Processing Systems
2021
Later among the works it cites.
2021
Later among the works it cites.
Wang, L., Yang, Z., and Wang, Z. (2021), “Provably efficient causal reinforcement learning with confounded observational data,” Advances in Neural Information Processing Systems
2021
Later among the works it cites.
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A. (2021), “Bellman-consistent pessimism for offline reinforcement learning,” Advances in neural information processing systems
2021
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
— (2022), “Stateful offline contextual policy evaluation and learning,” in International Conference on Artificial Intelligence and Statistics
2022
Later among the works it cites.
2022
Later among the works it cites.
Qi, Z., Tang, J., Fang, E., and Shi, C. (2022), “Offline personalized pricing with censored demand,” in Offline Personalized Pricing with Censored Demand: Qi, Zhengling— uTang, Jingwen— uFang, Ethan— uShi, Cong
2022
Later among the works it cites.
2022
Later among the works it cites.
Uehara, M. and Sun, W. (2022), “Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage,” in International Conference on Learning Representations
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Zhou, W., Zhu, R., and Qu, A. (2022), “Estimating optimal infinite horizon dynamic treatment regimes via pt-learning,” Journal of the American Statistical Association
2022
Later among the works it cites.
2023
Later among the works it cites.
OpenAI (2023), “GPT-4 Technical Report,”
2023
Later among the works it cites.
Zhou, Y., Qi, Z., Shi, C., and Li, L. (2023), “Optimizing Pessimism in Dynamic Treatment Regimes: A Bayesian Learning Approach,” in International Conference on Artificial Intelligence and Statistics
2023
Later among the works it cites.