Fetching the paper…
Reading the bibliography…
Offline reinforcement learning enables evaluation and optimization of sequential decisions from historical data, when it is not possible to deploy new policies online due to safety, cost, and other concerns.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
Stone CJ (1980) Optimal rates of convergence for nonparametric estimators. The annals of Statistics 1348–1360
1980
Earlier work this paper cites.
Tsitsiklis JN, Van Roy B (1996) Feature-based methods for large scale dynamic programming. Machine Learning 22(1):59–94
1996
Earlier work this paper cites.
Talluri K, van Ryzin G (1998) An analysis of bid-price controls for network revenue management. Management Science 44(11):1577–1593
1998
Earlier work this paper cites.
Spirtes P, Glymour C, Scheines R (2001) Causation, prediction, and search (The MIT press)
2001
Earlier work this paper cites.
Kakade S, Langford J (2002) Approximately optimal approximate reinforcement learning. Proceedings of the Nineteenth International Conference on Machine Learning , 267–274
2002
Earlier work this paper cites.
Murphy S (2003) Optimal dynamic treatment regimes. Journal of the Royal Statistical Society Series B: Statistical Methodology 65(2):331–355
2003
Earlier work this paper cites.
2004
Earlier work this paper cites.
Robins JM (2004) Optimal structural nested models for optimal sequential decisions. Proceedings of the Second Seattle Symposium in Biostatistics: analysis of correlated data , 189–326 (Springer)
2004
Earlier work this paper cites.
Tsybakov AB (2004) Optimal aggregation of classifiers in statistical learning. The Annals of Statistics 32(1):135–166
2004
Earlier work this paper cites.
Li L, Walsh TJ, Littman ML (2006) Towards a unified theory of state abstraction for mdps. AI&M 1(2):3
2006
Earlier work this paper cites.
Van Roy B (2006) Performance loss bounds for approximate value iteration with state aggregation. Mathematics of Operations Research 31(2):234–244
2006
Earlier work this paper cites.
Adelman D (2007) Dynamic bid prices in revenue management. Operations Research 55(4):647–661
2007
Earlier work this paper cites.
Fan J, Lv J (2008) Sure independence screening for ultrahigh dimensional feature space. Journal of the Royal Statistical Society Series B: Statistical Methodology 70(5):849–911
2008
Earlier work this paper cites.
Munos R, Szepesvári C (2008) Finite-time bounds for fitted value iteration. Journal of Machine Learning Research 9(5)
2008
Earlier work this paper cites.
Neumann G, Peters J (2008) Fitted q-iteration by advantage weighted regression. Advances in neural information processing systems 21
2008
Earlier work this paper cites.
Bickel PJ, Ritov Y, Tsybakov AB (2009) Simultaneous analysis of Lasso and Dantzig selector. The Annals of Statistics 37(4):1705–1732
2009
Earlier work this paper cites.
Meinshausen N, Yu B (2009) Lasso-type recovery of sparse representations for high-dimensional data. The Annals of Statistics 37(1):246–270
2009
Earlier work this paper cites.
Zhou S (2009) Thresholding procedures for high dimensional variable selection and statistical estimation. Advances in Neural Information Processing Systems 22 , 2304–2312
2009
Earlier work this paper cites.
2010
Earlier work this paper cites.
van de Geer S, Bühlmann P, Zhou S (2011) The adaptive and the thresholded Lasso for potentially misspecified models (and a lower bound for the Lasso). Electronic Journal of Statistics 5:688–749
2011
Earlier work this paper cites.
Chakraborty B, Murphy SA (2014) Dynamic treatment regimes. Annual review of statistics and its application 1:447–464
2014
Earlier work this paper cites.
Schulte PJ, Tsiatis AA, Laber EB, Davidian M (2014) Q-and a-learning methods for estimating optimal dynamic treatment regimes. Statistical science: a review journal of the Institute of Mathematical Statistics 29(4):640
2014
Earlier work this paper cites.
Thomas P, Theocharous G, Ghavamzadeh M (2015) High confidence policy improvement. International Conference on Machine Learning , 2380–2388
2015
Earlier work this paper cites.
2016
Cited alongside, same era.
Jiang N, Li L (2016) Doubly robust off-policy value evaluation for reinforcement learning. Proceedings of the 33rd International Conference on Machine Learning
2016
Cited alongside, same era.
Belghazi MI, Baratin A, Rajeshwar S, Ozair S, Bengio Y, Courville A, Hjelm D (2018) Mutual information neural estimation. International conference on machine learning , 531–540 (PMLR)
2018
Cited alongside, same era.
Chernozhukov V, Chetverikov D, Demirer M, Duflo E, Hansen C, Newey W, Robins J (2018) Double/debiased machine learning for treatment and structural parameters
2018
Cited alongside, same era.
Farias V, Li A, Peng T, Zheng A (2022) Markovian interference in experiments. Advances in Neural Information Processing Systems 35:535–549
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Foster DJ, Syrgkanis V (2023) Orthogonal statistical learning. The Annals of Statistics 51(3):879–908
2023
Later among the works it cites.
Fu Y, Fisher M (2023) The value of social media data in fashion forecasting. Manufacturing & Service Operations Management 25(3):1136–1154
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Liu Q, Li L, Tang Z, Zhou D (2018) Breaking the curse of horizon: Infinite-horizon off-policy estimation. Advances in Neural Information Processing Systems , 5356–5366
2018
Cited alongside, same era.
Wager S, Athey S (2018) Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association 113(523):1228–1242
2018
Cited alongside, same era.
Duan Y, Ke T, Wang M (2019) State aggregation learning from markov transition data. Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
Künzel SR, Sekhon JS, Bickel PJ, Yu B (2019) Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the national academy of sciences 116(10):4156–4165
2019
Cited alongside, same era.
Le H, Voloshin C, Yue Y (2019) Batch policy learning under constraints. International Conference on Machine Learning , 3703–3712 (PMLR)
2019
Cited alongside, same era.
Li F, Thomas LE, Li F (2019) Addressing extreme propensity scores via the overlap weights. American journal of epidemiology 188(1):250–257
2019
Cited alongside, same era.
Wainwright MJ (2019) High-dimensional statistics: A non-asymptotic viewpoint , volume 48 (Cambridge university press)
2019
Cited alongside, same era.
2023
Later among the works it cites.
Liu YR, Huang B, Zhu Z, Tian H, Gong M, Yu Y, Zhang K (2023) Learning world models with identifiable factorization. Advances in Neural Information Processing Systems , volume 36
2023
Later among the works it cites.
2023
Later among the works it cites.
Pan HR, Schölkopf B (2023) Learning endogenous representation in reinforcement learning via advantage estimation. Causal Representation Learning Workshop at NeurIPS 2023 , URL https://openreview.net/forum?id=aCiFCj4rEA
2023
Later among the works it cites.
Sinclair SR, Frujeri FV, Cheng CA, Marshall L, Barbalho HDO, Li J, Neville J, Menache I, Swaminathan A (2023) Hindsight learning for mdps with exogenous inputs. International Conference on Machine Learning , 31877–31914 (PMLR)
2023
Later among the works it cites.
Xie C, Yang W, Zhang Z (2023) Semiparametrically efficient off-policy evaluation in linear markov decision processes. International Conference on Machine Learning , 38227–38257 (PMLR)
2023
Later among the works it cites.
2024
Closest in time.
Hu Y, Kallus N, Uehara M (2024) Fast rates for the regret of offline reinforcement learning. Mathematics of Operations Research
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Zhou A (2024) Reward-relevance-filtered linear offline reinforcement learning. International Conference on Artificial Intelligence and Statistics , 3025–3033 (PMLR)
2024
Closest in time.
Bennouna A, Pachamanova D, Perakis G, Skali Lami O (2025) Learning the minimal representation of a continuous state-space markov decision process from transition data. Management Science 71(6):5162–5184
2025
Closest in time.
2025
Closest in time.
Hu Y, Li S, Wager S (2025) Optimal targeting in dynamic systems. arXiv preprint arXiv:2507.00312
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Cao D, Zhou A (2026) Structured difference-of-q via orthogonal learning. AISTATS
2026
Closest in time.
Javurek E, Melnychuk V, Schweisthal J, Hess K, Frauen D, Feuerriegel S (2026) An orthogonal learner for individualized outcomes in markov decision processes. International Conference on Learning Representations , volume 2026, 149774–149799
2026
Closest in time.
Jia N, Johari R, Si N, Zheng Z (2026) Experimentation for different scheduling policies on queues: Mixed differences-in-q estimators based on little’s law. Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2 , 2074–2085
2085
Closest in time.