Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) aims to learn an optimal policy from pre-collected data.
Basic properties of strong mixing conditions
Bradley, R. C. (1986) · 1986
Earlier work this paper cites.
Q-learning
Watkins, C. J. and P. Dayan (1992) · 1992
Earlier work this paper cites.
A note on the efficiency of sandwich covariance matrix estimation
Kauermann, G. and R. J. Carroll (2001) · 2001
Earlier work this paper cites.
Maximal inequalities and empirical central limit theorems
Dedecker, J. and S. Louhichi (2002) · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and J. Langford (2002) · 2002
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., A. Kumar, O. Nachum, G. Tucker, and S. Levine (2020) · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., P. Geurts, and L. Wehenkel (2005) · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., A. Kumar, G. Tucker, and J. Fu (2020) · 2005
Earlier work this paper cites.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M. (2005) · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and B. Recht (2007) · 2007
Earlier work this paper cites.
On the sample complexity of reinforcement learning with policy space generalization
Mou, W., Z. Wen, and X. Chen (2020) · 2008
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Rahimi, A. and B. Recht (2008) · 2008
Earlier work this paper cites.
Deep-brain stimulation for parkinson’s disease
Okun, M. S. (2012) · 2012
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Lee, D.-H. et al. (2013) · 2013
Earlier work this paper cites.
Adaptive deep brain stimulation in advanced parkinson disease
Little, S., A. Pogosyan, S. Neal, B. Zavala, L. Zrinzo, M. Hariz, T. Foltynie, P. Limousin, K. Ashkan, J. FitzGerald, A. L. Green, T. Z. Aziz, and P. Brown (2013, September) · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and J. Ba (2014) · 2014
Earlier work this paper cites.
Efficient and adaptive linear regression in semi-supervised settings
Chakrabortty, A. and T. Cai (2018) · 2018
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Haarnoja, T., A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and A. G. Barto (2018) · 2018
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and N. Jiang (2019) · 2019
Earlier work this paper cites.
Semi-supervised inference: General theory and estimation of means
Zhang, A., L. D. Brown, and T. T. Cai (2019) · 2019
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R., A. Rajeswaran, P. Netrapalli, and T. Joachims (2020) · 2020
Cited alongside, same era.
Semi-supervised reward learning for offline reinforcement learning
Konyushkova, K., K. Zolna, Y. Aytar, A. Novikov, S. Reed, S. Cabi, and N. de Freitas (2020) · 2020
Cited alongside, same era.
Provably good batch off-policy reinforcement learning without great exploration
Liu, Y., A. Swaminathan, A. Agarwal, and E. Brunskill (2020) · 2020
Cited alongside, same era.
Estimating dynamic treatment regimes in mobile health using v-learning
Luckett, D. J., E. B. Laber, A. R. Kahkoska, D. M. Maahs, E. Mayer-Davis, and M. R. Kosorok (2020) · 2020
Cited alongside, same era.
Semi-supervised learning for doubly robust offline policy evaluation
Steel: Singularity-aware reinforcement learning
Chen, X., Z. Qi, and R. Wan (2023) · 2023
Later among the works it cites.
The provable benefit of unsupervised data sharing for offline reinforcement learning
Hu, H., Y. Yang, Q. Zhao, and C. Zhang (2023) · 2023
Later among the works it cites.
How does semi-supervised learning with pseudo-labelers work? a case study
Kou, Y., Z. Chen, Y. Cao, and Q. Gu (2023) · 2023
Later among the works it cites.
Quasi-optimal reinforcement learning with continuous actions
Li, Y., W. Zhou, and R. Zhu (2023) · 2023
Later among the works it cites.
Online estimation and inference for robust policy evaluation in reinforcement learning
Liu, W., J. Tu, Y. Zhang, and X. Chen (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sonabend-W, A., N. Laha, R. Mukherjee, and T. Cai (2020) · 2020
Cited alongside, same era.
MOPO: Model-based offline policy optimization
Yu, T., G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma (2020) · 2020
Cited alongside, same era.
Sparse feature selection makes batch reinforcement learning more sample efficient
Hao, B., Y. Duan, T. Lattimore, C. Szepesvári, and M. Wang (2021) · 2021
Cited alongside, same era.
Is pessimism provably efficient for offline RL?
Jin, Y., Z. Yang, and Z. Wang (2021) · 2021
Cited alongside, same era.
Uehara, M., M. Imaizumi, N. Jiang, N. Kallus, W. Sun, and T. Xie (2021) · 2021
Cited alongside, same era.
Representation matters: Offline pretraining for sequential decision making
Yang, M. and O. Nachum (2021) · 2021
Cited alongside, same era.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Bai, C., L. Wang, Z. Yang, Z.-H. Deng, A. Garg, P. Liu, and Z. Wang (2022) · 2022
Cited alongside, same era.
Adaptive deep brain stimulation: From experimental evidence toward practical implementation
Neumann, W.-J., R. Gilron, S. Little, and G. Tinkhauser (2023) · 2023
Later among the works it cites.
Online bootstrap inference for policy evaluation in reinforcement learning
Ramprasad, P., Y. Li, Z. Yang, Z. Wang, W. W. Sun, and G. Cheng (2023) · 2023
Later among the works it cites.
Semi-supervised off-policy reinforcement learning and value estimation for dynamic treatment regimes
Sonabend-W, A., N. Laha, A. N. Ananthakrishnan, T. Cai, and R. Mukherjee (2023) · 2023
Later among the works it cites.
Projected state-action balancing weights for offline reinforcement learning
Wang, J., Z. Qi, and R. K. Wong (2023) · 2023
Later among the works it cites.
A survey on deep semi-supervised learning
Yang, X., Z. Song, I. King, and Z. Xu (2023) · 2023
Later among the works it cites.
Estimation and inference in distributional reinforcement learning
Zhang, L., Y. Peng, J. Liang, W. Yang, and Z. Zhang (2023) · 2023
Later among the works it cites.
Optimizing pessimism in dynamic treatment regimes: A bayesian learning approach
Zhou, Y., Z. Qi, C. Shi, and L. Li (2023) · 2023
Later among the works it cites.
Reinforcement learning in latent heterogeneous environments
Chen, E. Y., R. Song, and M. I. Jordan (2024) · 2024
Later among the works it cites.
Settling the sample complexity of model-based offline reinforcement learning
Li, G., L. Shi, Y. Chen, Y. Chi, and Y. Wei (2024) · 2024
Later among the works it cites.
Chronic adaptive deep brain stimulation versus conventional stimulation in parkinson’s disease: a blinded randomized feasibility trial
Oehrn, C. R., C. Palmisano, P. A. Starr, S. Little, et al. (2024) · 2024
Later among the works it cites.
Pessimistic causal reinforcement learning with mediators for confounded offline data
Wang, D., C. Shi, S. Luo, and W. W. Sun (2024) · 2024
Later among the works it cites.
Neural network approximation for pessimistic offline reinforcement learning
Wu, D., Y. Jiao, L. Shen, H. Yang, and X. Lu (2024) · 2024
Later among the works it cites.
Reward-relevance-filtered linear offline reinforcement learning
Zhou, A. (2024) · 2024
Later among the works it cites.
Estimating optimal infinite horizon dynamic treatment regimes via pt-learning
Zhou, W., R. Zhu, and A. Qu (2024) · 2024
Later among the works it cites.
Value enhancement of reinforcement learning via efficient and robust trust region optimization
Shi, C., Z. Qi, J. Wang, and F. Zhou (2024) · 2025
Closest in time.