Fetching the paper…
Reading the bibliography…
This paper endeavors to augment the robustness of offline reinforcement learning (RL) in scenarios laden with heavy-tailed rewards, a prevalent circumstance in real-world applications.
Doubly robust bias reduction in infinite horizon off-policy estimation
Tang, Z., Feng, Y., Li, L., Zhou, D., and Liu, Q. (2019) · 1910
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovskij, A. S. and Yudin, D. B. (1983) · 1983
Earlier work this paper cites.
Asymptotic properties of the bootstrap for heavy-tailed distributions
Hall, P. (1990) · 1990
Earlier work this paper cites.
The space complexity of approximating the frequency moments
Alon, N., Matias, Y., and Szegedy, M. (1996) · 1996
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Least-squares temporal difference learning
Boyan, J. A. (1999) · 1999
Earlier work this paper cites.
Alpha-stable modeling of noise and robust time-delay estimation in the presence of impulsive noise
Georgiou, P. G., Tsakalides, P., and Kyriakakis, C. (1999) · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D. (2000) · 2000
Earlier work this paper cites.
Image denoising: A nonlinear robust statistical approach
Hamza, A. B. and Krim, H. (2001) · 2001
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R. (2003) · 2003
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020) · 2004
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L. (2005) · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L. (2005) · 2005
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C. (2008) · 2008
Earlier work this paper cites.
A convergent o(n) algorithm for off-policy temporal-difference learning with linear function approximation
Sutton, R. S., Szepesvári, C., and Maei, H. R. (2008) · 2008
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. (2011) · 2011
Earlier work this paper cites.
Dynamic programming and optimal control: Volume I
Bertsekas, D. (2012) · 2012
Earlier work this paper cites.
Bandits with heavy tail
Bubeck, S., Cesa-Bianchi, N., and Lugosi, G. (2013) · 2013
Earlier work this paper cites.
Robust markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B. (2013) · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Geometric median and robust estimation in banach spaces
Minsker, S. (2015) · 2015
Earlier work this paper cites.
High-confidence off-policy evaluation
Thomas, P., Theocharous, G., and Ghavamzadeh, M. (2015) · 2015
Earlier work this paper cites.
Sub-gaussian mean estimators
Devroye, L., Lerasle, M., Lugosi, G., and Oliveira, R. I. (2016) · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L. (2016) · 2016
Earlier work this paper cites.
Robust mdps with k-rectangular uncertainty
Mannor, S., Mebel, O., and Xu, H. (2016) · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E. (2016) · 2016
Earlier work this paper cites.
A new process uncertainty robust student’st based kalman filter for sins/gps integration
Huang, Y. and Zhang, Y. (2017) · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.
More robust doubly robust off-policy evaluation
Farajtabar, M., Chow, Y., and Ghavamzadeh, M. (2018) · 2018
Earlier work this paper cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D. (2018) · 2018
Earlier work this paper cites.
Error modelling for multi-sensor measurements in infrastructure-free indoor navigation
Ruotsalainen, L., Kirkko-Jaakkola, M., Rantanen, J., and Mäkelä, M. (2018) · 2018
Earlier work this paper cites.
Almost optimal algorithms for linear stochastic bandits with heavy-tailed payoffs
Shao, H., Yu, X., King, I., and Lyu, M. R. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Earlier work this paper cites.
Exponentially weighted imitation learning for batched historical data
Wang, Q., Xiong, J., Han, L., Liu, H., Zhang, T., et al. (2018) · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N. (2019) · 2019
Cited alongside, same era.
Thompson sampling on symmetric alpha-stable bandits
Dubey, A. and Pentland, A. S. (2019) · 2019
Cited alongside, same era.
Importance sampling policy evaluation with an estimated behavior policy
Hanna, J., Niekum, S., and Stone, P. (2019) · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S. (2019) · 2019
Cited alongside, same era.
Batch policy learning under constraints
Le, H., Voloshin, C., and Yue, Y. (2019) · 2019
Cited alongside, same era.
Empirical study of off-policy policy evaluation for reinforcement learning
Voloshin, C., Le, H. M., Jiang, N., and Yue, Y. (2021) · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A. (2021) · 2021
Later among the works it cites.
Near-optimal offline reinforcement learning via double variance reduction
Yin, M., Bai, Y., and Wang, Y.-X. (2021) · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C. (2021) · 2021
Later among the works it cites.
Robust policy gradient against strong data corruption
Zhang, X., Chen, Y., Zhu, X., and Sun, W. (2021) · 2021
Later among the works it cites.
Median-of-means approach for repeated measures data
Zhang, Y. and Liu, P. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal algorithms for lipschitz bandits with heavy-tailed rewards
Lu, S., Wang, G., Hu, Y., and Zhang, L. (2019) · 2019
Cited alongside, same era.
Mean estimation and regression under heavy-tailed distributions: A survey
Lugosi, G. and Mendelson, S. (2019) · 2019
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O., Chow, Y., Dai, B., and Li, L. (2019) · 2019
Cited alongside, same era.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., and Wang, Y.-X. (2019) · 2019
Cited alongside, same era.
Coindice: Off-policy confidence interval estimation
Dai, B., Nachum, O., Chow, Y., Li, L., Szepesvari, C., and Schuurmans, D. (2020) · 2020
Cited alongside, same era.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y., Jia, Z., and Wang, M. (2020) · 2020
Cited alongside, same era.
Later among the works it cites.
Breaking the moments condition barrier: No-regret algorithm for bandits with super heavy-tailed payoffs
Zhong, H., Huang, J., Yang, L., and Wang, L. (2021) · 2021
Later among the works it cites.
No-regret reinforcement learning with heavy-tailed rewards
Zhuang, V. and Sui, Y. (2021) · 2021
Later among the works it cites.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Bai, C., Wang, L., Yang, Z., Deng, Z.-H., Garg, A., Liu, P., and Wang, Z. (2022) · 2022
Later among the works it cites.
On well-posedness and minimax optimal rates of nonparametric q-function estimation in off-policy evaluation
Chen, X. and Qi, Z. (2022) · 2022
Later among the works it cites.
On solutions of the distributional bellman equation
Gerstenberg, J., Neininger, R., and Spiegel, D. (2022) · 2022
Later among the works it cites.
Model-based offline reinforcement learning with pessimism-modulated dynamics belief
Guo, K., Shao, Y., and Geng, Y. (2022) · 2022
Later among the works it cites.
Doubly robust distributionally robust off-policy evaluation and learning
Kallus, N., Mao, X., Wang, K., and Zhou, Z. (2022) · 2022
Later among the works it cites.
Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning
Kallus, N. and Uehara, M. (2022) · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S. (2022) · 2022
Later among the works it cites.
Batch policy learning in average reward markov decision processes
Liao, P., Qi, Z., Wan, R., Klasnja, P., and Murphy, S. A. (2022) · 2022
Later among the works it cites.
Mildly conservative q-learning for offline reinforcement learning
Lyu, J., Ma, X., Li, X., and Lu, Z. (2022) · 2022
Later among the works it cites.
A sharp characterization of linear estimators for offline policy evaluation
Perdomo, J. C., Krishnamurthy, A., Bartlett, P., and Kakade, S. (2022) · 2022
Later among the works it cites.
A review of off-policy evaluation in reinforcement learning
Uehara, M., Shi, C., and Kallus, N. (2022) · 2022
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and Sun, W. (2022) · 2022
Later among the works it cites.
Reliable off-policy evaluation for reinforcement learning
Wang, J., Gao, R., and Zha, H. (2022) · 2022
Later among the works it cites.
Quantile off-policy evaluation via deep conditional generative learning
Xu, Y., Shi, C., Luo, S., Wang, L., and Song, R. (2022) · 2022
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
Zhan, W., Huang, B., Huang, A., Jiang, N., and Lee, J. (2022) · 2022
Later among the works it cites.
Corruption-robust offline reinforcement learning
Zhang, X., Chen, Y., Zhu, X., and Sun, W. (2022) · 2022
Later among the works it cites.
Steel: Singularity-aware reinforcement learning
Chen, X., Qi, Z., and Wan, R. (2023) · 2023
Closest in time.
Robust markov decision processes: Beyond rectangularity
Goyal, V. and Grand-Clement, J. (2023) · 2023
Closest in time.
Online estimation and inference for robust policy evaluation in reinforcement learning
Liu, W., Tu, J., Zhang, Y., and Chen, X. (2023) · 2023
Closest in time.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Prudencio, R. F., Maximo, M. R. O. A., and Colombini, E. L. (2023) · 2023
Closest in time.
The statistical benefits of quantile temporal-difference learning for value estimation
Rowland, M., Tang, Y., Lyle, C., Munos, R., Bellemare, M. G., and Dabney, W. (2023) · 2023
Closest in time.
Value enhancement of reinforcement learning via efficient and robust trust region optimization
Shi, C., Qi, Z., Wang, J., and Zhou, F. (2023) · 2023
Closest in time.
Offline RL with no OOD actions: In-sample learning via implicit value regularization
Xu, H., Jiang, L., Li, J., Yang, Z., Wang, Z., Chan, V. W. K., and Zhan, X. (2023) · 2023
Closest in time.
In-sample actor critic for offline reinforcement learning
Zhang, H., Mao, Y., Wang, B., He, S., Xu, Y., and Ji, X. (2023) · 2023
Closest in time.
Bi-level offline policy optimization with limited exploration
Zhou, W. (2023) · 2023
Closest in time.
Optimizing pessimism in dynamic treatment regimes: A bayesian learning approach
Zhou, Y., Qi, Z., Shi, C., and Li, L. (2023) · 2023
Closest in time.