Fetching the paper…
Reading the bibliography…
Dynamic decision-making under distributional shifts is of fundamental interest in theory and applications of reinforcement learning: The distribution of the environment in which the data is collected can differ from that of the environment in which the model is deployed.
Characterizing attacks on deep reinforcement learning
Pan, X., Xiao, C., He, W., Yang, S., Peng, J., Sun, M., Yi, J., Yang, Z., Liu, M., Li, B., et al. (2019) · 1907
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Minimax control of discrete-time stochastic systems
González-Trejo, J. I., Hernández-Lerma, O., and Hoyos-Reyes, L. F. (2002) · 2002
Earlier work this paper cites.
Learning rates for q-learning
Even-Dar, E., Mansour, Y., and Bartlett, P. (2003) · 2003
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. (2005) · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L. (2005) · 2005
Earlier work this paper cites.
Dataset shift in machine learning
Quinonero-Candela, J., Sugiyama, M., Schwaighofer, A., and Lawrence, N. D. (2008) · 2008
Earlier work this paper cites.
Distributionally robust markov decision processes
Xu, H. and Mannor, S. (2010) · 2010
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2013) · 2013
Earlier work this paper cites.
Kullback-leibler divergence constrained distributionally robust optimization
Hu, Z. and Hong, L. J. (2013) · 2013
Earlier work this paper cites.
Stochastic Approximation and Recursive Algorithms and Applications
Kushner, H. and Yin, G. (2013) · 2013
Earlier work this paper cites.
Robust markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B. (2013) · 2013
Earlier work this paper cites.
Lectures on Stochastic Programming: Modeling and Theory, Second Edition
Shapiro, A., Dentcheva, D., and Ruszczyński, A. (2014) · 2014
Earlier work this paper cites.
“18. s997: High dimensional statistics lecture notes
Rigollet, P. (2015) · 2015
Cited alongside, same era.
Tactics of adversarial attack on deep reinforcement learning agents
Lin, Y.-C., Hong, Z.-W., Liao, Y.-H., Shih, M.-L., Liu, M.-Y., and Sun, M. (2017) · 2017
Cited alongside, same era.
An introduction to deep reinforcement learning
François-Lavet, V., Henderson, P., Islam, R., Bellemare, M. G., Pineau, J., et al. (2018) · 2018
Cited alongside, same era.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Sidford, A., Wang, M., Wu, X., Yang, L., and Ye, Y. (2018) · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
Non-asymptotic analysis of biased stochastic approximation scheme
Towards theoretical understandings of robust markov decision processes: Sample complexity and asymptotics
Yang, W., Zhang, L., and Zhang, Z. (2021) · 2021
Later among the works it cites.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhou, Z., Zhou, Z., Bai, Q., Qiu, L., Blanchet, J., and Glynn, P. (2021) · 2021
Later among the works it cites.
Finite-sample analysis of nonlinear stochastic approximation with applications in reinforcement learning
Chen, Z., Zhang, S., Doan, T. T., Clarke, J.-P., and Maguluri, S. T. (2022) · 2022
Later among the works it cites.
Settling the sample complexity of model-based offline reinforcement learning
Li, G., Shi, L., Chen, Y., Chi, Y., and Wei, Y. (2022) · 2022
Later among the works it cites.
Distributionally robust q q -learning
Liu, Z., Bai, Q., Blanchet, J., Dong, P., Xu, W., Zhou, Z., and Zhou, Z. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karimi, B., Miasojedow, B., Moulines, E., and Wai, H.-T. (2019) · 2019
Cited alongside, same era.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A., Kakade, S., and Yang, L. F. (2020) · 2020
Cited alongside, same era.
Finite-sample analysis of contractive stochastic approximation using smooth convex envelopes
Chen, Z., Maguluri, S. T., Shakkottai, S., and Shanmugam, K. (2020) · 2020
Cited alongside, same era.
Distributionally robust policy evaluation and learning in offline contextual bandits
Si, N., Zhang, F., Zhou, Z., and Blanchet, J. (2020) · 2020
Cited alongside, same era.
Learning models with uniform performance via distributionally robust optimization
Duchi, J. C. and Namkoong, H. (2021) · 2021
Cited alongside, same era.
Instance-optimality in optimal value estimation: Adaptivity via variance-reduced q-learning
Khamaru, K., Xia, E., Wainwright, M. J., and Jordan, M. I. (2021) · 2021
Cited alongside, same era.
Is q-learning minimax optimal? a tight sample complexity analysis
Li, G., Cai, C., Chen, Y., Gu, Y., Wei, Y., and Chi, Y. (2021) · 2021
Cited alongside, same era.
Distributionally robust modeling of optimal control
Shapiro, A. (2022) · 2022
Later among the works it cites.
Distributionally robust model-based offline reinforcement learning with near-optimal sample complexity
Shi, L. and Chi, Y. (2022) · 2022
Later among the works it cites.
Finite-time error bounds of biased stochastic approximation with application to td-learning
Wang, G. (2022) · 2022
Later among the works it cites.
A statistical analysis of polyak-ruppert averaged q-learning
Li, X., Yang, W., Liang, J., Zhang, Z., and Jordan, M. I. (2023) · 2023
Closest in time.
Avoiding model estimation in robust markov decision processes with a generative model
Yang, W., Wang, H., Kozuno, T., Jordan, S. M., and Zhang, Z. (2023) · 2023
Closest in time.
The curious price of distributional robustness in reinforcement learning with a generative model
Shi, L., Li, G., Wei, Y., Chen, Y., Geist, M., and Chi, Y. (2024) · 2024
Closest in time.
On the foundation of distributionally robust reinforcement learning
Wang, S., Si, N., Blanchet, J., and Zhou, Z. (2024) · 2024
Closest in time.