Fetching the paper…
Reading the bibliography…
We study off-dynamics Reinforcement Learning (RL), where the policy training and deployment environments are different.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Nilim, A · 2005
Earlier work this paper cites.
Off-dynamics reinforcement learning: Training for transfer with domain classifiers
Eysenbach, B · 2006
Earlier work this paper cites.
The robustness-performance tradeoff in markov decision processes
Xu, H · 2006
Earlier work this paper cites.
What are the statistical limits of offline rl with linear function approximation?
Wang, R · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Earlier work this paper cites.
Robust markov decision processes
Wiesemann, W · 2013
Earlier work this paper cites.
Distributionally robust counterpart in markov decision processes
Yu, P · 2015
Earlier work this paper cites.
Robust mdps with k-rectangular uncertainty
Mannor, S · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Earlier work this paper cites.
Generalization and regularization in dqn
Farebrother, J · 2018
Earlier work this paper cites.
Is q-learning provably efficient?
Jin, C · 2018
Earlier work this paper cites.
Optimal treatment allocations in space and time for on-line control of an emerging infectious disease
Laber, E. B · 2018
Earlier work this paper cites.
Sim-to-real transfer of robotic control with dynamics randomization
Peng, X. B · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science
Vershynin, R · 2018
Earlier work this paper cites.
Provably efficient q-learning with low switching cost
Bai, Y · 2019
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Cited alongside, same era.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A · 2020
Cited alongside, same era.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Yang, L · 2020
Cited alongside, same era.
Frequentist regret bounds for randomized least-squares value iteration
Zanette, A · 2020
Cited alongside, same era.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey
Zhao, W · 2020
Cited alongside, same era.
Reward-free rl is no harder than reward-aware rl in linear markov decision processes
Wagenmaker, A. J · 2022
Later among the works it cites.
Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics
Yang, W · 2022
Later among the works it cites.
Computationally efficient horizon-free reinforcement learning for linear mixture mdps
Zhou, D · 2022
Later among the works it cites.
Blanchet, J · 2023
Later among the works it cites.
Robust markov decision processes: Beyond rectangularity
Goyal, V · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Logarithmic regret for reinforcement learning with linear function approximation
He, J · 2021
Cited alongside, same era.
Randomized exploration in reinforcement learning with general value function approximation
Ishfaq, H · 2021
Cited alongside, same era.
Simgan: Hybrid simulator identification for domain adaptation via adversarial reinforcement learning
Jiang, Y · 2021
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation under adaptivity constraints
Wang, T · 2021
Cited alongside, same era.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T · 2021
Cited alongside, same era.
Online policy optimization for robust mdp
Dong, J · 2022
Cited alongside, same era.
He, J · 2023
Later among the works it cites.
Nearly minimax optimal reinforcement learning with linear function approximation
Hu, P · 2023
Later among the works it cites.
Provable and practical: Efficient exploration in reinforcement learning via langevin monte carlo
Ishfaq, H · 2023
Later among the works it cites.
Deep spatial q-learning for infectious disease control
Liu, Z · 2023
Later among the works it cites.
The curious price of distributional robustness in reinforcement learning with a generative model
Shi, L · 2023
Later among the works it cites.
Improved sample complexity bounds for distributionally robust reinforcement learning
Xu, Z · 2023
Later among the works it cites.
Zhao, H · 2023
Later among the works it cites.
Randomized exploration in cooperative multi-agent reinforcement learning
Hsu, H.-L · 2024
Closest in time.
Lu, M · 2024
Closest in time.
Panaganti, K · 2024
Closest in time.
Wasserstein distributionally robust policy evaluation and learning for contextual bandits
Shen, Y · 2024
Closest in time.
Sample complexity of offline distributionally robust linear markov decision processes
Wang, H · 2024
Closest in time.