Fetching the paper…
Reading the bibliography…
We study the adversarial robustness in offline reinforcement learning.
Sample-optimal parametric q-learning using linearly additive features
Yang, L. F. and M. Wang (2019) · 1902
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., J. Fu, G. Tucker, and S. Levine (2019) · 1906
Earlier work this paper cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., S. M. Kakade, J. D. Lee, and G. Mahajan (2019) · 1908
Earlier work this paper cites.
Recent advances in algorithmic high-dimensional robust statistics
Diakonikolas, I. and D. M. Kane (2019) · 1911
Earlier work this paper cites.
Corruption robust exploration in episodic reinforcement learning
Lykouris, T., M. Simchowitz, A. Slivkins, and W. Sun (2019) · 1911
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Wu, Y., G. Tucker, and O. Nachum (2019) · 1911
Earlier work this paper cites.
The behavior of maximum likelihood estimates under nonstandard conditions
Huber, P. J. et al. (1967) · 1967
Earlier work this paper cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Siegel, N. Y., J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, N. Heess, and M. Riedmiller (2020) · 2002
Earlier work this paper cites.
Morel: Model-based offline reinforcement learning
Kidambi, R., A. Rajeswaran, P. Netrapalli, and T. Joachims (2020) · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., A. Kumar, G. Tucker, and J. Fu (2020) · 2005
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
Yu, T., G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma (2020) · 2005
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., A. Zhou, G. Tucker, and S. Levine (2020) · 2006
Earlier work this paper cites.
Robust linear regression: Optimal rates in polynomial time
Bakshi, A. and A. Prasad (2020) · 2007
Earlier work this paper cites.
Provably good batch reinforcement learning without great exploration
Liu, Y., A. Swaminathan, A. Agarwal, and E. Brunskill (2020) · 2007
Earlier work this paper cites.
Near optimal provable uniform convergence in off-policy evaluation for reinforcement learning
Yin, M., Y. Bai, and Y.-X. Wang (2020) · 2007
Earlier work this paper cites.
The importance of pessimism in fixed-dataset policy optimization
Buckman, J., C. Gelada, and M. G. Bellemare (2020) · 2009
Cited alongside, same era.
Online markov decision processes
Even-Dar, E., S. M. Kakade, and Y. Mansour (2009) · 2009
Cited alongside, same era.
Robust regression with covariate filtering: Heavy tails and adversarial contamination
Pensia, A., V. Jog, and P.-L. Loh (2020) · 2009
Cited alongside, same era.
Online markov decision processes under bandit feedback
Neu, G., A. Antos, A. György, and C. Szepesvári (2010) · 2010
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Jin, Y., Z. Yang, and Z. Wang (2020) · 2012
Cited alongside, same era.
Safe policy improvement with baseline bootstrapping
Laroche, R., P. Trichelair, and R. T. Des Combes (2019) · 2019
Later among the works it cites.
Online stochastic shortest path with bandit feedback and unknown transition function
Rosenberg, A. and Y. Mansour (2019) · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., D. Schuurmans, and M. Norouzi (2020) · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q., Z. Yang, C. Jin, and Z. Wang (2020) · 2020
Later among the works it cites.
Notes on tabular methods
Jiang, N. (2020) · 2020
Later among the works it cites.
Learning adversarial markov decision processes with bandit feedback and unknown transition
Jin, C., T. Jin, H. Luo, S. Sra, and T. Yu (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch reinforcement learning
Lange, S., T. Gabel, and M. Riedmiller (2012) · 2012
Cited alongside, same era.
The adversarial stochastic shortest path problem with unknown transition probabilities
Neu, G., A. Gyorgy, and C. Szepesvári (2012) · 2012
Cited alongside, same era.
Online learning in episodic markovian decision processes by relative entropy policy search
Zimin, A. and G. Neu (2013) · 2013
Cited alongside, same era.
Robust estimators in high dimensions without the computational intractability
Diakonikolas, I., G. Kamath, D. Kane, J. Li, A. Moitra, and A. Stewart (2016) · 2016
Cited alongside, same era.
Agnostic estimation of mean and covariance
Lai, K. A., A. B. Rao, and S. Vempala (2016) · 2016
Cited alongside, same era.
Talking to bots: Symbiotic agency and the case of tay
Neff, G. (2016) · 2016
Cited alongside, same era.
Learning from untrusted data
Charikar, M., J. Steinhardt, and G. Valiant (2017) · 2017
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Z. Yang, Z. Wang, and M. I. Jordan (2020) · 2020
Later among the works it cites.
Improved corruption robust algorithms for episodic reinforcement learning
Chen, Y., S. S. Du, and K. Jamieson (2021) · 2021
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., B. Zhu, C. Ma, J. Jiao, and S. Russell (2021) · 2021
Closest in time.
Representation matters: Offline pretraining for sequential decision making
Yang, M. and O. Nachum (2021) · 2021
Closest in time.
Near-optimal offline reinforcement learning via double variance reduction
Yin, M., Y. Bai, and Y.-X. Wang (2021) · 2021
Closest in time.
Combo: Conservative offline model-based policy optimization
Yu, T., A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn (2021) · 2021
Closest in time.
Cautiously optimistic policy optimization and exploration with linear function approximation
Zanette, A., C.-A. Cheng, and A. Agarwal (2021) · 2021
Closest in time.
Robust policy gradient against strong data corruption
Zhang, X., Y. Chen, X. Zhu, and W. Sun (2021) · 2021
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., D. Meger, and D. Precup (2019) · 2062
Closest in time.