Fetching the paper…
Reading the bibliography…
We present a novel observation about the behavior of offline reinforcement learning (RL) algorithms: on many benchmark datasets, offline RL can produce well-performing and safe policies even when trained with "wrong" reward labels, such as those that are zero everywhere or are negatives of the true rewards.
CRC press, 1999
E. Altman, Constrained Markov decision processes · 1999
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in International Conference on Machine Learning
1999
Earlier work this paper cites.
R. Munos, “Error bounds for approximate policy iteration,” in International Conference on Machine Learning
2003
Earlier work this paper cites.
2005
Earlier work this paper cites.
A. Hans, D. Schneegaß, A. M. Schäfer, and S. Udluft, “Safe exploration for reinforcement learning.,” in ESANN
2008
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey, et al
2008
Earlier work this paper cites.
M. Dudík, J. Langford, and L. Li, “Doubly robust policy evaluation and learning,” in Proceedings of the 28th International Conference on International Conference on Machine Learning
2011
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Hadfield-Menell, S. Milli, P. Abbeel, S. J. Russell, and A. Dragan, “Inverse reward design,” Advances in neural information processing systems
2017
Earlier work this paper cites.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in International conference on machine learning
2017
Earlier work this paper cites.
B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei, “Reward learning from human preferences and demonstrations in atari,” Advances in neural information processing systems
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
MIT press, 2018
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction · 2018
Earlier work this paper cites.
H. Le, C. Voloshin, and Y. Yue, “Batch policy learning under constraints,” in International Conference on Machine Learning
2019
Earlier work this paper cites.
J. Chen and N. Jiang, “Information-theoretic considerations in batch reinforcement learning,” in International Conference on Machine Learning
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
W. Yu, C. K. Liu, and G. Turk, “Policy transfer with strategy optimization,” in International Conference on Learning Representations
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” Advances in Neural Information Processing Systems
2020
Earlier work this paper cites.
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine, “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in Conference on robot learning
2020
Earlier work this paper cites.
Y. Liu, A. Swaminathan, A. Agarwal, and E. Brunskill, “Provably good batch off-policy reinforcement learning without great exploration,” Advances in neural information processing systems
2020
Earlier work this paper cites.
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma, “Mopo: Model-based offline policy optimization,” Advances in Neural Information Processing Systems
2020
Earlier work this paper cites.
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims, “Morel: Model-based offline reinforcement learning,” Advances in neural information processing systems
2020
Earlier work this paper cites.
A. Wachi and Y. Sui, “Safe reinforcement learning in constrained markov decision processes,” in International Conference on Machine Learning
2020
Earlier work this paper cites.
A. Singh, A. Yu, J. Yang, J. Zhang, A. Kumar, and S. Levine, “Cog: Connecting new skills to past experience with offline reinforcement learning,” in Conference on Robot Learning
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
Z. Zhu, K. Lin, B. Dai, and J. Zhou, “Off-policy imitation learning from observations,” Advances in Neural Information Processing Systems
2020
Cited alongside, same era.
B. Dai, O. Nachum, Y. Chow, L. Li, C. Szepesvári, and D. Schuurmans, “Coindice: Off-policy confidence interval estimation,” Advances in neural information processing systems
2020
Cited alongside, same era.
T. Xie, C.-A. Cheng, N. Jiang, P. Mineiro, and A. Agarwal, “Bellman-consistent pessimism for offline reinforcement learning,” Advances in neural information processing systems
2021
Cited alongside, same era.
2022
Later among the works it cites.
M. Rigter, B. Lacerda, and N. Hawes, “Rambo-RL: Robust adversarial model-based offline reinforcement learning,” Advances in neural information processing systems
2022
Later among the works it cites.
W. Zhan, B. Huang, A. Huang, N. Jiang, and J. Lee, “Offline reinforcement learning with realizability and single-policy concentrability,” in Conference on Learning Theory
2022
Later among the works it cites.
D. J. Foster, A. Krishnamurthy, D. Simchi-Levi, and Y. Xu, “Offline reinforcement learning: Fundamental barriers for value function approximation,” in Conference on Learning Theory
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch, “Decision transformer: Reinforcement learning via sequence modeling,” Advances in neural information processing systems
2021
Cited alongside, same era.
Y. Jin, Z. Yang, and Z. Wang, “Is pessimism provably efficient for offline rl?,” in International Conference on Machine Learning
2021
Cited alongside, same era.
J. Chang, M. Uehara, D. Sreenivas, R. Kidambi, and W. Sun, “Mitigating covariate shift in imitation learning via offline data with partial coverage,” Advances in Neural Information Processing Systems
2021
Cited alongside, same era.
Z. Zhou, Z. Zhou, Q. Bai, L. Qiu, J. Blanchet, and P. Glynn, “Finite-sample regret bound for distributionally robust offline tabular reinforcement learning,” in International Conference on Artificial Intelligence and Statistics
2021
Cited alongside, same era.
S. Fujimoto and S. S. Gu, “A minimalist approach to offline reinforcement learning,” Advances in neural information processing systems
2021
Cited alongside, same era.
P. Rashidinejad, B. Zhu, C. Ma, J. Jiao, and S. Russell, “Bridging offline reinforcement learning and imitation learning: A tale of pessimism,” Advances in Neural Information Processing Systems
2021
Cited alongside, same era.
T. Yu, A. Kumar, R. Rafailov, A. Rajeswaran, S. Levine, and C. Finn, “Combo: Conservative offline model-based policy optimization,” Advances in neural information processing systems
2021
Cited alongside, same era.
S. Paternain, M. Calvo-Fullana, L. F. Chamon, and A. Ribeiro, “Safe policies for reinforcement learning via primal-dual methods,” IEEE Transactions on Automatic Control
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Kumar, J. Hong, A. Singh, and S. Levine, “Should i run offline reinforcement learning or behavioral cloning?,” in International Conference on Learning Representations
2022
Later among the works it cites.
T. Yu, A. Kumar, Y. Chebotar, K. Hausman, C. Finn, and S. Levine, “How to leverage unlabeled data in offline reinforcement learning,” in International Conference on Machine Learning
2022
Later among the works it cites.
H. Xu, X. Zhan, H. Yin, and H. Qin, “Discriminator-weighted offline imitation learning from suboptimal demonstrations,” in International Conference on Machine Learning
2022
Later among the works it cites.
R. Yang, C. Bai, X. Ma, Z. Wang, C. Zhang, and L. Han, “Rorl: Robust offline reinforcement learning via conservative smoothing,” in Advances in Neural Information Processing Systems
2022
Later among the works it cites.
2022
Later among the works it cites.
N. Polosky, B. C. Da Silva, M. Fiterau, and J. Jagannath, “Constrained offline policy optimization,” in International Conference on Machine Learning
2022
Later among the works it cites.
S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin, “Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,” in Conference on Robot Learning
2022
Later among the works it cites.
F. Wu, L. Li, H. Zhang, B. Kailkhura, K. Kenthapadi, D. Zhao, and B. Li, “Copa: Certifying robust policies for offline reinforcement learning against poisoning attacks,” in International Conference on Learning Representations
2023
Closest in time.
D. Shin, A. D. Dragan, and D. S. Brown, “Benchmarks and algorithms for offline preference-based reward learning,” Transactions on Machine Learning Research
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Z. Liu, Z. Guo, H. Lin, Y. Yao, J. Zhu, Z. Cen, H. Hu, W. Yu, T. Zhang, J. Tan, et al
2023
Closest in time.
H. Hu, Y. Yang, Q. Zhao, and C. Zhang, “The provable benefits of unsupervised data sharing for offline reinforcement learning,” in International Conference on Learning Representations
2023
Closest in time.
2023
Closest in time.
J. Li, X. Hu, H. Xu, J. Liu, X. Zhan, Q.-S. Jia, and Y.-Q. Zhang, “Mind the gap: Offline policy optimization for imperfect rewards,” in International Conference on Learning Representations
2023
Closest in time.
S. Yue, G. Wang, W. Shao, Z. Zhang, S. Lin, J. Ren, and J. Zhang, “CLARE: Conservative model-based reward learning for offline inverse reinforcement learning,” in International Conference on Learning Representations
2023
Closest in time.
2023
Closest in time.
Y. Chen, X. Zhang, K. Zhang, M. Wang, and X. Zhu, “Byzantine-robust online and offline distributed reinforcement learning,” in International Conference on Artificial Intelligence and Statistics
2023
Closest in time.
2023
Closest in time.
W. Xu, Y. Ma, K. Xu, H. Bastani, and O. Bastani, “Uniformly conservative exploration in reinforcement learning,” in International Conference on Artificial Intelligence and Statistics
2023
Closest in time.