Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) presents a promising approach for learning reinforced policies from offline datasets without the need for costly or unsafe interactions with the environment.
Statistical estimates and transformed beta-variables
Gunnar Blom · 1958
Earlier work this paper cites.
Measures of multivariate skewness and kurtosis with applications
Kanti V Mardia · 1970
Earlier work this paper cites.
Robust regression: asymptotics, conjectures and monte carlo
Peter J Huber · 1973
Earlier work this paper cites.
Algorithm as 177: Expected normal order statistics (exact and approximate)
JP Royston · 1982
Earlier work this paper cites.
Robust estimation of a location parameter
Peter J Huber · 1992
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Robustness in markov decision problems with uncertain transition matrices
Arnab Nilim and Laurent Ghaoui · 2003
Earlier work this paper cites.
Robust statistics , volume 523
Peter J Huber · 2004
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Robust imitation learning from noisy demonstrations
Voot Tangkaratt, Nontawat Charoenphakdee, and Masashi Sugiyama · 2010
Earlier work this paper cites.
Bandits with heavy tail
Sébastien Bubeck, Nicolo Cesa-Bianchi, and Gábor Lugosi · 2013
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Earlier work this paper cites.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc Bellemare, and Rémi Munos · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Fast bellman updates for robust mdps
Chin Pang Ho, Marek Petrik, and Wolfram Wiesemann · 2018
Earlier work this paper cites.
Almost optimal algorithms for linear stochastic bandits with heavy-tailed payoffs
Han Shao, Xiaotian Yu, Irwin King, and Michael R Lyu · 2018
Earlier work this paper cites.
Exponentially weighted imitation learning for batched historical data
Qing Wang, Jiechao Xiong, Lei Han, Han Liu, Tong Zhang, et al · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Earlier work this paper cites.
Robust reinforcement learning for continuous control with model misspecification
Daniel J Mankowitz, Nir Levine, Rae Jeong, Yuanyuan Shi, Jackie Kay, Abbas Abdolmaleki, Jost Tobias Springenberg, Timothy Mann, Todd Hester, and Martin Riedmiller · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Earlier work this paper cites.
Action robust reinforcement learning and applications in continuous control
Chen Tessler, Yonathan Efroni, and Shie Mannor · 2019
Earlier work this paper cites.
Imitation learning from imperfect demonstration
Yueh-Hua Wu, Nontawat Charoenphakdee, Han Bao, Voot Tangkaratt, and Masashi Sugiyama · 2019
Earlier work this paper cites.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Earlier work this paper cites.
Lanqing Li, Rui Yang, and Dijun Luo · 2020
Earlier work this paper cites.
A general framework for uncertainty estimation in deep learning
Antonio Loquercio, Mattia Segu, and Davide Scaramuzza · 2020
Earlier work this paper cites.
Awac: Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine · 2020
Cited alongside, same era.
Behavioral cloning from noisy demonstrations
Fumihiro Sasaki and Ryota Yamashina · 2020
Cited alongside, same era.
Adaptive huber regression
Qiang Sun, Wen-Xin Zhou, and Jianqing Fan · 2020
Cited alongside, same era.
Reinforcement learning with perturbed rewards
Jingkang Wang, Yang Liu, and Bo Li · 2020
Cited alongside, same era.
Error bounds of imitating policies and environments
Tian Xu, Ziniu Li, and Yang Yu · 2020
Cited alongside, same era.
Nearly optimal regret for stochastic linear bandits with heavy-tailed payoffs
Provable sim-to-real transfer in continuous domain with partial observations
Jiachen Hu, Han Zhong, Chi Jin, and Liwei Wang · 2022
Later among the works it cites.
Adaptive best-of-both-worlds algorithm for heavy-tailed multi-armed bandits
Jiatai Huang, Yan Dai, and Longbo Huang · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
Michael Janner, Yilun Du, Joshua Tenenbaum, and Sergey Levine · 2022
Later among the works it cites.
Settling the sample complexity of model-based offline reinforcement learning
Gen Li, Laixi Shi, Yuxin Chen, Yuejie Chi, and Yuting Wei · 2022
Later among the works it cites.
Robust imitation learning from corrupted demonstrations
Liu Liu, Ziyang Tang, Lanqing Li, and Dijun Luo · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bo Xue, Guanghui Wang, Yimu Wang, and Lijun Zhang · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Cited alongside, same era.
Robust deep reinforcement learning against adversarial perturbations on state observations
Huan Zhang, Hongge Chen, Chaowei Xiao, Bo Li, Mingyan Liu, Duane Boning, and Cho-Jui Hsieh · 2020
Cited alongside, same era.
No-regret reinforcement learning with heavy-tailed rewards
Vincent Zhuang and Yanan Sui · 2020
Cited alongside, same era.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song · 2021
Cited alongside, same era.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Johan Samir Obando Ceron and Pablo Samuel Castro · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
Later among the works it cites.
Robust reinforcement learning: A review of foundations and recent advances
Janosch Moos, Kay Hansel, Hany Abdulsamad, Svenja Stark, Debora Clever, and Jan Peters · 2022
Later among the works it cites.
Robust reinforcement learning using offline data
Kishan Panaganti, Zaiyan Xu, Dileep Kalathil, and Mohammad Ghavamzadeh · 2022
Later among the works it cites.
Robust losses for learning value functions
Andrew Patterson, Victor Liao, and Martha White · 2022
Later among the works it cites.
Laixi Shi and Yuejie Chi · 2022
Later among the works it cites.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Laixi Shi, Gen Li, Yuting Wei, Yuxin Chen, and Yuejie Chi · 2022
Later among the works it cites.
CORL: Research-oriented deep offline reinforcement learning library
Denis Tarasov, Alexander Nikulin, Dmitry Akimov, Vladislav Kurenkov, and Sergey Kolesnikov · 2022
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Zhendong Wang, Jonathan J Hunt, and Mingyuan Zhou · 2022
Later among the works it cites.
A model selection approach for corruption robust reinforcement learning
Chen-Yu Wei, Christoph Dann, and Julian Zimmert · 2022
Later among the works it cites.
Copa: Certifying robust policies for offline reinforcement learning against poisoning attacks
Fan Wu, Linyi Li, Chejian Xu, Huan Zhang, Bhavya Kailkhura, Krishnaram Kenthapadi, Ding Zhao, and Bo Li · 2022
Later among the works it cites.
Wei Xiong, Han Zhong, Chengshuai Shi, Cong Shen, Liwei Wang, and Tong Zhang · 2022
Later among the works it cites.
Corruption-robust offline reinforcement learning
Xuezhou Zhang, Yiding Chen, Xiaojin Zhu, and Wen Sun · 2022
Later among the works it cites.
Pessimistic minimax value iteration: Provably efficient equilibrium learning from offline datasets
Han Zhong, Wei Xiong, Jiyuan Tan, Liwei Wang, Tong Zhang, Zhaoran Wang, and Zhuoran Yang · 2022
Later among the works it cites.
Jose Blanchet, Miao Lu, Tong Zhang, and Han Zhong · 2023
Closest in time.
Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions
Yevgen Chebotar, Quan Vuong, Alex Irpan, Karol Hausman, Fei Xia, Yao Lu, Aviral Kumar, Tianhe Yu, Alexander Herzog, Karl Pertsch, et al · 2023
Closest in time.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Philippe Hansen-Estruch, Ilya Kostrikov, Michael Janner, Jakub Grudzien Kuba, and Sergey Levine · 2023
Closest in time.
Jiayi Huang, Han Zhong, Liwei Wang, and Lin F Yang · 2023
Closest in time.
Heavy-tailed linear bandit with huber regression
Minhyun Kang and Gi-Soo Kim · 2023
Closest in time.
Survival instinct in offline reinforcement learning
Anqi Li, Dipendra Misra, Andrey Kolobov, and Ching-An Cheng · 2023
Closest in time.
Anti-exploration by random network distillation
Alexander Nikulin, Vladislav Kurenkov, Denis Tarasov, and Sergey Kolesnikov · 2023
Closest in time.
Hao Sun · 2023
Closest in time.
Towards robust offline-to-online reinforcement learning via uncertainty and smoothness
Xiaoyu Wen, Xudong Yu, Rui Yang, Chenjia Bai, and Zhen Wang · 2023
Closest in time.
Offline rl with no ood actions: In-sample learning via implicit value regularization
Haoran Xu, Li Jiang, Jianxiong Li, Zhuoran Yang, Zhaoran Wang, Victor Wai Kin Chan, and Xianyuan Zhan · 2023
Closest in time.
Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline rl
Taku Yamagata, Ahmed Khalil, and Raul Santos-Rodriguez · 2023
Closest in time.
What is essential for unseen goal generalization of offline goal-conditioned rl?
Rui Yang, Lin Yong, Xiaoteng Ma, Hao Hu, Chongjie Zhang, and Tong Zhang · 2023
Closest in time.
Accountability in offline reinforcement learning: Explaining decisions with a corpus of examples
Hao Sun, Alihan Hüyük, Daniel Jarrett, and Mihaela van der Schaar · 2024
Closest in time.
Secrets of rlhf in large language models part ii: Reward modeling
Binghai Wang, Rui Zheng, Lu Chen, Yan Liu, Shihan Dou, Caishuang Huang, Wei Shen, Senjie Jin, Enyu Zhou, Chenyu Shi, et al · 2024
Closest in time.