Fetching the paper…
Reading the bibliography…
Progress in offline reinforcement learning (RL) has been impeded by ambiguous problem definitions and entangled algorithmic designs, resulting in inconsistent implementations, insufficient ablations, and unfair evaluations.
Striving for simplicity in off-policy deep reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 1907
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems, November 2020
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2005
Earlier work this paper cites.
MOPO: Model-based Offline Policy Optimization, November 2020
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2005
Earlier work this paper cites.
MOReL : Model-Based Offline Reinforcement Learning, March 2021
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2005
Earlier work this paper cites.
Conservative Q-Learning for Offline Reinforcement Learning, August 2020
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2006
Earlier work this paper cites.
Hyperparameter Selection for Offline Reinforcement Learning, July 2020
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Caglar Gulcehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2007
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Spinning Up in Deep Reinforcement Learning, 2018
Joshua Achiam · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Addressing Function Approximation Error in Actor-Critic Methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
Deployment-efficient reinforcement learning via model-based offline optimization
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, and Shixiang Gu · 2020
Cited alongside, same era.
Planning to explore via self-supervised world models
Showing Your Offline Reinforcement Learning Work: Online Evaluation Budget Matters, June 2022
Vladislav Kurenkov and Sergey Kolesnikov · 2022
Later among the works it cites.
COMBO: Conservative Offline Model-Based Policy Optimization, January 2022
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn · 2022
Later among the works it cites.
CORL: Research-oriented deep offline reinforcement learning library
Denis Tarasov, Alexander Nikulin, Dmitry Akimov, Vladislav Kurenkov, and Sergey Kolesnikov · 2022
Later among the works it cites.
No more pesky hyperparameters: Offline hyperparameter tuning for rl
Han Wang, Archit Sakhadeo, Adam White, James Bell, Vincent Liu, Xutong Zhao, Puer Liu, Tadashi Kozuno, Alona Fyshe, and Martha White · 2022
Later among the works it cites.
Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
d3rlpy: An offline deep reinforcement library
Takuma Seno · 2020
Cited alongside, same era.
A Minimalist Approach to Offline Reinforcement Learning, December 2021
Scott Fujimoto and Shixiang Shane Gu · 2021
Cited alongside, same era.
Active offline policy selection
Ksenia Konyushova, Yutian Chen, Thomas Paine, Caglar Gulcehre, Cosmin Paduraru, Daniel J Mankowitz, Misha Denil, and Nando de Freitas · 2021
Cited alongside, same era.
Uncertainty-Based Offline Reinforcement Learning with Diversified Q-Ensemble, October 2021
Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song · 2021
Cited alongside, same era.
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Cited alongside, same era.
Muesli: Combining improvements in policy optimization
Matteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez, Simon Schmitt, Laurent Sifre, Theophane Weber, David Silver, and Hado Van Hasselt · 2021
Cited alongside, same era.
Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, Jeff Braga, Dipam Chakraborty, Kinal Mehta, and JoÃĢo GM AraÚjo · 2022
Later among the works it cites.
Revisiting the Minimalist Approach to Offline Reinforcement Learning, October 2023
Denis Tarasov, Vladislav Kurenkov, Alexander Nikulin, and Sergey Kolesnikov · 2023
Later among the works it cites.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Rafael Figueiredo Prudencio, Marcos ROA Maximo, and Esther Luna Colombini · 2023
Later among the works it cites.
A strong baseline for batch imitation learning
Matthew Smith, Lucas Maystre, Zhenwen Dai, and Kamil Ciosek · 2023
Later among the works it cites.
Offlinerl-kit: An elegant pytorch offline reinforcement learning library
Yihao Sun · 2023
Later among the works it cites.
Dual rl: Unification and new methods for reinforcement and imitation learning
Harshit Sikchi, Qinqing Zheng, Amy Zhang, and Scott Niekum · 2023
Later among the works it cites.
Policy-guided diffusion, 2024
Matthew Thomas Jackson, Michael Tryfan Matthews, Cong Lu, Benjamin Ellis, Shimon Whiteson, and Jakob Foerster · 2024
Later among the works it cites.
Jax-corl: Clean sigle-file implementations of offline rl algorithms in jax
Soichiro Nishimori · 2024
Later among the works it cites.
rejax, 2024
Jarek Liesen, Chris Lu, and Robert Lange · 2024
Later among the works it cites.
The Generalization Gap in Offline Reinforcement Learning, March 2024
Ishita Mediratta, Qingfei You, Minqi Jiang, and Roberta Raileanu · 2024
Later among the works it cites.