Fetching the paper…
Reading the bibliography…
Recent Offline Reinforcement Learning methods have succeeded in learning high-performance policies from fixed datasets of experience.
“SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards”, 2019
Siddharth Reddy, Anca. Dragan and Sergey Levine · 1905
Earlier work this paper cites.
“An Optimistic Perspective on Offline Reinforcement Learning”, 2020
Rishabh Agarwal, Dale Schuurmans and Mohammad Norouzi · 1907
Earlier work this paper cites.
“BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning”, 2020
Xinyue Chen et al · 1910
Earlier work this paper cites.
“Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning”, 2019
Xue Peng, Aviral Kumar, Grace Zhang and Sergey Levine · 1910
Earlier work this paper cites.
“Maxmin Q-learning: Controlling the Estimation Bias of Q-learning”, 2020
Qingfeng Lan, Yangchen Pan, Alona Fyshe and Martha White · 2002
Earlier work this paper cites.
“DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction”, 2020
Aviral Kumar, Abhishek Gupta and Sergey Levine · 2003
Earlier work this paper cites.
“Reinforcement learning as classification: Leveraging modern classifiers”
Michail Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
“D4RL: Datasets for Deep Data-Driven Reinforcement Learning”, 2021
Justin Fu et al · 2004
Earlier work this paper cites.
“Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO”, 2020
Logan Engstrom et al · 2005
Earlier work this paper cites.
Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin and Dmitry Vetrov · 2005
Earlier work this paper cites.
“Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems”, 2020
Sergey Levine, Aviral Kumar, George Tucker and Justin Fu · 2005
Earlier work this paper cites.
“MOPO: Model-based Offline Policy Optimization”, 2020
Tianhe Yu et al · 2005
Earlier work this paper cites.
“RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning”, 2021
Caglar Gulcehre et al · 2006
Earlier work this paper cites.
“Conservative Q-Learning for Offline Reinforcement Learning”, 2020
Aviral Kumar, Aurick Zhou, George Tucker and Sergey Levine · 2006
Earlier work this paper cites.
“Accelerating Online Reinforcement Learning with Offline Datasets”, 2020
Ashvin Nair, Murtaza Dalal, Abhishek Gupta and Sergey Levine · 2006
Earlier work this paper cites.
“Critic Regularized Regression”, 2020
Ziyu Wang et al · 2006
Earlier work this paper cites.
“SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement Learning”, 2020
Kimin Lee, Michael Laskin, Aravind Srinivas and Pieter Abbeel · 2007
Earlier work this paper cites.
“Offline Learning from Demonstrations and Unlabeled Experience”, 2020
Konrad Zolna et al · 2011
Cited alongside, same era.
“Continuous control with deep reinforcement learning”, 2015
Timothy. Lillicrap et al · 2015
Cited alongside, same era.
“ImageNet Large Scale Visual Recognition Challenge”, 2015
Olga Russakovsky et al · 2015
Cited alongside, same era.
Greg Brockman et al · 2016
Cited alongside, same era.
“Learning values across many orders of magnitude”, 2016
Hado van Hasselt et al · 2016
“An Algorithmic Perspective on Imitation Learning”
Takayuki Osa et al · 2018
Later among the works it cites.
“Exponentially Weighted Imitation Learning for Batched Historical Data”
Qing Wang et al · 2018
Later among the works it cites.
“Learning Self-Imitating Diverse Policies”, 2019
Tanmay Gangwani, Qiang Liu and Jian Peng · 2019
Later among the works it cites.
“Deep Reinforcement Learning that Matters”, 2019
Peter Henderson et al · 2019
Later among the works it cites.
“Stabilizing off-policy q-learning via bootstrapping error reduction”
Aviral Kumar, Justin Fu, George Tucker and Sergey Levine · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Prioritized Experience Replay”, 2016
Tom Schaul, John Quan, Ioannis Antonoglou and David Silver · 2016
Cited alongside, same era.
“Averaged-DQN: Variance Reduction and Stabilization for Deep Reinforcement Learning”, 2017
Oron Anschel, Nir Baram and Nahum Shimkin · 2017
Cited alongside, same era.
“Improving Stochastic Policy Gradients in Continuous Control with Deep Reinforcement Learning using the Beta Distribution”
Po-Wei Chou, Daniel Maturana and Sebastian Scherer · 2017
Cited alongside, same era.
“OpenAI Baselines”
Prafulla Dhariwal et al · 2017
Cited alongside, same era.
“On Calibration of Modern Neural Networks”, 2017
Chuan Guo, Geoff Pleiss, Yu Sun and Kilian. Weinberger · 2017
Cited alongside, same era.
“Spinning Up in Deep Reinforcement Learning”, 2018
Joshua Achiam · 2018
Cited alongside, same era.
“Addressing function approximation error in actor-critic methods”
Scott Fujimoto, Herke Hoof and David Meger · 2018
Cited alongside, same era.
“Deep Reinforcement Learning in the Real World” Workshop on New Directions in Reinforcement Learning and Control, 2019
Sergey Levine · 2019
Later among the works it cites.
“Imitation learning from imperfect demonstration”
Yueh-Hua Wu et al · 2019
Later among the works it cites.
“Self-Imitation Advantage Learning”
Johan Ferret, Olivier Pietquin and Matthieu Geist · 2020
Later among the works it cites.
“ deep_control
Jake Grigsby · 2020
Later among the works it cites.
“Variational Imitation Learning with Diverse-quality Demonstrations”
Voot Tangkaratt, Bo Han, Mohammad Khan and Masashi Sugiyama · 2020
Later among the works it cites.
“Uncertainty Weighted Offline Reinforcement Learning”, 2020
Yue Wu et al · 2020
Later among the works it cites.
“Soft Actor-Critic (SAC) implementation in PyTorch”
Denis Yarats and Ilya Kostrikov · 2020
Later among the works it cites.
“Decision Transformer: Reinforcement Learning via Sequence Modeling”, 2021
Lili Chen et al · 2021
Closest in time.
“Randomized Ensembled Double Q-Learning: Learning Fast Without a Model”, 2021
Xinyue Chen, Che Wang, Zijian Zhou and Keith Ross · 2021
Closest in time.
“NeoRL: A Near Real-World Benchmark for Offline Reinforcement Learning”, 2021
Rongjun Qin et al · 2021
Closest in time.
Mamshad Rizve, Kevin Duarte, Yogesh Rawat and Mubarak Shah · 2021
Closest in time.