Fetching the paper…
Reading the bibliography…
In offline reinforcement learning (RL), an RL agent learns to solve a task using only a fixed dataset of previously collected data.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Earlier work this paper cites.
Cad2rl: Real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine · 2016
Earlier work this paper cites.
Improved learning of dynamics models for control
Arun Venkatraman, Roberto Capobianco, Lerrel Pinto, Martial Hebert, Daniele Nardi, and J Andrew Bagnell · 2016
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Imagination-augmented agents for deep reinforcement learning
Sébastien Racanière, Théophane Weber, David Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Earlier work this paper cites.
Dher: Hindsight experience replay for dynamic goals
Meng Fang, Cheng Zhou, Bei Shi, Boqing Gong, Jia Xu, and Tong Zhang · 2018
Earlier work this paper cites.
An environment for autonomous driving decision-making
Edouard Leurent · 2018
Earlier work this paper cites.
Run, skeleton, run: skeletal model in a physics-based simulation
Sergey Kolesnikov Mikhail Pavlov and Sergey M. Plis · 2018
Earlier work this paper cites.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Earlier work this paper cites.
On learning symmetric locomotion
Farzad Abdolhosseini, Hung Yu Ling, Zhaoming Xie, Xue Bin Peng, and Michiel Van de Panne · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, G. Tucker, and Sergey Levine · 2019
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation
L. Guan, Mudit Verma, Sihang Guo, Ruohan Zhang, and Subbarao Kambhampati · 2020
Earlier work this paper cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez, Konrad Zolna, Rishabh Agarwal, Josh S Merel, Daniel J Mankowitz, Cosmin Paduraru, et al · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Reinforcement learning with augmented data
Misha Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Sample-efficient reinforcement learning via counterfactual-based data augmentation
C. Lu, B. Huang, K. Wang, J. M. Hernández-Lobato, K. Zhang, and B. Schölkopf · 2020
Cited alongside, same era.
Automatic data augmentation for generalization in reinforcement learning
Roberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2021
Later among the works it cites.
Offline reinforcement learning with soft behavior regularization
Haoran Xu, Xianyuan Zhan, Jianxiong Li, and Honglei Yin · 2021
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Plas: Latent action space for offline reinforcement learning
Wenxuan Zhou, Sujay Bajracharya, and David Held · 2021
Later among the works it cites.
S2p: State-conditioned image synthesis for data augmentation in offline reinforcement learning
Daesol Cho, Dongseok Shim, and H. Jin Kim · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine · 2020
Cited alongside, same era.
Counterfactual data augmentation using locally factored dynamics
Silviu Pitis, Elliot Creager, and Animesh Garg · 2020
Cited alongside, same era.
Improving generalization in reinforcement learning with mixture regularization
Kaixin Wang, Bingyi Kang, Jie Shao, and Jiashi Feng · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Cited alongside, same era.
Emaq: Expected-max q-learning operator for simple yet effective offline and online rl
Seyed Kamyar Seyed Ghasemipour, Dale Schuurmans, and Shixiang Shane Gu · 2021
Cited alongside, same era.
Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation
L. Guan, Mudit Verma, Sihang Guo, Ruohan Zhang, and Subbarao Kambhampati · 2021
Cited alongside, same era.
Discovering faster matrix multiplication algorithms with reinforcement learning
Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J R Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, et al · 2022
Later among the works it cites.
Selective data augmentation for improving the performance of offline reinforcement learning
Jungwoo Han and Jinwhan Kim · 2022
Later among the works it cites.
Model-based trajectory stitching for improved offline reinforcement learning
Charles A Hepburn and Giovanni Montana · 2022
Later among the works it cites.
A swapping target q-value technique for data augmentation in offline reinforcement learning
Ho-Taek Joo, In-Chang Baek, and Kyung-Joong Kim · 2022
Later among the works it cites.
Should I run offline reinforcement learning or behavioral cloning?
Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine · 2022
Later among the works it cites.
Data augmentation for manipulation
Peter Mitrano and Dmitry Berenson · 2022
Later among the works it cites.
Mocoda: Model-based counterfactual data augmentation
Silviu Pitis, Elliot Creager, Ajay Mandlekar, and Animesh Garg · 2022
Later among the works it cites.
Bootstrapped transformer for offline reinforcement learning
Kerong Wang, Hanye Zhao, Xufang Luo, Kan Ren, Weinan Zhang, and Dongsheng Li · 2022
Later among the works it cites.
Koopman q-learning: Offline reinforcement learning via symmetries of dynamics
Matthias Weissenbacher, Samarth Sinha, Animesh Garg, and Kawahara Yoshinobu · 2022
Later among the works it cites.
Minimizing human assistance: Augmenting a single demonstration for deep reinforcement learning
Abraham George, Alison Bartsch, and Amir Barati Farimani · 2023
Closest in time.
Understanding when dynamics-invariant data augmentations benefit model-free reinforcement learning updates
Nicholas E. Corrado and Josiah P. Hanna · 2024
Closest in time.