Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) refers to the problem of learning policies from a static dataset of environment interactions.
Some asymptotic theory for the bootstrap
Peter J Bickel and David A Freedman · 1981
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2005
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
Sascha Lange and Martin Riedmiller · 2010
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin A. Riedmiller · 2012
Earlier work this paper cites.
The movielens datasets: History and context
F Maxwell Harper and Joseph A Konstan · 2015
Earlier work this paper cites.
Batch learning from logged bandit feedback through counterfactual risk minimization
Adith Swaminathan and Thorsten Joachims · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
Chelsea Finn, Xin Yu Tan, Yan Duan, Trevor Darrell, Sergey Levine, and Pieter Abbeel · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Earlier work this paper cites.
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2017
Earlier work this paper cites.
Deep predictive policy training using reinforcement learning
Ali Ghadirzadeh, Atsuto Maki, Danica Kragic, and Mårten Björkman · 2017
Earlier work this paper cites.
Stochastic variational video prediction
Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H. Campbell, and Sergey Levine · 2018
Earlier work this paper cites.
Robust locally-linear controllable embedding
Ershad Banijamali, Rui Shu, Hung Bui, Ali Ghodsi, et al · 2018
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Frederik Ebert, Chelsea Finn, Sudeep Dasari, Annie Xie, Alex Lee, and Sergey Levine · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2018
Earlier work this paper cites.
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Earlier work this paper cites.
Visual reinforcement learning with imagined goals
Ashvin V Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Earlier work this paper cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2018
Earlier work this paper cites.
Multiple interactions made easy (mime): Large scale demonstrations data for imitation
Pratyusha Sharma, Lekha Mohan, Lerrel Pinto, and Abhinav Gupta · 2018
Cited alongside, same era.
Unsupervised predictive memory in a goal-directed agent
Greg Wayne, Chia-Chun Hung, David Amos, Mehdi Mirza, Arun Ahuja1, Agnieszka, Grabska-Barwinska, Jack Rae1, Piotr Mirowski, Joel Z. Leibo, Adam Santoro, Mevlana Gemici, Malcolm Reynolds, Tim Harley, Josh Abramson, Shakir Mohamed, Danilo Rezende, David Saxton, Adam Cain, Chloe Hillier, David Silver, Koray Kavukcuoglu, Matt Botvinick, Demis Hassabis, and Timothy Lillicrap · 2018
Cited alongside, same era.
Striving for simplicity in off-policy deep reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2019
Cited alongside, same era.
Robel: Robotics benchmarks for learning with low-cost robots, 2019
Michael Ahn, Henry Zhu, Kristian Hartikainen, Hugo Ponte, Abhishek Gupta, Sergey Levine, and Vikash Kumar · 2019
Cited alongside, same era.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Closest in time.
Batch exploration with examples for scalable robotic reinforcement learning, 2020
Annie S. Chen, HyunJi Nam, Suraj Nair, and Chelsea Finn · 2020
Closest in time.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Closest in time.
Rl unplugged: Benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gómez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel Mankowitz, Cosmin Paduraru, et al · 2020
Closest in time.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robonet: Large-scale multi-robot learning
Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn · 2019
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2019
Cited alongside, same era.
Contrastive learning of structured world models
Thomas Kipf, Elise van der Pol, and Max Welling · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Cited alongside, same era.
Closest in time.
Variational recurrent models for solving partially observable control tasks
Dongqi Han, Kenji Doya, and Jun Tani · 2020
Closest in time.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Closest in time.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2020
Closest in time.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Closest in time.
Reinforcement learning with augmented data
Michael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas · 2020
Closest in time.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Alex X. Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Closest in time.
Provably good batch reinforcement learning without great exploration
Yao Liu, A. Swaminathan, A. Agarwal, and Emma Brunskill · 2020
Closest in time.
Deployment-efficient reinforcement learning via model-based offline optimization
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, and Shixiang Gu · 2020
Closest in time.
A game theoretic framework for model based reinforcement learning
Aravind Rajeswaran, Igor Mordatch, and Vikash Kumar · 2020
Closest in time.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, and Martin Riedmiller · 2020
Closest in time.
The surprising effectiveness of linear models for visual foresight in object pile manipulation
HJ Suh and Russ Tedrake · 2020
Closest in time.
Overcoming model bias for robust offline deep reinforcement learning
Phillip Swazinna, Steffen Udluft, and Thomas Runkler · 2020
Closest in time.
dm control: Software and tasks for continuous control, 2020
Yuval Tassa, Saran Tunyasuvunakool, Alistair Muldal, Yotam Doron, Siqi Liu, Steven Bohez, Josh Merel, Tom Erez, Timothy Lillicrap, and Nicolas Heess · 2020
Closest in time.
Gendice: Generalized offline estimation of stationary values, 2020
Ruiyi Zhang, Bo Dai, Li Lihong, and Dale Schuurmans · 2020
Closest in time.