Fetching the paper…
Reading the bibliography…
In offline RL, constraining the learned policy to remain close to the data is essential to prevent the policy from outputting out-of-distribution (OOD) actions with erroneously overestimated values.
Multi-task batch reinforcement learning with metric learning
Quan Vuong, Shuang Liu, Minghua Liu, Kamil Ciosek, Hao Su, and Henrik Iskov Christensen · 1909
Earlier work this paper cites.
Bail: Best-action imitation learning for batch deep reinforcement learning, 2019
Xinyue Chen, Zijian Zhou, Zheng Wang, Che Wang, Yanqiu Wu, and Keith Ross · 1910
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
A Mathematical Introduction to Robotic Manipulation
Richard M. Murray, S. Shankar Sastry, and Li Zexiang · 1994
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2004
Earlier work this paper cites.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2006
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Least squares generative adversarial networks, 2016
Xudong Mao, Qing Li, Haoran Xie, Raymond Y. K. Lau, Zhen Wang, and Stephen Paul Smolley · 2016
Earlier work this paper cites.
Amortised map inference for image super-resolution, 2016
Casper Kaae Sønderby, Jose Caballero, Lucas Theis, Wenzhe Shi, and Ferenc Huszár · 2016
Earlier work this paper cites.
Multi-modal imitation learning from unstructured demonstrations using generative adversarial nets
Karol Hausman, Yevgen Chebotar, Stefan Schaal, Gaurav S. Sukhatme, and Joseph J. Lim · 2017
Earlier work this paper cites.
Infogail: Interpretable imitation learning from visual demonstrations
Yunzhu Li, Jiaming Song, and Stefano Ermon · 2017
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration, 2018
Scott Fujimoto, David Meger, and Doina Precup · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2019
Cited alongside, same era.
Data-driven deep reinforcement learning
Aviral Kumar · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
Dmitry Kalashnikov, Jacob Varley, Yevgen Chebotar, Benjamin Swanson, Rico Jonschkowski, Chelsea Finn, Sergey Levine, and Karol Hausman · 2021
Later among the works it cites.
Zhihan Liu, Yufeng Zhang, Zuyue Fu, Zhuoran Yang, and Zhaoran Wang · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, and Martin Riedmiller · 2020
Cited alongside, same era.
Ziyu Wang, Alexander Novikov, Konrad Żołna, Jost Tobias Springenberg, Scott Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, et al · 2020
Cited alongside, same era.
Chenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhihong Deng, Animesh Garg, Peng Liu, and Zhaoran Wang · 2022
Closest in time.
Latent-variable advantage-weighted policy optimization for offline rl
Xi Chen, Ali Ghadirzadeh, Tianhe Yu, Yuan Gao, Jianhao Wang, Wenzhe Li, Bin Liang, Chelsea Finn, and Chongjie Zhang · 2022
Closest in time.
Density estimation for conservative q-learning, 2022
Paul Daoudi, Merwan Barlier, Ludovic Dos Santos, and Aladin Virmaux · 2022
Closest in time.
Generalized decision transformer for offline hindsight information matching
Hiroki Furuta, Yutaka Matsuo, and Shixiang Shane Gu · 2022
Closest in time.
GPT-critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems
Youngsoo Jang, Jongmin Lee, and Kee-Eung Kim · 2022
Closest in time.
When should we prefer offline reinforcement learning over behavioral cloning?, 2022
Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine · 2022
Closest in time.
Robust imitation learning from corrupted demonstrations, 2022
Liu Liu, Ziyang Tang, Lanqing Li, and Dijun Luo · 2022
Closest in time.
Offline pre-trained multi-agent decision transformer, 2022
Linghui Meng, Muning Wen, Yaodong Yang, chenyang le, Xi yun Li, Haifeng Zhang, Ying Wen, Weinan Zhang, Jun Wang, and Bo XU · 2022
Closest in time.
Context-aware language modeling for goal-oriented dialogue systems, 2022
Charlie Snell, Mengjiao Yang, Justin Fu, Yi Su, and Sergey Levine · 2022
Closest in time.