Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) enables learning control policies by utilizing only prior experience, without any online interaction.
Eligibility traces for off-policy policy evaluation
D. Precup · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
D. Precup, R. S. Sutton, and S. Dasgupta · 2001
Earlier work this paper cites.
Error bounds for approximate policy iteration
R. Munos · 2003
Earlier work this paper cites.
Off-policy evaluation in Markov decision processes
C. Paduraru · 2012
Earlier work this paper cites.
Batch reinforcement learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2015
Earlier work this paper cites.
Learning robot in-hand manipulation with tactile features
H. van Hoof, T. Hermans, G. Neumann, and J. Peters · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Learning dexterous manipulation policies from experience and imitation
V. Kumar, A. Gupta, E. Todorov, and S. Levine · 2016
Earlier work this paper cites.
Deep variational information bottleneck
A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Visual closed-loop control for pouring liquids
C. Schenck and D. Fox · 2017
Earlier work this paper cites.
Collective robot reinforcement learning with distributed asynchronous guided policy search
A. Yahya, A. Li, M. Kalakrishnan, Y. Chebotar, and S. Levine · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Earlier work this paper cites.
Learning synergies between pushing and grasping with self-supervised deep reinforcement learning
A. Zeng, S. Song, S. Welker, J. Lee, A. Rodriguez, and T. Funkhouser · 2018
Earlier work this paper cites.
Learning dexterous in-hand manipulation
OpenAI · 2018
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2018
Cited alongside, same era.
Sim-to-real reinforcement learning for deformable object manipulation
J. Matas, S. James, and A. J. Davison · 2018
Cited alongside, same era.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
F. Ebert, C. Finn, S. Dasari, A. Xie, A. Lee, and S. Levine · 2018
Cited alongside, same era.
Interpretable latent spaces for learning from demonstration
Y. Hristov, A. Lascarides, and S. Ramamoorthy · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2018
Reinforcement learning via fenchel-rockafellar duality
O. Nachum and B. Dai · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Efficient adaptation for end-to-end vision-based robotic manipulation
R. Julian, B. Swanson, G. S. Sukhatme, S. Levine, C. Finn, and K. Hausman · 2020
Later among the works it cites.
Model-based visual planning with self-supervised functional distances
S. Tian, S. Nair, F. Ebert, S. Dasari, B. Eysenbach, C. Finn, and S. Levine · 2020
Later among the works it cites.
Visual imitation made easy
S. Young, D. Gandhi, S. Tulsiani, A. Gupta, P. Abbeel, and L. Pinto · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Van Hoof, and D. Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Emergence of invariance and disentanglement in deep representations
A. Achille and S. Soatto · 2018
Cited alongside, same era.
Behavior regularized offline reinforcement learning
Y. Wu, G. Tucker, and O. Nachum · 2019
Cited alongside, same era.
A framework for data-driven robotics
S. Cabi, S. G. Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Żołna, Y. Aytar, D. Budden, M. Vecerik, et al · 2019
Cited alongside, same era.
Improvisation through physical understanding: Using novel objects as tools with visual foresight
A. Xie, F. Ebert, S. Levine, and C. Finn · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Cited alongside, same era.
Later among the works it cites.
Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data
A. Mandlekar, F. Ramos, B. Boots, S. Savarese, L. Fei-Fei, A. Garg, and D. Fox · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
A. Nair, M. Dalal, A. Gupta, and S. Levine · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Zou, S. Levine, C. Finn, and T. Ma · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Later among the works it cites.
Representations for stable off-policy reinforcement learning
D. Ghosh and M. G. Bellemare · 2020
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Closest in time.
Actionable models: Unsupervised offline reinforcement learning of robotic skills
Y. Chebotar, K. Hausman, Y. Lu, T. Xiao, D. Kalashnikov, J. Varley, A. Irpan, B. Eysenbach, R. Julian, C. Finn, and S. Levine · 2021
Closest in time.
Benchmarks for deep off-policy evaluation
J. Fu, M. Norouzi, O. Nachum, G. Tucker, ziyu wang, A. Novikov, M. Yang, M. R. Zhang, Y. Chen, A. Kumar, C. Paduraru, S. Levine, and T. Paine · 2021
Closest in time.
Coarse-to-fine imitation learning: Robot manipulation from a single demonstration
E. Johns · 2021
Closest in time.
Offline reinforcement learning with fisher divergence critic regularization
I. Kostrikov, J. Tompson, R. Fergus, and O. Nachum · 2021
Closest in time.
Continuous doubly constrained batch reinforcement learning
R. Fakoor, J. Mueller, P. Chaudhari, and A. J. Smola · 2021
Closest in time.
Offline reinforcement learning from images with latent space models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2021
Closest in time.
What can i do here? learning new skills by imagining visual affordances
A. Khazatsky, A. Nair, D. Jing, and S. Levine · 2021
Closest in time.
Neorl: A near real-world benchmark for offline reinforcement learning
R. Qin, S. Gao, X. Zhang, Z. Xu, S. Huang, Z. Li, W. Zhang, and Y. Yu · 2021
Closest in time.
A minimalist approach to offline reinforcement learning
S. Fujimoto and S. S. Gu · 2021
Closest in time.