Fetching the paper…
Reading the bibliography…
Everyday tasks of long-horizon and comprising a sequence of multiple implicit subtasks still impose a major challenge in offline robot control.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
I. Higgins, L. Matthey, A. Pal, C. P. Burgess, X. Glorot, M. M. Botvinick, S. Mohamed, and A. Lerchner · 2017
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma · 2017
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, et al · 2018
Earlier work this paper cites.
Learning by playing solving sparse reward tasks from scratch
M. Riedmiller, R. Hafner, T. Lampe, M. Neunert, J. Degrave, T. van de Wiele, V. Mnih, N. Heess, and J. T. Springenberg · 2018
Earlier work this paper cites.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2018
Earlier work this paper cites.
Neural task programming: Learning to generalize across hierarchical tasks
D. Xu, S. Nair, Y. Zhu, J. Gao, A. Garg, L. Fei-Fei, and S. Savarese · 2018
Earlier work this paper cites.
Learning an embedding space for transferable robot skills
K. Hausman, J. T. Springenberg, Z. Wang, N. Heess, and M. Riedmiller · 2018
Earlier work this paper cites.
Learning to walk via deep reinforcement learning
T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Earlier work this paper cites.
Mcp: Learning composable hierarchical control with multiplicative compositional policies
X. B. Peng, M. Chang, G. Zhang, P. Abbeel, and S. Levine · 2019
Earlier work this paper cites.
Time-agnostic prediction: Predicting predictable video frames
D. Jayaraman, F. Ebert, A. A. Efros, and S. Levine · 2019
Earlier work this paper cites.
Dynamics learning with cascaded variational inference for multi-step manipulation
K. Fang, Y. Zhu, A. Garg, S. Savarese, and L. Fei-Fei · 2019
Earlier work this paper cites.
Planning with goal-conditioned policies
S. Nasiriany, V. Pong, S. Lin, and S. Levine · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Earlier work this paper cites.
Learning actionable representations with goal conditioned policies
D. Ghosh, A. Gupta, and S. Levine · 2019
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims · 2020
Cited alongside, same era.
Learning latent plans from play
C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet · 2020
Model-based reinforcement learning via latent-space collocation
O. Rybkin, C. Zhu, A. Nagabandi, K. Daniilidis, I. Mordatch, and S. Levine · 2021
Later among the works it cites.
Broadly-exploring, local-policy trees for long-horizon task planning
B. Ichter, P. Sermanet, and C. Lynch · 2021
Later among the works it cites.
Value function spaces: Skill-centric state abstractions for long-horizon reasoning
D. Shah, P. Xu, Y. Lu, T. Xiao, A. T. Toshev, S. Levine, et al · 2021
Later among the works it cites.
Parrot: Data-driven behavioral priors for reinforcement learning
A. Singh, H. Liu, G. Zhou, A. Yu, N. Rhinehart, and S. Levine · 2021
Later among the works it cites.
OPAL: offline primitive discovery for accelerating offline reinforcement learning
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2021
Later among the works it cites.
Model-based visual planning with self-supervised functional distances
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hierarchical foresight: Self-supervised learning of long-horizon tasks via visual subgoal generation
S. Nair and C. Finn · 2020
Cited alongside, same era.
Accelerating reinforcement learning with learned skill priors
K. Pertsch, Y. Lee, and J. J. Lim · 2020
Cited alongside, same era.
Learning robot skills with temporal variational inference
T. Shankar and A. Gupta · 2020
Cited alongside, same era.
GTI: learning to generalize across long-horizon tasks from human demonstrations
A. Mandlekar, D. Xu, R. Martín-Martín, S. Savarese, and L. Fei-Fei · 2020
Cited alongside, same era.
Multi-agent manipulation via locomotion using hierarchical sim2real
O. Nachum, M. Ahn, H. Ponte, S. S. Gu, and V. Kumar · 2020
Cited alongside, same era.
Adversarial skill networks: Unsupervised robot skill learning from videos
O. Mees, M. Merklinger, G. Kalweit, and W. Burgard · 2020
Cited alongside, same era.
S. Tian, S. Nair, F. Ebert, S. Dasari, B. Eysenbach, C. Finn, and S. Levine · 2021
Later among the works it cites.
Composing pick-and-place tasks by grounding language
O. Mees and W. Burgard · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
I. Kostrikov, A. Nair, and S. Levine · 2022
Closest in time.
Should i run offline reinforcement learning or behavioral cloning?
A. Kumar, J. Hong, A. Singh, and S. Levine · 2022
Closest in time.
Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard · 2022
Closest in time.
Demonstration-bootstrapped autonomous practicing via multi-task reinforcement learning
A. Gupta, C. Lynch, B. Kinman, G. Peake, S. Levine, and K. Hausman · 2022
Closest in time.
Skill-based model-based reinforcement learning
L. X. Shi, J. J. Lim, and Y. Lee · 2022
Closest in time.
Planning to practice: Efficient online fine-tuning by composing goals in latent space
K. Fang, P. Yin, A. Nair, and S. Levine · 2022
Closest in time.
Affordance learning from play for sample-efficient policy learning
J. Borja-Diaz, O. Mees, G. Kalweit, L. Hermann, J. Boedecker, and W. Burgard · 2022
Closest in time.
Hierarchical policies for cluttered-scene grasping with latent plans
L. Wang, X. Meng, Y. Xiang, and D. Fox · 2022
Closest in time.
What matters in language conditioned robotic imitation learning over unstructured data
O. Mees, L. Hermann, and W. Burgard · 2022
Closest in time.
R3m: A universal visual representation for robot manipulation
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Closest in time.