Fetching the paper…
Reading the bibliography…
In this paper, we study the problem of learning a repertoire of low-level skills from raw images that can be sequenced to complete long-horizon visuomotor tasks.
Connectionist ai, symbolic ai, and the brain
P. Smolensky · 1987
Earlier work this paper cites.
Skillful control under uncertainty via direct reinforcement learning
V. Gullapalli · 1995
Earlier work this paper cites.
Biped dynamic walking using reinforcement learning
H. Benbrahim and J. A. Franklin · 1997
Earlier work this paper cites.
Pddl–the planning domain definition language–version 1.2
D. McDermott, M. Ghallab, A. Howe, C. Knoblock, A. Ram, M. Veloso, D. Weld, and D. Wilkins · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
N. Kohl and P. Stone · 2004
Earlier work this paper cites.
Stochastic policy gradient reinforcement learning on a simple 3d biped
R. Tedrake, T. Zhang, and H. Seung · 2004
Earlier work this paper cites.
In Proceedings of the Twenty-First International Conference on Machine Learning , ICML ’04, page 1, 2004
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Maximum margin planning
N. D. Ratliff, J. A. Bagnell, and M. A. Zinkevich · 2006
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Manipulation primitives—a universal interface between sensor-based motion control and robot programming
T. Kröger, B. Finkemeyer, and F. M. Wahl · 2010
Earlier work this paper cites.
Learning to control a low-cost manipulator using data-efficient reinforcement learning
M. Deisenroth, C. Rasmussen, and D. Fox · 2011
Earlier work this paper cites.
Hierarchical task and motion planning in the now
L. P. Kaelbling and T. Lozano-Perez · 2011
Earlier work this paper cites.
Autonomous reinforcement learning on raw visual input data in a real world application
S. Lange, M. Riedmiller, and A. Voigtländer · 2012
Earlier work this paper cites.
Combined task and motion planning through an extensible planner-independent interface layer
S. Srivastava, E. Fang, L. Riano, R. Chitnis, S. Russell, and P. Abbeel · 2014
Earlier work this paper cites.
The option-critic architecture
P. Bacon, J. Harb, and D. Precup · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
C. Finn, I. Goodfellow, and S. Levine · 2016
Earlier work this paper cites.
Maximum entropy deep inverse reinforcement learning, 2016
M. Wulfmeier, P. Ondruska, and I. Posner · 2016
Earlier work this paper cites.
Guided cost learning: Deep inverse optimal control via policy optimization
C. Finn, S. Levine, and P. Abbeel · 2016
Earlier work this paper cites.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Earlier work this paper cites.
Deep predictive policy training using reinforcement learning
A. Ghadirzadeh, A. Maki, D. Kragic, and M. Björkman · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. S. Sastry, and S. A. Seshia · 2017
Earlier work this paper cites.
DDCO: discovery of deep continuous options forrobot learning from demonstrations
S. Krishnan, R. Fox, I. Stoica, and K. Goldberg · 2017
Cited alongside, same era.
Learning composable models of parameterized skills
L. Kaelbling and T. Lozano-Perez · 2017
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2018
Cited alongside, same era.
From skills to symbols: Learning symbolic representations for abstract high-level planning
G. Konidaris, L. Kaelbling, and T. Lozano-Perez · 2018
Cited alongside, same era.
Differentiable physics and stable modes for tool-use and manipulation planning
M. Toussaint, K. R. Allen, K. Smith, and J. Tenenbaum · 2018
Cited alongside, same era.
Neural task graphs: Generalizing to unseen tasks from a single video demonstration
D.-A. Huang, S. Nair, D. Xu, Y. Zhu, A. Garg, L. Fei-Fei, S. Savarese, and J. C. Niebles · 2019
Later among the works it cites.
Reasoning about physical interactions with object-oriented prediction and planning
M. Janner, S. Levine, W. Freeman, J. Tenenbaum, C. Finn, and J. Wu · 2019
Later among the works it cites.
Learning robotic manipulation through visual planning and acting
A. Wang, T. Kurutach, K. Liu, P. Abbeel, and A. Tamar · 2019
Later among the works it cites.
Time-agnostic prediction: Predicting predictable video frames
D. Jayaraman, F. Ebert, A. Efros, and S. Levine · 2019
Later among the works it cites.
Solving rubik’s cube with a robot hand
OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q. Yuan, W. Zaremba, and L. Zhang · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke, et al · 2018
Cited alongside, same era.
Learning synergies between pushing and grasping with self-supervised deep reinforcement learning
A. Zeng, S. Song, S. Welker, J. Lee, A. Rodriguez, and T. Funkhouser · 2018
Cited alongside, same era.
Visual reinforcement learning with imagined goals
A. V. Nair, V. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Cited alongside, same era.
Learning robust rewards with adverserial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Cited alongside, same era.
Semi-parametric topological memory for navigation, 2018
N. Savinov, A. Dosovitskiy, and V. Koltun · 2018
Cited alongside, same era.
Few-shot goal inference for visuomotor learning and planning
A. Xie, A. Singh, S. Levine, and C. Finn · 2018
Cited alongside, same era.
Variational inverse control with events: A general framework for data-driven reward definition
J. Fu, A. Singh, D. Ghosh, L. Yang, and S. Levine · 2018
Cited alongside, same era.
Later among the works it cites.
The surprising effectiveness of linear models for visual foresight in object pile manipulation
H. Suh and R. Tedrake · 2020
Later among the works it cites.
Information theoretic model predictive q-learning
M. Bhardwaj, A. Handa, D. Fox, and B. Boots · 2020
Later among the works it cites.
Offline reinforcement learning from images with latent space models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2020
Later among the works it cites.
The ingredients of real-world robotic reinforcement learning
H. Zhu, J. Yu, A. Gupta, D. Shah, K. Hartikainen, A. Singh, V. Kumar, and S. Levine · 2020
Later among the works it cites.
Safe imitation learning via fast bayesian reward inference from preferences, 2020
D. S. Brown, R. Coleman, R. Srinivasan, and S. Niekum · 2020
Later among the works it cites.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg · 2020
Later among the works it cites.
Hierarchical foresight: Self-supervised learning of long-horizon tasks via visual subgoal generation
S. Nair and C. Finn · 2020
Later among the works it cites.
Batch exploration with examples for scalable robotic reinforcement learning
A. S. Chen, H. Nam, S. Nair, and C. Finn · 2020
Later among the works it cites.
Dynamics-aware unsupervised discovery of skills, 2020
A. Sharma, S. Gu, S. Levine, V. Kumar, and K. Hausman · 2020
Later among the works it cites.
Learning to generalize across long-horizon tasks from human demonstrations
A. Mandlekar, D. Xu, R. Martín-Martín, S. Savarese, and L. Fei-Fei · 2020
Later among the works it cites.
Skew-explore: Learn faster in continuous spaces with sparse rewards, 2020
X. Chen, Y. Gao, A. Ghadirzadeh, M. Bjorkman, G. Castellano, and P. Jensfelt · 2020
Later among the works it cites.
Towards general and autonomous learning of core skills: A case study in locomotion
R. Hafner, T. Hertweck, P. Kloppner, M. Bloesch, M. Neunert, M. Wulfmeier, S. Tunyasuvunakool, N. Heess, and M. A. Riedmiller · 2020
Later among the works it cites.
Hallucinative topological memory for zero-shot visual planning, 2020
K. Liu, T. Kurutach, C. Tung, P. Abbeel, and A. Tamar · 2020
Later among the works it cites.
Broadly-exploring, local-policy trees for long-horizon task planning, 2020
B. Ichter, P. Sermanet, and C. Lynch · 2020
Later among the works it cites.
Long-horizon visual planning with goal-conditioned hierarchical predictors, 2020
K. Pertsch, O. Rybkin, F. Ebert, C. Finn, D. Jayaraman, and S. Levine · 2020
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Closest in time.
Model-based visual planning with self-supervised functional distances
S. Tian, S. Nair, F. Ebert, S. Dasari, B. Eysenbach, C. Finn, and S. Levine · 2021
Closest in time.
Learning generalizable robotic reward functions from ”in-the-wild” human videos, 2021
A. S. Chen, S. Nair, and C. Finn · 2021
Closest in time.
Greedy hierarchical variational autoencoders for large-scale video prediction
B. Wu, S. Nair, R. Martin-Martin, L. Fei-Fei*, and C. Finn* · 2021
Closest in time.
Replacing rewards with examples: Example-based policy search via recursive classification, 2021
B. Eysenbach, S. Levine, and R. Salakhutdinov · 2021
Closest in time.
Action priors for large action spaces in robotics
O. Biza, D. Wang, R. W. Platt, J.-W. van de Meent, and L. L. S. Wong · 2021
Closest in time.