Fetching the paper…
Reading the bibliography…
Recent work in deep reinforcement learning (RL) has produced algorithms capable of mastering challenging games such as Go, chess, or shogi.
IKEA furniture assembly environment for long-horizon complex manipulation tasks
Y. Lee, E. S. Hu, Z. Yang, A. Yin, and J. J. Lim · 1911
Earlier work this paper cites.
Algorithmic motion planning
C.-K. Yap · 1986
Earlier work this paper cites.
Elephants don’t play chess
R. A. Brooks · 1990
Earlier work this paper cites.
Combining specialized reasoners and general purpose planners: a case study
S. Kambhampati, M. R. Cutkosky, M. Tenenbaum, and S. Lee · 1991
Earlier work this paper cites.
Robot Motion Planning
J.-C. Latombe · 1991
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
Motion planning in the presence of movable obstacles
G. Wilfong · 1991
Earlier work this paper cites.
Efficient exploration in reinforcement learning
S. B. Thrun · 1992
Earlier work this paper cites.
Feudal reinforcement learning
P. Dayan and G. E. Hinton · 1993
Earlier work this paper cites.
Tromp-Taylor Go rules
J. Tromp and B. Taylor · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Sokoban: A challenging single-agent search problem
A. Junghanns and J. Schaeffer · 1997
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Manipulation planning with probabilistic roadmaps
T. Siméon, J.-P. Laumond, J. Corés, and A. Sahbani · 2004
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
A. L. Strehl and M. L. Littman · 2008
Earlier work this paper cites.
GNUGo 3.8
D. Bump, G. Farneback, A. Bayer, and more · 2009
Earlier work this paper cites.
Sampling-based motion and symbolic action planning with geometric and differential constraints
E. Plaku and G. D. Hager · 2010
Earlier work this paper cites.
The Motion Grammar for physical human-robot games
N. T. Dantam, P. Kolhe, and M. Stilman · 2011
Earlier work this paper cites.
Hierarchical task and motion planning in the now
L. P. Kaelbling and T. Lozano-Pérez · 2011
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Combined task and motion planning through an extensible planner-independent interface layer
S. Srivastava, E. Fang, L. Riano, R. Chitnis, S. Russell, and P. Abbeel · 2014
Earlier work this paper cites.
Simultaneously learning at different levels of abstraction
B. Quack, F. Wörgötter, and A. Agostini · 2015
Earlier work this paper cites.
Logic-geometric programming: an optimization-based approach to combined task and motion planning
M. Toussaint · 2015
Earlier work this paper cites.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, J. Schrittwieser, K. Anderson, S. York, M. Cant, A. Cain, A. Bolton, S. Gaffney, H. King, D. Hassabis, S. Legg, and S. Petersen · 2016
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Guided search for task and motion plans using learned heuristics
R. Chitnis, D. Hadfield-Menell, A. Gupta, S. Srivastava, E. Groshev, C. Lin, and P. Abbeel · 2016
Earlier work this paper cites.
Incremental task and motion planning: A constraint-based approach
N. T. Dantam, Z. K. Kingston, S. Chaudhuri, and L. E. Kavraki · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. v. d. Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Cited alongside, same era.
Kinova modular robot arms for service robotics applications
A. Campeau-Lecours, H. Lamontagne, S. Latour, P. Fauteux, V. Maheu, F. Boucher, C. Deguire, and L.-J. C. L’Ecuyer · 2017
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. de Las Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. P. Lillicrap, and M. A. Riedmiller · 2018
Later among the works it cites.
Unsupervised predictive memory in a goal-directed agent
G. Wayne, C. Hung, D. Amos, M. Mirza, A. Ahuja, A. Grabska-Barwinska, J. W. Rae, P. Mirowski, J. Z. Leibo, A. Santoro, M. Gemici, M. Reynolds, T. Harley, J. Abramson, S. Mohamed, D. J. Rezende, D. Saxton, A. Cain, C. Hillier, D. Silver, K. Kavukcuoglu, M. Botvinick, D. Hassabis, and T. P. Lillicrap · 2018
Later among the works it cites.
Artificial Intelligence and Games
G. N. Yannakakis and J. Togelius · 2018
Later among the works it cites.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
One-shot imitation learning
Y. Duan, M. Andrychowicz, B. C. Stadie, J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Deep visual foresight for planning robotic motion
C. Finn and S. Levine · 2017
Cited alongside, same era.
Automated curriculum learning for neural networks
A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu · 2017
Cited alongside, same era.
Variational intrinsic control
K. Gregor, D. J. Rezende, and D. Wierstra · 2017
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments
N. Heess, D. TB, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. M. A. Eslami, M. A. Riedmiller, and D. Silver · 2017
Cited alongside, same era.
Data-efficient deep reinforcement learning for dexterous manipulation
I. Popov, N. Heess, T. P. Lillicrap, R. Hafner, G. Barth-Maron, M. Vecerík, T. Lampe, Y. Tassa, T. Erez, and M. A. Riedmiller · 2017
Cited alongside, same era.
I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, et al · 2019
Later among the works it cites.
RUDDER: Return decomposition for delayed rewards
J. A. Arjona-Medina, M. Gillhofer, M. Widrich, T. Unterthiner, J. Brandstetter, and S. Hochreiter · 2019
Later among the works it cites.
See, feel, act: Hierarchical learning for complex manipulation skills with multisensory fusion
N. Fazeli, M. Oller, J. Wu, Z. Wu, J. Tenenbaum, and A. Rodriguez · 2019
Later among the works it cites.
Shaping belief states with generative environment models for RL
K. Gregor, D. J. Rezende, F. Besse, Y. Wu, H. Merzic, and A. van den Oord · 2019
Later among the works it cites.
An investigation of model-free planning
A. Guez, M. Mirza, K. Gregor, R. Kabra, S. Racanière, T. Weber, D. Raposo, A. Santoro, L. Orseau, T. Eccles, G. Wayne, D. Silver, and T. P. Lillicrap · 2019
Later among the works it cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman · 2019
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
C.-C. Hung, T. Lillicrap, J. Abramson, Y. Wu, M. Mirza, F. Carnevale, A. Ahuja, and G. Wayne · 2019
Later among the works it cites.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Later among the works it cites.
Learning value functions with relational state representations for guiding task-and-motion planning
B. Kim and L. Shimanuki · 2019
Later among the works it cites.
Adversarial actor-critic method for task and motion planning problems using planning experience
B. Kim, L. P. Kaelbling, and T. Lozano-Pérez · 2019
Later among the works it cites.
OpenSpiel: A framework for reinforcement learning in games
M. Lanctot, E. Lockhart, J.-B. Lespiau, V. Zambaldi, S. Upadhyay, J. Pérolat, S. Srinivasan, F. Timbers, K. Tuyls, S. Omidshafiei, et al · 2019
Later among the works it cites.
Emergent coordination through competition
S. Liu, G. Lever, J. Merel, S. Tunyasuvunakool, N. Heess, and T. Graepel · 2019
Later among the works it cites.
Hierarchical visuomotor control of humanoids
J. Merel, A. Ahuja, V. Pham, S. Tunyasuvunakool, S. Liu, D. Tirumala, N. Heess, and G. Wayne · 2019
Later among the works it cites.
Near-optimal representation learning for hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Meta-World: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2019
Later among the works it cites.
Deep visual reasoning: Learning to predict action sequences for task and motion planning from an initial scene image
D. Driess, H. Jung-Su, and M. Toussaint · 2020
Closest in time.
RLBench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Closest in time.
Deep modular reinforcement learning for physically embedded abstract reasoning
P. Karkus, M. Mirza, A. Guez, A. Jaegle, T. Lillicrap, L. Buesing, N. Heess, and T. Weber · 2020
Closest in time.
Reusable neural skill embeddings for vision-guided whole body movement and object manipulation
J. Merel, S. Tunyasuvunakool, A. Ahuja, Y. Tassa, L. Hasenclever, V. Pham, T. Erez, G. Wayne, and N. Heess · 2020
Closest in time.
Making efficient use of demonstrations to solve hard exploration problems
T. L. Paine, B. S. Caglar Gulcehre, M. Denil, M. Hoffman, H. Soyer, R. Tanburn, S. Kapturowski, N. Rabinowitz, D. Williams, G. Barth-Maron, Z. Wang, N. de Freitas, and W. Team · 2020
Closest in time.
V-MPO: On-policy maximum a posteriori policy optimization for discrete and continuous control
H. F. Song, A. Abdolmaleki, J. T. Springenberg, A. Clark, H. Soyer, J. W. Rae, S. Noury, A. Ahuja, S. Liu, D. Tirumala, N. Heess, D. Belov, M. Riedmiller, and M. M. Botvinick · 2020
Closest in time.
dm_control: Software and tasks for continuous control
Y. Tassa, S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, and N. Heess · 2020
Closest in time.
Exploring model-based planning with policy networks
T. Wang and J. Ba · 2020
Closest in time.