Fetching the paper…
Reading the bibliography…
While recent advances in artificial intelligence have achieved human-level performance in environments like Starcraft and Go, many physical reasoning tasks remain challenging for modern algorithms.
Complexity results for sas+ planning
Christer Bäckström and Bernhard Nebel · 1995
Earlier work this paper cites.
PDDL - The Planning Domain Definition Language
M. Ghallab, A. Howe, C. Knoblock, D. Mcdermott, A. Ram, M. Veloso, D. Weld, and D. Wilkins · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Sokoban and other motion planning problems
Dorit Dor and Uri Zwick · 1999
Earlier work this paper cites.
The ff planning system: Fast plan generation through heuristic search
Jörg Hoffmann and Bernhard Nebel · 2001
Earlier work this paper cites.
Probabilistic robotics
Sebastian Thrun · 2002
Earlier work this paper cites.
Manipulation planning with probabilistic roadmaps
Thierry Siméon, Jean-Paul Laumond, Juan Cortés, and Anis Sahbani · 2004
Earlier work this paper cites.
The fast downward planning system
Malte Helmert · 2006
Earlier work this paper cites.
Domain-independent construction of pattern database heuristics for cost-optimal planning
Patrik Haslum, Adi Botea, Malte Helmert, Blai Bonet, Sven Koenig, et al · 2007
Earlier work this paper cites.
Manipulation planning among movable obstacles
Mike Stilman, Jan-Ullrich Schamburek, James Kuffner, and Tamim Asfour · 2007
Earlier work this paper cites.
Hierarchical path planning for multi-size agents in heterogeneous environments
Daniel Harabor and Adi Botea · 2008
Earlier work this paper cites.
Unifying the causal graph and additive heuristics
Malte Helmert and Héctor Geffner · 2008
Earlier work this paper cites.
Planning among movable obstacles with artificial constraints
Mike Stilman and James Kuffner · 2008
Earlier work this paper cites.
Hierarchical planning in the now
Leslie Pack Kaelbling and Tomás Lozano-Pérez · 2010
Earlier work this paper cites.
The lama planner: Guiding cost-based anytime planning with landmarks
Silvia Richter and Matthias Westphal · 2010
Earlier work this paper cites.
A framework for push-grasping in clutter
Mehmet Dogar and Siddhartha Srinivasa · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches, 2016
Jacob Andreas, Dan Klein, and Sergey Levine · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation, 2016
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Hindsight experience replay, 2017
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Differentiable physics and stable modes for tool-use and manipulation planning
Marc A Toussaint, Kelsey Rebecca Allen, Kevin A Smith, and Joshua B Tenenbaum · 2018
Later among the works it cites.
An investigation of model-free planning
Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racanière, Théophane Weber, David Raposo, Adam Santoro, Laurent Orseau, et al · 2019
Later among the works it cites.
Hindsight credit assignment
Anna Harutyunyan, Will Dabney, Thomas Mesnard, Mohammad Gheshlaghi Azar, Bilal Piot, Nicolas Heess, Hado P van Hasselt, Gregory Wayne, Satinder Singh, Doina Precup, et al · 2019
Later among the works it cites.
Sub-policy adaptation for hierarchical reinforcement learning, 2019
Alexander C. Li, Carlos Florensa, Ignasi Clavera, and Pieter Abbeel · 2019
Later among the works it cites.
Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning
Caelan Reed Garrett, Tomás Lozano-Pérez, and Leslie Pack Kaelbling · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks, 2017
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Cited alongside, same era.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Ken Kansky, Tom Silver, David A Mély, Mohamed Eldawy, Miguel Lázaro-Gredilla, Xinghua Lou, Nimrod Dorfman, Szymon Sidor, Scott Phoenix, and Dileep George · 2017
Cited alongside, same era.
Best-first width search: Exploration and exploitation in classical planning
Nir Lipovetzky and Hector Geffner · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Cited alongside, same era.
Jessica B Hamrick, Abram L Friesen, Feryal Behbahani, Arthur Guez, Fabio Viola, Sims Witherspoon, Thomas Anthony, Lars Buesing, Petar Veličković, and Théophane Weber · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, et al · 2020
Later among the works it cites.
Decoupling exploration and exploitation for meta-reinforcement learning without sacrifices, 2020
Evan Zheran Liu, Aditi Raghunathan, Percy Liang, and Chelsea Finn · 2020
Later among the works it cites.
Counterfactual credit assignment in model-free reinforcement learning, 2020
Thomas Mesnard, Théophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Tom Stepleton, Nicolas Heess, Arthur Guez, Éric Moulines, Marcus Hutter, Lars Buesing, and Rémi Munos · 2020
Later among the works it cites.
Curriculum learning for reinforcement learning domains: A framework and survey
Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E Taylor, and Peter Stone · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
The fess algorithm: A feature based approach to single-agent search
Yaron Shoham and Jonathan Schaeffer · 2020
Later among the works it cites.
There is no turning back: A self-supervised approach for reversibility-aware reinforcement learning, 2021
Nathan Grinsztajn, Johan Ferret, Olivier Pietquin, Philippe Preux, and Matthieu Geist · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
Open Ended Learning Team, Adam Stooke, Anuj Mahajan, Catarina Barros, Charlie Deck, Jakob Bauer, Jakub Sygnowski, Maja Trebacz, et al · 2021
Later among the works it cites.
Lisa: Learning interpretable skill abstractions from language, 2022
Divyansh Garg, Skanda Vaidyanath, Kuno Kim, Jiaming Song, and Stefano Ermon · 2022
Later among the works it cites.
Deep reinforcement learning with credit assignment for combinatorial optimization
Dong Yan, Jiayi Weng, Shiyu Huang, Chongxuan Li, Yichi Zhou, Hang Su, and Jun Zhu · 2022
Later among the works it cites.