Fetching the paper…
Reading the bibliography…
A hallmark of intelligence is the ability to deduce general principles from examples, which are correct beyond the range of those observed.
Strips: A new approach to the application of theorem proving to problem solving
Richard E Fikes and Nils J Nilsson · 1971
Earlier work this paper cites.
On the complexity of blocks-world planning
Naresh Gupta and Dana S Nau · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Pddl-the planning domain definition language, 1998
Drew McDermott, Malik Ghallab, Adele Howe, Craig Knoblock, Ashwin Ram, Manuela Veloso, Daniel Weld, and David Wilkins · 1998
Earlier work this paper cites.
The ff planning system: Fast plan generation through heuristic search
Jörg Hoffmann and Bernhard Nebel · 2001
Earlier work this paper cites.
Optimizing search engines using clickthrough data
Thorsten Joachims · 2002
Earlier work this paper cites.
Learning-assisted automated planning: looking back, taking stock, going forward
Terry Zimmerman and Subbarao Kambhampati · 2003
Earlier work this paper cites.
The fast downward planning system
Malte Helmert · 2006
Earlier work this paper cites.
Learning heuristic functions from relaxed plans
Sung Wook Yoon, Alan Fern, and Robert Givan · 2006
Earlier work this paper cites.
The first learning track of the international planning competition
Alan Fern, Roni Khardon, and Prasad Tadepalli · 2011
Earlier work this paper cites.
Machine learning methods for planning
Steven Minton · 2014
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Learning to rank for synthesizing planning heuristics
Caelan Reed Garrett, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio · 2017
Later among the works it cites.
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al · 2018
Later among the works it cites.
Learning generalized reactive policies using deep neural networks
Edward Groshev, Aviv Tamar, Maxwell Goldstein, Siddharth Srivastava, and Pieter Abbeel · 2018
Later among the works it cites.
Action schema networks: Generalised policies with deep learning
Sam Toyer, Felipe Trevizan, Sylvie Thiébaux, and Lexing Xie · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Thinking fast and slow with deep learning and tree search
Thomas Anthony, Zheng Tian, and David Barber · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Later among the works it cites.
Solving np-hard problems on graphs by reinforcement learning without domain knowledge
Kenshin Abe, Zijian Xu, Issei Sato, and Masashi Sugiyama · 2019
Later among the works it cites.
An investigation of model-free planning
Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racanière, Théophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, et al · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Learning domain-independent planning heuristics with hypergraph networks
William Shen, Felipe Trevizan, and Sylvie Thiébaux · 2019
Later among the works it cites.