Fetching the paper…
Reading the bibliography…
An agent that has well understood the environment should be able to apply its skills for any given goals, leading to the fundamental problem of learning the Universal Value Function Approximator (UVFA).
A formal basis for the heuristic determination of minimum cost paths
Peter E Hart, Nils J Nilsson, and Bertram Raphael · 1968
Earlier work this paper cites.
Efficient memory-based learning for robot control
Andrew William Moore · 1990
Earlier work this paper cites.
Probabilistic roadmaps for path planning in high-dimensional configuration spaces
Lydia Kavraki, Petr Svestka, and Mark H Overmars · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner, et al · 1998
Earlier work this paper cites.
Rapidly-exploring random trees: A new tool for path planning
Steven M LaValle · 1998
Earlier work this paper cites.
Geometric Control of Mechanical Systems
Francesco Bullo and Andrew D. Lewis · 2004
Earlier work this paper cites.
Sparse multidimensional scaling using landmark points
Vin De Silva and Joshua B Tenenbaum · 2004
Earlier work this paper cites.
Computing the shortest path: A search meets graph theory
Andrew V Goldberg and Chris Harrelson · 2005
Earlier work this paper cites.
k-means++: The advantages of careful seeding
David Arthur and Sergei Vassilvitskii · 2007
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2018
Later among the works it cites.
Modeling the long term future in model-based reinforcement learning
Nan Rosemary Ke, Amanpreet Singh, Ahmed Touati, Anirudh Goyal, Yoshua Bengio, Devi Parikh, and Dhruv Batra · 2018
Later among the works it cites.
Robot motion planning in learned latent spaces
Brian Ichter and Marco Pavone · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Cited alongside, same era.
Model-based planning with discrete and continuous actions
Mikael Henaff, William F Whitney, and Yann LeCun · 2017
Cited alongside, same era.
Hierarchical reinforcement learning with hindsight
Andrew Levy, Robert Platt, and Kate Saenko · 2018
Cited alongside, same era.
Universal agent for disentangling environments and tasks
Jiayuan Mao, Honghua Dong, and Joseph J Lim · 2018
Cited alongside, same era.
Temporal difference models: Model-free deep RL for model-based control
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine · 2018
Cited alongside, same era.
Ahmed H. Qureshi and Michael C. Yip · 2018
Later among the works it cites.
Prm-rl: Long-range robotic navigation tasks by combining reinforcement learning and sampling-based planning
Aleksandra Faust, Kenneth Oslund, Oscar Ramirez, Anthony Francis, Lydia Tapia, Marek Fiser, and James Davidson · 2018
Later among the works it cites.
Semi-parametric topological memory for navigation
Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun · 2018
Later among the works it cites.
Composable planning with attributes
Amy Zhang, Adam Lerer, Sainbayar Sukhbaatar, Rob Fergus, and Arthur Szlam · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang Shane Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, Vikash Kumar, and Wojciech Zaremba · 2018
Later among the works it cites.
Unsupervised visuomotor control through distributional planning networks
Tianhe Yu, Gleb Shevchuk, Dorsa Sadigh, and Chelsea Finn · 2019
Closest in time.
Towards learning abstract representations for locomotion planning in high-dimensional state spaces
Tobias Klamt and Sven Behnke · 2019
Closest in time.