Fetching the paper…
Reading the bibliography…
To operate effectively in the real world, agents should be able to act from high-dimensional raw sensory input such as images and achieve diverse goals across long time-horizons.
Benchmarking safe exploration in deep reinforcement learning
A. Ray, J. Achiam, and D. Amodei · 1910
Earlier work this paper cites.
Cognitive maps in rats and men
E. C. Tolman · 1948
Earlier work this paper cites.
A note on two problems in connexion with graphs
E. W. Dijkstra · 1959
Earlier work this paper cites.
Experiments with the graph traverser program
J. E. Doran and D. Michie · 1966
Earlier work this paper cites.
A formal basis for the heuristic determination of minimum cost paths
P. E. Hart, N. J. Nilsson, and B. Raphael · 1968
Earlier work this paper cites.
Direct trajectory optimization using nonlinear programming and collocation
C. R. Hargraves and S. W. Paris · 1987
Earlier work this paper cites.
Model predictive control: theory and practice—a survey
C. E. Garcia, D. M. Prett, and M. Morari · 1989
Earlier work this paper cites.
Learning to achieve goals
L. P. Kaelbling · 1993
Earlier work this paper cites.
Overcoming incomplete perception with util distinction memory
A. McCallum · 1993
Earlier work this paper cites.
On the hardness of approximating minimization problems
C. Lund and M. Yannakakis · 1994
Earlier work this paper cites.
A sub-constant error-probability low-degree test, and a sub-constant error-probability pcp characterization of np
R. Raz and S. Safra · 1997
Earlier work this paper cites.
A threshold of ln n for approximating set cover
U. Feige · 1998
Earlier work this paper cites.
Navigation and acquisition of spatial knowledge in a virtual maze
S. Gillner and H. A. Mallot · 1998
Earlier work this paper cites.
Rapidly-exploring random trees: A new tool for path planning, 1998
S. M. LaValle · 1998
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
R. Rubinstein · 1999
Earlier work this paper cites.
Human spatial representation: Insights from animals
R. F. Wang and E. S. Spelke · 2002
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
R. Givan, T. Dean, and M. Greig · 2003
Earlier work this paper cites.
Metrics for finite markov decision processes
N. Ferns, P. Panangaden, and D. Precup · 2004
Cited alongside, same era.
Do humans integrate routes into a cognitive map? map-versus landmark-based navigation of novel shortcuts
P. Foo, W. H. Warren, A. Duchon, and M. J. Tarr · 2005
Cited alongside, same era.
Simultaneous localization and mapping (slam): Part ii
T. Bailey and H. Durrant-Whyte · 2006
Cited alongside, same era.
Simultaneous localization and mapping: part i
H. Durrant-Whyte and T. Bailey · 2006
Cited alongside, same era.
Towards a unified theory of state abstraction for mdps
L. Li, T. J. Walsh, and M. L. Littman · 2006
Cited alongside, same era.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Distributed distributional deterministic policy gradients
G. Barth-Maron, M. W. Hoffman, D. Budden, W. Dabney, D. Horgan, D. Tb, A. Muldal, N. Heess, and T. Lillicrap · 2018
Later among the works it cites.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al · 2018
Later among the works it cites.
Notes on state abstractions, 2018
N. Jiang · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
A. V. Nair, V. Pong, M. Dalal, S. Bahl, S. Lin, and S. Levine · 2018
Later among the works it cites.
Temporal difference models: Model-free deep RL for model-based control
V. Pong, S. Gu, M. Dalal, and S. Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Vizdoom competitions: Playing doom from pixels
M. Wydmuch, M. Kempka, and W. Jaśkowski · 2018
Later among the works it cites.
Composable planning with attributes
A. Zhang, A. Lerer, S. Sukhbaatar, R. Fergus, and A. Szlam · 2018
Later among the works it cites.
A behavioral approach to visual navigation with graph localization networks
K. Chen, J. P. de Vicente, G. Sepulveda, F. Xia, A. Soto, M. Vázquez, and S. Savarese · 2019
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
B. Eysenbach, R. R. Salakhutdinov, and S. Levine · 2019
Later among the works it cites.
Mapping state space using landmarks for universal goal reaching
Z. Huang, F. Liu, and H. Su · 2019
Later among the works it cites.
Benchmarking classic and learned navigation in complex 3d environments
D. Mishkin, A. Dosovitskiy, and V. Koltun · 2019
Later among the works it cites.
Semi-parametric Topological Memory for Navigation
N. Savinov, A. Dosovitskiy, and V. Koltun · 2019
Later among the works it cites.
Learning World Graphs to Accelerate Hierarchical Reinforcement Learning
W. Shang, A. Trott, S. Zheng, C. Xiong, and R. Socher · 2019
Later among the works it cites.
Neural topological slam for visual navigation
D. S. Chaplot, R. Salakhutdinov, A. Gupta, and S. Gupta · 2020
Closest in time.
Hallucinative topological memory for zero-shot visual planning
K. Liu, T. Kurutach, C. Tung, P. Abbeel, and A. Tamar · 2020
Closest in time.
Scaling local control to large-scale topological navigation
X. Meng, N. Ratliff, Y. Xiang, and D. Fox · 2020
Closest in time.