Fetching the paper…
Reading the bibliography…
Hierarchies of temporally decoupled policies present a promising approach for enabling structured exploration in complex long-term planning problems.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1993
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart J Russell · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Best-response multiagent learning in non-stationary environments
Michael Weinberg and Jeffrey S Rosenschein · 2004
Earlier work this paper cites.
Multi-agent reinforcement learning: A survey
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2006
Earlier work this paper cites.
Traffic jams without bottlenecks—experimental evidence for the physical mechanism of the formation of a jam
Yuki Sugiyama, Minoru Fukui, Macoto Kikuchi, Katsuya Hasebe, Akihiro Nakayama, Katsuhiro Nishinari, Shin-ichi Tadaki, and Satoshi Yukawa · 2008
Earlier work this paper cites.
Policy gradient coagent networks
Philip S Thomas · 2011
Earlier work this paper cites.
Conjugate markov decision processes
Philip S Thomas and Andrew G Barto · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Constrained optimization and Lagrange multiplier methods
Dimitri P Bertsekas · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Jakob N. Foerster, Yannis M. Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Cited alongside, same era.
Multiagent cooperation and competition with deep reinforcement learning
Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Later among the works it cites.
Dissipating stop-and-go waves in closed and open networks via deep reinforcement learning
Abdul Rahman Kreidieh, Cathy Wu, and Alexandre M Bayen · 2018
Later among the works it cites.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang Shane Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2016
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Cited alongside, same era.
Stabilising experience replay for deep multi-agent reinforcement learning
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip HS Torr, Pushmeet Kohli, and Shimon Whiteson · 2017
Cited alongside, same era.
Andrew Levy, Robert Platt, and Kate Saenko · 2017
Cited alongside, same era.
Deeploco: Dynamic locomotion skills using hierarchical deep reinforcement learning
Xue Bin Peng, Glen Berseth, KangKang Yin, and Michiel van de Panne · 2017
Cited alongside, same era.
Flow: Architecture and benchmarking for reinforcement learning in traffic control
Cathy Wu, Aboudy Kreidieh, Kanaad Parvate, Eugene Vinitsky, and Alexandre M Bayen
Cited in the paper.
Emergent behaviors in mixed-autonomy traffic
Cathy Wu, Aboudy Kreidieh, Eugene Vinitsky, and Alexandre M Bayen
Cited in the paper.
Benchmarks for reinforcement learning in mixed-autonomy traffic
Eugene Vinitsky, Aboudy Kreidieh, Luc Le Flem, Nishant Kheterpal, Kathy Jang, Fangyu Wu, Richard Liaw, Eric Liang, and Alexandre M Bayen · 2018
Later among the works it cites.
Biases for emergent communication in multi-agent reinforcement learning
Tom Eccles, Yoram Bachrach, Guy Lever, Angeliki Lazaridou, and Thore Graepel · 2019
Closest in time.
Sub-policy adaptation for hierarchical reinforcement learning
Alexander C Li, Carlos Florensa, Ignasi Clavera, and Pieter Abbeel · 2019
Closest in time.
Planning with goal-conditioned policies
Soroush Nasiriany, Vitchyr H Pong, Steven Lin, and Sergey Levine · 2019
Closest in time.