Fetching the paper…
Reading the bibliography…
In this paper, we investigate learning temporal abstractions in cooperative multi-agent systems, using the options framework (Sutton et al, 1999).
Towards an economic theory of organization and information
Jacob Marschak · 1954
Earlier work this paper cites.
Team decision problems
R. Radner · 1962
Earlier work this paper cites.
Discrete dynamic programming
David Blackwell · 1962
Earlier work this paper cites.
A bayesian approach to problems in stochastic estimation and control
Y. C. Ho and R. C. K. Lee · 1964
Earlier work this paper cites.
Stochastic Systems: Estimation, Identification and Adaptive Control
P. R. Kumar and Pravin Varaiya · 1986
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G. Dietterich · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Learning to cooperate via policy search
Leonid Peshkin, Kee-Eung Kim, Nicolas Meuleau, and Leslie Pack Kaelbling · 2000
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Daniel S. Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Andrew G. Barto and Sridhar Mahadevan · 2003
Cited alongside, same era.
Bayesian filtering: From kalman filters to particle filters, and beyond
Zhe Chen · 2003
Cited alongside, same era.
Hierarchical multi-agent reinforcement learning
Mohammad Ghavamzadeh, Sridhar Mahadevan, and Rajbala Makar · 2006
Cited alongside, same era.
Optimal performance of networked control systems with non-classical information structures
A. Mahajan and D. Teneketzis · 2009
Cited alongside, same era.
Optimal and approximate q-value functions for decentralized pomdps
Frans A. Oliehoek, Matthijs T. J. Spaan, and Nikos A. Vlassis · 2011
Cited alongside, same era.
Information structures in optimal decentralized control
Structural Results for Partially Observed Markov Decision Processes
Vikram Krishnamurthy · 2015
Later among the works it cites.
Graph-based cross entropy method for solving multi-robot decentralized POMDPs
S. Omidshafiei, A. Agha-mohammadi, C. Amato, S. Liu, J. P. How, and J. Vian · 2016
Later among the works it cites.
Learning multiagent communication with backpropagation
Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2016
Later among the works it cites.
A Concise Introduction to Decentralized POMDPs
Frans A. Oliehoek and Christopher Amato · 2016
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aditya Mahajan, Nuno C Martins, Michael C Rotkowitz, and Serdar Yuksel · 2012
Cited alongside, same era.
Temporal Difference Learning
Klaas Apostol · 2012
Cited alongside, same era.
Decentralized stochastic control with partial history sharing: A common information approach
A. Nayyar, A. Mahajan, and D. Teneketzis · 2013
Cited alongside, same era.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Cited alongside, same era.
Decentralized control of partially observable markov decision processes using belief space macro-actions
S. Omidshafiei, A. Agha-mohammadi, C. Amato, and J. P. How · 2015
Cited alongside, same era.
Harm van Seijen, Mehdi Fatemi, Joshua Romoff, Romain Laroche, Tavian Barnes, and Jeffrey Tsang · 2017
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2017
Later among the works it cites.
Bayesian action decoder for deep multi-agent reinforcement learning
Jakob N. Foerster, Francis Song, Edward Hughes, Neil Burch, Iain Dunning, Shimon Whiteson, Matthew Botvinick, and Michael Bowling · 2018
Later among the works it cites.
Minimalistic gridworld environment for openai gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Later among the works it cites.
Multi-agent hierarchical reinforcement learning with dynamic termination
Dongge Han, Wendelin Boehmer, Michael Wooldridge, and Alex Rogers · 2019
Closest in time.