Fetching the paper…
Reading the bibliography…
Coagent policy gradient algorithms (CPGAs) are reinforcement learning algorithms for training a class of stochastic neural networks called coagent networks.
The hedonistic neuron: A theory of memory, learning, and intelligence
Klopf, A. H · 1982
Earlier work this paper cites.
Learning by statistical cooperation of self-interested neuron-like computing elements
Barto, A. G · 1985
Earlier work this paper cites.
Learning Automata: an Introduction
Narendra, K. S. and Thathachar, M. A · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
Bertsekas, D. P. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Coordinated reinforcement learning
Guestrin, C., Lagoudakis, M., and Parr, R · 2002
Earlier work this paper cites.
Conditional random fields for multi-agent reinforcement learning
Zhang, X., Aberdeen, D., and Vishwanathan, S · 2007
Cited alongside, same era.
Multi-agent reinforcement learning: An overview
Buşoniu, L., Babuška, R., and De Schutter, B · 2010
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., and White, A · 2011
Cited alongside, same era.
Policy gradient coagent networks
Thomas, P. S · 2011
Cited alongside, same era.
Conjugate markov decision processes
Thomas, P. S. and Barto, A. G · 2011
Cited alongside, same era.
Motor primitive discovery
Thomas, P. S. and Barto, A. G · 2012
Cited alongside, same era.
Optimal rewards for cooperative agents
Liu, B., Singh, S., Lewis, R. L., and Qin, S · 2014
Later among the works it cites.
Gradient estimation using stochastic computation graphs
Schulman, J., Heess, N., Weber, T., and Abbeel, P · 2015
Later among the works it cites.
Artificial Intelligence: A Modern Approach
Russell, S. J. and Norvig, P · 2016
Later among the works it cites.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Later among the works it cites.
Learning abstract options
Riemer, M., Liu, M., and Tesauro, G · 2018
Later among the works it cites.
Dac: The double actor-critic architecture for learning options
Zhang, S. and Whiteson, S · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…