Fetching the paper…
Reading the bibliography…
We propose a unified mechanism for achieving coordination and communication in Multi-Agent Reinforcement Learning (MARL), through rewarding agents for having causal influence over other agents' actions.
The dance language and orientation of bees
von Frisch, K · 1969
Earlier work this paper cites.
Hockey helmets, concealed weapons, and daylight saving: A study of binary choices with externalities
Schelling, T. C · 1973
Earlier work this paper cites.
Strategic information transmission
Crawford, V. P. and Sobel, J · 1982
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Gers, F. A., Schmidhuber, J., and Cummins, F · 1999
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Singh, S. P., Barto, A. G., and Chentanez, N · 2004
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
Klyubin, A. S., Polani, D., and Nehaniv, C. L · 2005
Earlier work this paper cites.
Discovering communication
Oudeyer, P.-Y. and Kaplan, F · 2006
Earlier work this paper cites.
Maximization of potential information flow as a universal utility for collective behaviour
Capdepuy, P., Polani, D., and Nehaniv, C. L · 2007
Earlier work this paper cites.
Humans have evolved specialized skills of social cognition: The cultural intelligence hypothesis
Herrmann, E., Call, J., Hernàndez-Lloreda, M. V., Hare, B., and Tomasello, M · 2007
Earlier work this paper cites.
Why we cooperate
Tomasello, M · 2009
Earlier work this paper cites.
Expectations in counterfactual and theory of mind reasoning
Ferguson, H. J., Scheepers, C., and Sanford, A. J · 2010
Earlier work this paper cites.
Differentiating information transfer and causal effect
Lizier, J. T. and Prokopenko, M · 2010
Earlier work this paper cites.
How is human cooperation different?
Melis, A. P. and Semmann, D · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Schmidhuber, J · 2010
Earlier work this paper cites.
Emerging social awareness: Exploring intrinsic motivation in multiagent learning
Sequeira, P., Melo, F. S., Prada, R., and Paiva, A · 2011
Earlier work this paper cites.
Social learning and evolution: the cultural intelligence hypothesis
van Schaik, C. P. and Burkart, J. M · 2011
Earlier work this paper cites.
Structural counterfactuals: A brief introduction
Pearl, J · 2013
Earlier work this paper cites.
Research review: social motivation and oxytocin in autism–implications for joint attention development and intervention
Stavropoulos, K. K. and Carver, L. J · 2013
Cited alongside, same era.
Emotional multiagent reinforcement learning in social dilemmas
Yu, C., Zhang, M., and Ren, F · 2013
Cited alongside, same era.
Potential-based difference rewards for multiagent reinforcement learning
Devlin, S., Yliniemi, L., Kudenko, D., and Tumer, K · 2014
Cited alongside, same era.
Sapiens: A brief history of humankind
Harari, Y. N · 2014
Cited alongside, same era.
Optimal rewards for cooperative agents
Liu, B., Singh, S., Lewis, R. L., and Qin, S · 2014
Cited alongside, same era.
The Secret of Our Success: How culture is driving human evolution, domesticating our species, and making us smart
Henrich, J · 2015
Multi-agent reinforcement learning in sequential social dilemmas
Leibo, J. Z., Zambaldi, V., Lanctot, M., Marecki, J., and Graepel, T · 2017
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Later among the works it cites.
A multi-agent reinforcement learning model of common-pool resource appropriation
Perolat, J., Leibo, J. Z., Zambaldi, V., Beattie, C., Tuyls, K., and Graepel, T · 2017
Later among the works it cites.
Consequentialist conditional cooperation in social dilemmas with imperfect information
Peysakhovich, A. and Lerer, A · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S. and Rezende, D. J · 2015
Cited alongside, same era.
Untangling brain-wide dynamics in consciousness by cross-embedding
Tajima, S., Yanagawa, T., Fujii, N., and Toyoizumi, T · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, J., Assael, I. A., de Freitas, N., and Whiteson, S · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
How evolution may work through curiosity-driven developmental process
Oudeyer, P.-Y. and Smith, L. B · 2016
Cited alongside, same era.
Causal inference in statistics: a primer
Pearl, J., Glymour, M., and Jewell, N. P · 2016
Cited alongside, same era.
Barton, S. L., Waytowich, N. R., Zaroukian, E., and Asher, D. E · 2018
Closest in time.
Emergence of communication in an interactive world with consistent speakers
Bogin, B., Geva, M., and Berant, J · 2018
Closest in time.
Emergent communication through negotiation
Cao, K., Lazaridou, A., Lanctot, M., Leibo, J. Z., Tuyls, K., and Clark, S · 2018
Closest in time.
Compositional obverter communication learning from raw visual input
Choi, E., Lazaridou, A., and de Freitas, N · 2018
Closest in time.
Learning with opponent-learning awareness
Foerster, J., Chen, R. Y., Al-Shedivat, M., Whiteson, S., Abbeel, P., and Mordatch, I · 2018
Closest in time.
New and surprising ways to be mean. adversarial npcs with coupled empowerment minimisation
Guckelsberger, C., Salge, C., and Togelius, J · 2018
Closest in time.
Inequity aversion improves cooperation in intertemporal social dilemmas
Hughes, E., Leibo, J. Z., Phillips, M. G., Tuyls, K., Duéñez-Guzmán, E. A., Castañeda, A. G., Dunning, I., Zhu, T., McKee, K. R., Koster, R., et al · 2018
Closest in time.
Emergence of linguistic communication from referential games with symbolic and pixel input
Lazaridou, A., Hermann, K. M., Tuyls, K., and Clark, S · 2018
Closest in time.
Prosocial learning agents solve generalized stag hunts better than selfish ones
Peysakhovich, A. and Lerer, A · 2018
Closest in time.
Rabinowitz, N. C., Perbet, F., Song, H. F., Zhang, C., Eslami, S., and Botvinick, M · 2018
Closest in time.
Learning to share and hide intentions using information regularization
Strouse, D., Kleiman-Weiner, M., Tenenbaum, J., Botvinick, M., and Schwab, D · 2018
Closest in time.