Fetching the paper…
Reading the bibliography…
Advances in artificial intelligence often stem from the development of new environments that abstract real-world situations into a form where research can be done conveniently.
R. Wang, J. Lehman, J. Clune, and K. O. Stanley · 1901
Earlier work this paper cites.
J. Z. Leibo, E. Hughes, M. Lanctot, and T. Graepel · 1903
Earlier work this paper cites.
The foundations of statistics
L. J. Savage · 1951
Earlier work this paper cites.
Stochastic Games
L. S. Shapley · 1953
Earlier work this paper cites.
I, pencil
L. E. Read · 1958
Earlier work this paper cites.
Dynamic programming and markov processes
R. A. Howard · 1960
Earlier work this paper cites.
Games with incomplete information played by “Bayesian” players, I–III part I. The basic model
J. C. Harsanyi · 1967
Earlier work this paper cites.
The tragedy of the commons
G. Hardin · 1968
Earlier work this paper cites.
Hockey helmets, concealed weapons, and daylight saving: A study of binary choices with externalities
T. C. Schelling · 1973
Earlier work this paper cites.
The Evolution of Cooperation
R. M. Axelrod · 1984
Earlier work this paper cites.
The economics of rights, cooperation and welfare
R. Sugden et al · 1986
Earlier work this paper cites.
The firm, the market, and the law
R. H. Coase · 1988
Earlier work this paper cites.
A general theory of equilibrium selection in games
J. C. Harsanyi, R. Selten, et al · 1988
Earlier work this paper cites.
The symbol grounding problem
S. Harnad · 1990
Earlier work this paper cites.
Governing the Commons: The Evolution of Institutions for Collective Action
E. Ostrom · 1990
Earlier work this paper cites.
Decentralized, dispersed exchange without an auctioneer
P. Albin and D. K. Foley · 1992
Earlier work this paper cites.
Rational learning leads to Nash equilibrium
E. Kalai and E. Lehrer · 1993
Earlier work this paper cites.
Game theory and the social contract: just playing , volume 2
K. G. Binmore et al · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Economics
P. A. Samuelson and W. D. Nordhaus · 1995
Earlier work this paper cites.
Evolution of the social contract
B. Skyrms · 1996
Earlier work this paper cites.
The social brain hypothesis
R. I. Dunbar · 1998
Earlier work this paper cites.
Social dilemmas: The anatomy of cooperation
P. Kollock · 1998
Earlier work this paper cites.
Artificial ecosystem selection
W. Swenson, D. S. Wilson, and R. Elias · 2000
Earlier work this paper cites.
Neural consequences of enviromental enrichment
H. Van Praag, G. Kempermann, and F. H. Gage · 2000
Earlier work this paper cites.
Bilateral trade and ‘small-world’ networks
A. Wilhite · 2001
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. M. Kakade · 2003
Earlier work this paper cites.
How types of goods and property rights jointly affect collective action
E. Ostrom · 2003
Earlier work this paper cites.
A day of great illumination: BF Skinner’s discovery of shaping
G. B. Peterson · 2004
Earlier work this paper cites.
Spatial strategies and territoriality in the maine lobster industry
J. M. Acheson and R. J. Gardner · 2005
Earlier work this paper cites.
Understanding institutional diversity
E. Ostrom · 2005
Earlier work this paper cites.
Nice guys finish first: The competitive altruism hypothesis
C. L. Hardy and M. van Vugt · 2006
Earlier work this paper cites.
Five rules for the evolution of cooperation
M. A. Nowak · 2006
Earlier work this paper cites.
Agent-based computational economics: A constructive approach to economic theory
L. Tesfatsion · 2006
Earlier work this paper cites.
Rational decisions in large worlds
K. Binmore · 2007
Earlier work this paper cites.
If multi-agent learning is the answer, what is the question?
Y. Shoham, R. Powers, and T. Grenager · 2007
Earlier work this paper cites.
Core knowledge
E. S. Spelke and K. D. Kinzler · 2007
Earlier work this paper cites.
Turfs in the lab: institutional innovation in real-time dynamic spatial commons
M. A. Janssen and E. Ostrom · 2008
Earlier work this paper cites.
Game theory evolving
H. Gintis · 2009
Earlier work this paper cites.
Emergence of networks in distance-constrained trade
K. Venkat and W. Wakeland · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
The cultural niche: Why social learning is essential for human adaptation
R. Boyd, P. J. Richerson, and J. Henrich · 2011
Earlier work this paper cites.
What is money? an alternative to searle’s institutional facts
J. Smit, F. Buekens, and S. Du Plessis · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Property rights
I. Segal and M. D. Whinston · 2013
Earlier work this paper cites.
Territoriality as a driver of fishers’ spatial behavior in the northumberland lobster fishery
R. A. Turner, T. Gray, N. V. Polunin, and S. M. Stead · 2013
Earlier work this paper cites.
The bounds of reason
H. Gintis · 2014
Earlier work this paper cites.
The missing link: Ab models and dynamic microsimulation
M. Richiardi · 2014
Cited alongside, same era.
Heads-up limit hold’em poker is solved
M. Bowling, N. Burch, M. Johanson, and O. Tammelin · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
C. Finn, P. Christiano, P. Abbeel, and S. Levine · 2016
Cited alongside, same era.
Using natural language for reward shaping in reinforcement learning
P. Goyal, S. Niekum, and R. J. Mooney · 2019
Later among the works it cites.
Précis of cognitive gadgets: The cultural evolution of thinking
C. Heyes · 2019
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
C.-C. Hung, T. Lillicrap, J. Abramson, Y. Wu, M. Mirza, F. Carnevale, A. Ahuja, and G. Wayne · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castaneda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, et al · 2019
Later among the works it cites.
Obstacle tower: a generalization challenge in vision, control, and planning
A. Juliani, A. Khalifa, V.-P. Berges, J. Harper, E. Teng, H. Henry, A. Crespi, J. Togelius, and D. Lange · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to communicate with deep multi-agent reinforcement learning
J. Foerster, I. A. Assael, N. De Freitas, and S. Whiteson · 2016
Cited alongside, same era.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Cited alongside, same era.
Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction
M. Kleiman-Weiner, M. K. Ho, J. L. Austerweil, M. L. Littman, and J. B. Tenenbaum · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, and K. Silver, David andKavukcuoglu · 2016
Cited alongside, same era.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Cited alongside, same era.
Building machines that learn and think for themselves: Commentary on Lake et al., Behavioral and Brain Sciences, 2017
M. Botvinick, D. G. Barrett, P. Battaglia, N. de Freitas, D. Kumaran, J. Z. Leibo, T. Lillicrap, J. Modayil, S. Mohamed, N. C. Rabinowitz, D. J. Rezende, A. Santoro, T. Schaul, C. Summerfield, G. Wayne, T. Weber, D. Wierstra, S. Legg, and D. Hassabis · 2017
Cited alongside, same era.
Why are there so many explanations for primate brain evolution?
R. Dunbar and S. Shultz · 2017
Cited alongside, same era.
M. Karl, P. Becker-Ehmck, M. Soelch, D. Benbouzid, P. v. d. Smagt, and J. Bayer · 2019
Later among the works it cites.
Deep exploration via randomized value functions
I. Osband, B. Van Roy, D. J. Russo, and Z. Wen · 2019
Later among the works it cites.
Language is power: Representing states using natural language in reinforcement learning
E. Schwartz, G. Tennenholtz, C. Tessler, and S. Mannor · 2019
Later among the works it cites.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
H. F. Song, A. Abdolmaleki, J. T. Springenberg, A. Clark, H. Soyer, J. W. Rae, S. Noury, A. Ahuja, S. Liu, D. Tirumala, et al · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Imitating interactive intelligence
J. Abramson, A. Ahuja, I. Barr, A. Brussee, F. Carnevale, M. Cassin, R. Chhaparia, S. Clark, B. Damoc, A. Dudzik, et al · 2020
Later among the works it cites.
Emergent reciprocity and team formation from randomized uncertain social preferences
B. Baker · 2020
Later among the works it cites.
The Hanabi challenge: A new frontier for AI research
N. Bard, J. N. Foerster, S. Chandar, N. Burch, M. Lanctot, H. F. Song, E. Parisotto, V. Dumoulin, S. Moitra, E. Hughes, et al · 2020
Later among the works it cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Later among the works it cites.
Exploring zero-shot emergent communication in embodied multi-agent populations
K. Bullard, F. Meier, D. Kiela, J. Pineau, and J. Foerster · 2020
Later among the works it cites.
The animal-ai testbed and competition
M. Crosby, B. Beyret, M. Shanahan, J. Hernández-Orallo, L. Cheke, and M. Halina · 2020
Later among the works it cites.
Real world games look like spinning tops
W. M. Czarnecki, G. Gidel, B. Tracey, K. Tuyls, S. Omidshafiei, D. Balduzzi, and M. Jaderberg · 2020
Later among the works it cites.
D3c: Reducing the price of anarchy in multi-agent learning
I. Gemp, K. R. McKee, R. Everett, E. A. Duéñez-Guzmán, Y. Bachrach, D. Balduzzi, and A. Tacchetti · 2020
Later among the works it cites.
“other-play” for zero-shot coordination
H. Hu, A. Lerer, A. Peysakhovich, and J. Foerster · 2020
Later among the works it cites.
Model-free conventions in multi-agent reinforcement learning with heterogeneous preferences
R. Köster, K. R. McKee, R. Everett, L. Weidinger, W. S. Isaac, E. Hughes, E. A. Duéñez-Guzmán, T. Graepel, M. Botvinick, and J. Z. Leibo · 2020
Later among the works it cites.
Transforming task representations to perform novel tasks
A. K. Lampinen and J. L. McClelland · 2020
Later among the works it cites.
Principles of economics
N. G. Mankiw · 2020
Later among the works it cites.
Methodological issues of spatial agent-based models
S. Manson, L. An, K. C. Clarke, A. Heppenstall, J. Koch, B. Krzyzanowski, F. Morgan, D. O’Sullivan, B. C. Runck, E. Shook, et al · 2020
Later among the works it cites.
Social diversity and social preferences in mixed-motive reinforcement learning
K. R. McKee, I. Gemp, B. McWilliams, E. A. Duèñez-Guzmán, E. Hughes, and J. Z. Leibo · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
E. Parisotto, F. Song, J. Rae, R. Pascanu, C. Gulcehre, S. Jayakumar, M. Jaderberg, R. L. Kaufman, A. Clark, S. Noury, et al · 2020
Later among the works it cites.
Increasing generality in machine learning through procedural content generation
S. Risi and J. Togelius · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, et al · 2020
Later among the works it cites.
Too many cooks: Coordinating multi-agent collaboration through inverse planning
S. A. Wu, R. E. Wang, J. A. Evans, J. Tenenbaum, D. C. Parkes, and M. Kleiman-Weiner · 2020
Later among the works it cites.
The AI economist: Improving equality and productivity with ai-driven tax policies
S. Zheng, A. Trott, S. Srinivasa, N. Naik, M. Gruesbeck, D. C. Parkes, and R. Socher · 2020
Later among the works it cites.
Offline learning from demonstrations and unlabeled experience
K. Zolna, A. Novikov, K. Konyushkova, C. Gulcehre, Z. Wang, Y. Aytar, M. Denil, N. de Freitas, and S. Reed · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Later among the works it cites.
Muesli: Combining improvements in policy optimization
M. Hessel, I. Danihelka, F. Viola, A. Guez, S. Schmitt, L. Sifre, T. Weber, D. Silver, and H. Van Hasselt · 2021
Later among the works it cites.
Scalable evaluation of multi-agent reinforcement learning with Melting Pot
J. Z. Leibo, E. A. Duéñez-Guzmán, A. Vezhnevets, J. P. Agapiou, P. Sunehag, R. Koster, J. Matyas, C. Beattie, I. Mordatch, and T. Graepel · 2021
Later among the works it cites.
Deep reinforcement learning models the emergent dynamics of human cooperation
K. R. McKee, E. Hughes, T. O. Zhu, M. J. Chadwick, R. Koster, A. G. Castaneda, C. Beattie, T. Graepel, M. Botvinick, and J. Z. Leibo · 2021
Later among the works it cites.
Ella: Exploration through learned language abstraction
S. Mirchandani, S. Karamcheti, and D. Sadigh · 2021
Later among the works it cites.
Emergent communication under competition
M. Noukhovitch, T. LaCroix, A. Lazaridou, and A. Courville · 2021
Later among the works it cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation
O. OpenAI, M. Plappert, R. Sampedro, T. Xu, I. Akkaya, V. Kosaraju, P. Welinder, R. D’Sa, A. Petron, H. P. d. O. Pinto, et al · 2021
Later among the works it cites.
Reward is enough
D. Silver, S. Singh, D. Precup, and R. S. Sutton · 2021
Later among the works it cites.
Normative disagreement as a challenge for cooperative ai
J. Stastny, M. Riché, A. Lyzhov, J. Treutlein, A. Dafoe, and J. Clifton · 2021
Later among the works it cites.
Collaborating with humans without human data
D. Strouse, K. McKee, M. Botvinick, E. Hughes, and R. Everett · 2021
Later among the works it cites.
Agent-based computational economics: Overview and brief history
L. Tesfatsion · 2021
Later among the works it cites.
Discovery of options via meta-learned subgoals
V. Veeriah, T. Zahavy, M. Hessel, Z. Xu, J. Oh, I. Kemaev, H. P. van Hasselt, D. Silver, and S. Singh · 2021
Later among the works it cites.
E. Vinitsky, R. Köster, J. P. Agapiou, E. Duéñez-Guzmán, A. S. Vezhnevets, and J. Z. Leibo · 2021
Later among the works it cites.
Few-shot language coordination by modeling theory of mind
H. Zhu, G. Neubig, and Y. Bisk · 2021
Later among the works it cites.
Improving intrinsic exploration with language abstractions
J. Mu, V. Zhong, R. Raileanu, M. Jiang, N. Goodman, T. Rocktäschel, and E. Grefenstette · 2022
Closest in time.
Abstraction for deep reinforcement learning
M. Shanahan and M. Mitchell · 2022
Closest in time.
Semantic exploration from language abstractions and pretrained representations
A. C. Tam, N. C. Rabinowitz, A. K. Lampinen, N. A. Roy, S. C. Y. Chan, D. Strouse, J. X. Wang, A. Banino, and F. Hill · 2022
Closest in time.
Ace research area: Learning and the embodied mind, 2022
L. Tesfatsion · 2022
Closest in time.