Fetching the paper…
Reading the bibliography…
Cultural transmission is the domain-general social skill that allows agents to acquire and use information from each other in real-time with high fidelity and recall.
The laws of imitation
G. De Tarde · 1903
Earlier work this paper cites.
Meta-learning of sequential strategies
P. A. Ortega, J. X. Wang, M. Rowland, T. Genewein, Z. Kurth-Nelson, R. Pascanu, N. Heess, J. Veness, A. Pritzel, P. Sprechmann, S. M. Jayakumar, T. McGrath, K. J. Miller, M. G. Azar, I. Osband, N. C. Rabinowitz, A. György, S. Chiappa, S. Osindero, Y. W. Teh, H. van Hasselt, N. de Freitas, M. Botvinick, and S. Legg · 1905
Earlier work this paper cites.
Options as responses: Grounding behavioural hierarchies in multi-agent RL
A. S. Vezhnevets, Y. Wu, R. Leblond, and J. Z. Leibo · 1906
Earlier work this paper cites.
Emergent tool use from multi-agent autocurricula
B. Baker, I. Kanitscheider, T. M. Markov, Y. Wu, G. Powell, B. McGrew, and I. Mordatch · 1909
Earlier work this paper cites.
The nature of the child’s tie to his mother
J. Bowlby · 1958
Earlier work this paper cites.
The nature of love
H. F. Harlow · 1958
Earlier work this paper cites.
Dynamic programming and markov processes
R. A. Howard · 1960
Earlier work this paper cites.
Theory of fluid and crystallized intelligence: A critical experiment
R. B. Cattell · 1963
Earlier work this paper cites.
Observational learning in budgerigars
B. V. Dawson and B. Foss · 1965
Earlier work this paper cites.
Social learning theory , volume 1
A. Bandura and R. H. Walters · 1977
Earlier work this paper cites.
Mind in society: The development of higher psychological processes
L. S. Vygotsky · 1980
Earlier work this paper cites.
Intention, plans, and practical reason
M. Bratman · 1987
Earlier work this paper cites.
Culture and the evolutionary process
R. Boyd and P. J. Richerson · 1988
Earlier work this paper cites.
Imitation in animals: History, definition, and interpretation of data from the psychological laboratory
B. G. Galef Jr · 1988
Earlier work this paper cites.
Imitation, objects, tools, and the rudiments of language in human ontogeny
A. N. Meltzoff · 1988
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Meme and variations: A computational model of cultural evolution
L. Gabora · 1995
Earlier work this paper cites.
Learning tasks from a single demonstration
C. Atkeson and S. Schaal · 1997
Earlier work this paper cites.
Online learning and stochastic approximations, 1998
L. Bottou · 1998
Earlier work this paper cites.
Imitation of the sequential structure of actions by chimpanzees (pan troglodytes)
A. Whiten · 1998
Earlier work this paper cites.
The meme machine , volume 25
S. Blackmore · 2000
Earlier work this paper cites.
The evolution of prestige: Freely conferred deference as a mechanism for enhancing the benefits of cultural transmission
J. Henrich and F. J. Gil-White · 2001
Earlier work this paper cites.
Trends, rhythms, and aberrations in global climate 65 ma to present
J. Zachos, M. Pagani, L. Sloan, E. Thomas, and K. Billups · 2001
Earlier work this paper cites.
Robots that imitate humans
C. Breazeal and B. Scassellati · 2002
Earlier work this paper cites.
A model of (often mixed) stereotype content: competence and warmth respectively follow from perceived status and competition
S. T. Fiske, A. J. Cuddy, P. Glick, and J. Xu · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
M. Kearns and S. Singh · 2002
Earlier work this paper cites.
Social diversity and social preferences in mixed-motive reinforcement learning
K. R. McKee, I. M. Gemp, B. McWilliams, E. A. Duéñez-Guzmán, E. Hughes, and J. Z. Leibo · 2002
Earlier work this paper cites.
The correspondence problem
C. L. Nehaniv, K. Dautenhahn, et al · 2002
Earlier work this paper cites.
Libnoise
J. Bevins · 2003
Earlier work this paper cites.
Distributed, predictive perception of actions: a biologically inspired robotics architecture for imitation and learning
Y. Demiris and M. Johnson · 2003
Earlier work this paper cites.
Learning from and about others: Towards using imitation to bootstrap the social understanding of others by robots
C. Breazeal, D. Buchsbaum, J. Gray, D. Gatenby, and B. Blumberg · 2005
Earlier work this paper cites.
Evolution on a restless planet: Were environmental variability and environmental change major drivers of human evolution
P. J. Richerson, R. L. Bettinger, and R. Boyd · 2005
Earlier work this paper cites.
Solving deep memory pomdps with recurrent policy gradients
D. Wierstra, A. Foerster, J. Peters, and J. Schmidhuber · 2007
Earlier work this paper cites.
Cumulative cultural evolution in the laboratory: An experimental approach to the origins of structure in human language
S. Kirby, H. Cornish, and K. Smith · 2008
Earlier work this paper cites.
Learning concepts and categories: Is spacing the “enemy of induction”?
N. Kornell and R. A. Bjork · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Learning quadrupedal locomotion over challenging terrain
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2010
Earlier work this paper cites.
Biological motion preference in humans at birth: Role of dynamic and configural properties
L. Bardi, L. Regolin, and F. Simion · 2011
Earlier work this paper cites.
Cultural transmission and flexibility of partial migration patterns in a long-lived bird, the great bustard otis tarda
C. Palacín, J. C. Alonso, J. A. Alonso, M. Magaña, and C. A. Martín · 2011
Earlier work this paper cites.
Emergent complexity and zero-shot transfer via unsupervised environment design
M. Dennis, N. Jaques, E. Vinitsky, A. M. Bayen, S. Russell, A. Critch, and S. Levine · 2012
Earlier work this paper cites.
Task rules, working memory, and fluid intelligence
J. Duncan, M. Schramm, R. Thompson, and I. Dumontheil · 2012
Earlier work this paper cites.
A survey of the ontogeny of tool use: From sensorimotor experience to planning
F. Guerin, N. Krüger, and D. Kraft · 2012
Earlier work this paper cites.
A navigation mesh for dynamic environments
W. van Toll, A. Iv, and R. Geraerts · 2012
Earlier work this paper cites.
Social learning of migratory performance
T. Mueller, R. B. O’Hara, S. J. Converse, R. P. Urbanek, and W. F. Fagan · 2013
Earlier work this paper cites.
The first time ever i saw your feet: Inversion effect in newborns’ sensitivity to biological motion
L. Bardi, L. Regolin, and F. Simion · 2014
Earlier work this paper cites.
Experimentally induced innovations lead to persistent culture via conformity in wild birds
L. M. Aplin, D. R. Farine, J. Morand-Ferron, A. Cockburn, A. Thornton, and B. C. Sheldon · 2015
Earlier work this paper cites.
The secret of our success
J. Henrich · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain and Y. Bengio · 2016
Cited alongside, same era.
Understanding the multiple factors governing social learning and the diffusion of innovations
L. Aplin · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation, 2016
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
Rl^2: Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
Homo imitans? seven reasons why imitation couldn’t possibly be associative
C. Heyes · 2016
Cited alongside, same era.
Noisy networks for exploration, 2019
M. Fortunato, M. G. Azar, B. Piot, J. Menick, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, C. Blundell, and S. Legg · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castaneda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, et al · 2019
Later among the works it cites.
Survey of dropout methods for deep neural networks
A. Labach, H. Salehinejad, and S. Valaee · 2019
Later among the works it cites.
Autocurricula and the emergence of innovation from social interaction: A manifesto for multi-agent intelligence research, 2019
J. Z. Leibo, E. Hughes, M. Lanctot, and T. Graepel · 2019
Later among the works it cites.
dm_env: A python interface for reinforcement learning environments, 2019
A. Muldal, Y. Doron, J. Aslanides, T. Harley, T. Ward, and S. Liu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Learning to navigate in complex environments
P. Mirowski, R. Pascanu, F. Viola, H. Soyer, A. J. Ballard, A. Banino, M. Denil, R. Goroshin, L. Sifre, K. Kavukcuoglu, et al · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. G. Bellemare · 2016
Cited alongside, same era.
The developmental origins of selective social learning
D. Poulin-Dubois and P. Brosseau-Liard · 2016
Cited alongside, same era.
Policy distillation, 2016
A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2016
Cited alongside, same era.
Loss is its own reward: Self-supervision for reinforcement learning
E. Shelhamer, P. Mahmoudieh, M. Argus, and T. Darrell · 2016
Cited alongside, same era.
Solving rubik’s cube with a robot hand, 2019
OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q. Yuan, W. Zaremba, and L. Zhang · 2019
Later among the works it cites.
Deep exploration via randomized value functions
I. Osband, B. Van Roy, D. J. Russo, Z. Wen, et al · 2019
Later among the works it cites.
Cumulative cultural evolution in a non-copying task in children and guinea baboons
C. Saldana, J. Fagot, S. Kirby, K. Smith, and N. Claidière · 2019
Later among the works it cites.
Becoming human: A theory of ontogeny
M. Tomasello · 2019
Later among the works it cites.
Recent advances in imitation learning from observation
F. Torabi, G. Warnell, and P. Stone · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
Paired open-ended trailblazer (poet): Endlessly generating increasingly complex and diverse learning environments and their solutions, 2019
R. Wang, J. Lehman, J. Clune, and K. O. Stanley · 2019
Later among the works it cites.
Deep reinforcement learning for navigation in aaa video games
E. Alonso, M. Peter, D. Goumard, and J. Romoff · 2020
Later among the works it cites.
The DeepMind JAX Ecosystem, 2020
I. Babuschkin, K. Baumli, A. Bell, S. Bhupatiraju, J. Bruce, P. Buchlovsky, D. Budden, T. Cai, A. Clark, I. Danihelka, C. Fantacci, J. Godwin, C. Jones, T. Hennigan, M. Hessel, S. Kapturowski, T. Keck, I. Kemaev, M. King, L. Martens, V. Mikulik, T. Norman, J. Quan, G. Papamakarios, R. Ring, F. Ruiz, A. Sanchez, R. Schneider, E. Sezener, S. Spencer, S. Srinivasan, W. Stokowiec, and F. Viola · 2020
Later among the works it cites.
Never give up: Learning directed exploration strategies, 2020
A. P. Badia, P. Sprechmann, A. Vitvitskyi, D. Guo, B. Piot, S. Kapturowski, O. Tieleman, M. Arjovsky, A. Pritzel, A. Bolt, and C. Blundell · 2020
Later among the works it cites.
The hanabi challenge: A new frontier for ai research
N. Bard, J. N. Foerster, S. Chandar, N. Burch, M. Lanctot, H. F. Song, E. Parisotto, V. Dumoulin, S. Moitra, E. Hughes, et al · 2020
Later among the works it cites.
Wayfinding: the Art and Science of How We Find and Lose Our Way
M. Bond · 2020
Later among the works it cites.
Generating and adapting to diverse ad-hoc cooperation agents in hanabi
R. Canaan, X. Gao, J. Togelius, A. Nealen, and S. Menzel · 2020
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman · 2020
Later among the works it cites.
“other-play” for zero-shot coordination
H. Hu, A. Lerer, A. Peysakhovich, and J. Foerster · 2020
Later among the works it cites.
The next decade in ai: four steps towards robust artificial intelligence
G. Marcus · 2020
Later among the works it cites.
Meta-trained agents implement bayes-optimal agents
V. Mikulik, G. Delétang, T. McGrath, T. Genewein, M. Martic, S. Legg, and P. A. Ortega · 2020
Later among the works it cites.
Hindsight and sequential rationality of correlated play
D. Morrill, R. D’Orazio, R. Sarfati, M. Lanctot, J. R. Wright, A. Greenwald, and M. Bowling · 2020
Later among the works it cites.
Multi-agent social reinforcement learning improves generalization
K. Ndousse, D. Eck, S. Levine, and N. Jaques · 2020
Later among the works it cites.
Effective diversity in population based reinforcement learning
J. Parker-Holder, A. Pacchiano, K. M. Choromanski, and S. J. Roberts · 2020
Later among the works it cites.
Generalization guarantees for imitation learning
A. Z. Ren, S. Veer, and A. Majumdar · 2020
Later among the works it cites.
Hierarchical reinforcement learning for efficient exploration and transfer
L. Steccanella, S. Totaro, D. Allonsius, and A. Jonsson · 2020
Later among the works it cites.
Using unity to help solve intelligence, 2020
T. Ward, A. Bolt, N. Hemmings, S. Carter, M. Sanchez, R. Barreira, S. Noury, K. Anderson, J. Lemmon, J. Coe, P. Trochim, T. Handley, and A. Bolton · 2020
Later among the works it cites.
Learning to interactively learn and assist
M. Woodward, C. Finn, and K. Hausman · 2020
Later among the works it cites.
Procedural generalization by planning with self-supervised world models
A. Anand, J. Walker, Y. Li, E. Vértes, J. Schrittwieser, S. Ozair, T. Weber, and J. B. Hamrick · 2021
Later among the works it cites.
Quasi-equivalence discovery for zero-shot emergent communication
K. Bullard, D. Kiela, F. Meier, J. Pineau, and J. Foerster · 2021
Later among the works it cites.
Reverb: A framework for experience replay, 2021
A. Cassirer, G. Barth-Maron, E. Brevdo, S. Ramos, T. Boyd, T. Sottiaux, and M. Kroiss · 2021
Later among the works it cites.
Learned belief search: Efficiently improving policies in partially observable settings
H. Hu, A. Lerer, N. Brown, and J. Foerster · 2021
Later among the works it cites.
Replay-guided adversarial environment design
M. Jiang, M. Dennis, J. Parker-Holder, J. Foerster, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
Decoupling exploration and exploitation for meta-reinforcement learning without sacrifices, 2021
E. Z. Liu, A. Raghunathan, P. Liang, and C. Finn · 2021
Later among the works it cites.
Trajectory diversity for zero-shot coordination
A. Lupu, B. Cui, H. Hu, and J. Foerster · 2021
Later among the works it cites.
Offline meta-reinforcement learning with advantage weighting
E. Mitchell, R. Rafailov, X. B. Peng, S. Levine, and C. Finn · 2021
Later among the works it cites.
Policy manifold search: Exploring the manifold hypothesis for diversity-based neuroevolution
N. Rakicevic, A. Cully, and P. Kormushev · 2021
Later among the works it cites.
Reward is enough
D. Silver, S. Singh, D. Precup, and R. S. Sutton · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents, 2021
A. Stooke, A. Mahajan, C. Barros, C. Deck, J. Bauer, J. Sygnowski, M. Trebacz, M. Jaderberg, M. Mathieu, N. McAleese, N. Bradley-Schmieg, N. Wong, N. Porcel, R. Raileanu, S. Hughes-Fitt, V. Dalibard, and W. M. Czarnecki · 2021
Later among the works it cites.
Collaborating with humans without human data
D. Strouse, K. R. McKee, M. Botvinick, E. Hughes, and R. Everett · 2021
Later among the works it cites.
Growing knowledge culturally across generations to solve novel, complex tasks
M. H. Tessler, P. A. Tsividis, J. Madeano, B. Harper, and J. B. Tenenbaum · 2021
Later among the works it cites.
Launchpad: A programming model for distributed machine learning research
F. Yang, G. Barth-Maron, P. Stańczyk, M. Hoffman, S. Liu, M. Kroiss, A. Pope, and A. Rrustemi · 2021
Later among the works it cites.
Few-shot language coordination by modeling theory of mind
H. Zhu, G. Neubig, and Y. Bisk · 2021
Later among the works it cites.
Social learning in swarm robotics
N. Bredeche and N. Fontbonne · 2022
Closest in time.
Artificial evolution of robot bodies and control: on the interaction between evolution, learning and culture
E. Hart and L. K. Le Goff · 2022
Closest in time.
Objective misgeneralization: A consequential robustness failure
R. Shah, V. Varma, R. Kumar, M. Phuong, V. Krakovna, J. Uesato, and Z. Kenton · 2022
Closest in time.
Experiments in artificial culture: from noisy imitation to storytelling robots
A. F. Winfield and S. Blackmore · 2022
Closest in time.