Fetching the paper…
Reading the bibliography…
Building autonomous machines that can explore open-ended environments, discover possible interactions and build repertoires of skills is a general objective of artificial intelligence.
Self-supervised learning of image embedding for continuous control
Florensa, C., Degrave, J., Heess, N., Springenberg, J. T., and Riedmiller, M. (2019) · 1901
Earlier work this paper cites.
Actrce: Augmenting experience via teacher’s advice for multi-goal reinforcement learning
Chan, H., Wu, Y., Kiros, J., Fidler, S., and Ba, J. (2019) · 1902
Earlier work this paper cites.
Curiosity-driven multi-criteria hindsight experience replay
Lanier, J. B., McAleer, S., and Baldi, P. (2019) · 1906
Earlier work this paper cites.
Automated curricula through setter-solver interactions
Racanière, S., Lampinen, A., Santoro, A., Reichert, D., Firoiu, V., and Lillicrap, T. (2019) · 1909
Earlier work this paper cites.
Smirl: Surprise minimizing rl in dynamic environments
Berseth, G., Geng, D., Devin, C., Finn, C., Jayaraman, D., and Levine, S. (2019) · 1912
Earlier work this paper cites.
Thought and Language
Vygotsky, L. S. (1934) · 1934
Earlier work this paper cites.
Language, Thought, and Reality: Selected Writings of Benjamin Lee Whorf
Whorf, B. L. (1956) · 1956
Earlier work this paper cites.
Curiosity and exploration
Berlyne, D. E. (1966) · 1966
Earlier work this paper cites.
The Role of Tutoring in Problem Solving
Wood, D., Bruner, J. S., and Ross, G. (1976) · 1976
Earlier work this paper cites.
Sequential Thought Processes in Pdp Models
Rumelhart, D. E., Smolensky, P., McClelland, J. L., and Hinton, G. (1986) · 1986
Earlier work this paper cites.
Making the World Differentiable: On Using Self-Supervised Fully Recurrent N eu al Networks for Dynamic Reinforcement Learning and Planning in Non-Stationary Environments.
Schmidhuber, J. (1990) · 1990
Earlier work this paper cites.
Curious model-building control systems
Schmidhuber, J. (1991a) · 1991
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P., and Hinton, G. E. (1993) · 1993
Earlier work this paper cites.
Learning and development in neural networks: the importance of starting small
Elman, J. L. (1993) · 1993
Earlier work this paper cites.
Learning to achieve goals
Kaelbling, L. P. (1993) · 1993
Earlier work this paper cites.
Why Children Talk to Themselves
Berk, L. E. (1994) · 1994
Earlier work this paper cites.
The Helmholtz Machine
Dayan, P., Hinton, G. E., Neal, R. M., and Zemel, R. S. (1995) · 1995
Earlier work this paper cites.
Goal-driven learning
Ram, A., Leake, D. B., and Leake, D. (1995) · 1995
Earlier work this paper cites.
Multitask learning
Caruana, R. (1997) · 1997
Earlier work this paper cites.
Being There: Putting Brain, Body, and World Together Again
Clark, A. (1998) · 1998
Earlier work this paper cites.
Intra-option learning about temporally abstract actions.
Sutton, R. S., Precup, D., and Singh, S. P. (1998) · 1998
Earlier work this paper cites.
The scientist in the crib: Minds, brains, and how children learn
Gopnik, A., Meltzoff, A. N., and Kuhl, P. K. (1999) · 1999
Earlier work this paper cites.
Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
The Cultural Origins of Human Cognition
Tomasello, M. (1999) · 1999
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
McGovern, A., and Barto, A. G. (2001) · 2001
Earlier work this paper cites.
The Epigenesis of Meaning in Human Beings, and Possibly in Robots
Zlatev, J. (2001) · 2001
Earlier work this paper cites.
Modularity, Language, and the Flexibility of Thought
Carruthers, P. (2002) · 2002
Earlier work this paper cites.
From Embodied to Socially Embedded Agents – Implications for Interaction-Aware Robots
Dautenhahn, K., Ogden, B., and Quick, T. (2002) · 2002
Earlier work this paper cites.
Social Situatedness of Natural and Artificial Intelligence: Vygotsky and Beyond
Lindblom, J., and Ziemke, T. (2003) · 2003
Earlier work this paper cites.
Maximizing Learning Progress: An Internal Reward System for Development
Kaplan, F., and Oudeyer, P.-Y. (2004) · 2004
Earlier work this paper cites.
Using relative novelty to identify useful temporal abstractions in reinforcement learning
Simsek, Ö., and Barto, A. G. (2004) · 2004
Earlier work this paper cites.
Temporal-difference networks
Sutton, R. S., and Tanner, B. (2004) · 2004
Earlier work this paper cites.
Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text
Hill, F., Mokra, S., Wong, N., and Harley, T. (2020b) · 2005
Earlier work this paper cites.
Lynch, C., and Sermanet, P. (2020) · 2005
Earlier work this paper cites.
Language-conditioned goal generation: a new approach to language grounding for rl
Colas, C., Akakzia, A., Oudeyer, P.-Y., Chetouani, M., and Sigaud, O. (2020) · 2006
Earlier work this paper cites.
In search of the neural circuits of intrinsic motivation
Kaplan, F., and Oudeyer, P.-Y. (2007) · 2007
Earlier work this paper cites.
Intrinsic Motivation Systems for Autonomous Mental Development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V. (2007) · 2007
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Oudeyer, P.-Y., and Kaplan, F. (2007) · 2007
Earlier work this paper cites.
The goal construct in psychology
Elliot, A. J., and Fryer, J. W. (2008) · 2008
Earlier work this paper cites.
Grimgep: Learning progress for robust goal sampling in visual deep reinforcement learning
Kovač, G., Laversanne-Finot, A., and Oudeyer, P.-Y. (2020) · 2008
Earlier work this paper cites.
Skill characterization based on betweenness
Simsek, Ö., and Barto, A. G. (2008) · 2008
Earlier work this paper cites.
Cognitive developmental robotics: A survey
Asada, M., Hosoda, K., Kuniyoshi, Y., Ishiguro, H., Inui, T., Yoshikawa, Y., Ogino, M., and Yoshida, C. (2009) · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey.
Taylor, M. E., and Stone, P. (2009) · 2009
Earlier work this paper cites.
Constructing a Language
Tomasello, M. (2009) · 2009
Earlier work this paper cites.
Intrinsically motivated goal exploration for active motor learning in robots: A case study
Baranes, A., and Oudeyer, P.-Y. (2010) · 2010
Earlier work this paper cites.
An Empowerment-based Solution to Robotic Manipulation Tasks with Sparse Rewards
Dai, S., Xu, W., Hofmann, A., and Williams, B. (2020) · 2010
Earlier work this paper cites.
Goal babbling permits direct learning of inverse kinematics
Rolf, M., Steil, J. J., and Gienger, M. (2010) · 2010
Earlier work this paper cites.
Parameter-exploring policy gradients
Sehnke, F., Osendorfer, C., Rückstieß, T., Graves, A., Peters, J., and Schmidhuber, J. (2010) · 2010
Earlier work this paper cites.
Intrinsically motivated reinforcement learning: An evolutionary perspective
Singh, S., Lewis, R. L., Barto, A. G., and Sorg, J. (2010) · 2010
Earlier work this paper cites.
Evolving a diversity of virtual creatures through novelty search and local competition
Lehman, J., and Stanley, K. O. (2011) · 2011
Earlier work this paper cites.
Towards a Vygotskyan Cognitive Robotics: The Role of Language as a Cognitive Tool
Mirolli, M., and Parisi, D. (2011) · 2011
Earlier work this paper cites.
Model Learning for Robot Control: A Survey
Nguyen-Tuong, D., and Peters, J. (2011) · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Lopes, M., Lang, T., Toussaint, M., and Oudeyer, P. (2012) · 2012
Earlier work this paper cites.
What Do Words Do? Toward a Theory of Language-Augmented Thought
Lupyan, G. (2012) · 2012
Earlier work this paper cites.
Active learning of inverse models with intrinsically motivated goal exploration in robots
Baranes, A., and Oudeyer, P.-Y. (2013) · 2013
Earlier work this paper cites.
Robust active binocular vision through intrinsically motivated learning
Lonini, L., Forestier, S., Teulière, C., Zhao, Y., Shi, B. E., and Triesch, J. (2013) · 2013
Earlier work this paper cites.
Information driven self-organization of complex robotic behaviors
Martius, G., Der, R., and Ay, N. (2013) · 2013
Earlier work this paper cites.
Efficient exploratory learning of inverse kinematics on a bionic elephant trunk
Rolf, M., and Steil, J. J. (2013) · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Self-organization of early vocal development in infants and machines: The role of intrinsic motivation
Moulin-Frier, C., Nguyen, S. M., and Oudeyer, P.-Y. (2014) · 2014
Cited alongside, same era.
Socially guided intrinsic motivation for robot learning of motor skills
Nguyen, M., and Oudeyer, P.-Y. (2014) · 2014
Cited alongside, same era.
Natural evolution strategies
Wierstra, D., Schaul, T., Glasmachers, T., Sun, Y., Peters, J., and Schmidhuber, J. (2014) · 2014
Cited alongside, same era.
Developmental Robotics: From Babies to Robots
Cangelosi, A., and Schlesinger, M. (2015) · 2015
Cited alongside, same era.
The psychology and neuroscience of curiosity
Kidd, C., and Hayden, B. Y. (2015) · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Later among the works it cites.
Goal-conditioned imitation learning
Ding, Y., Florensa, C., Abbeel, P., and Phielipp, M. (2019) · 2019
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S. (2019) · 2019
Later among the works it cites.
Open-Endedness for the Sake of Open-Endedness
Hintze, A. (2019) · 2019
Later among the works it cites.
Language as an abstraction for hierarchical deep reinforcement learning
Jiang, Y., Gu, S., Murphy, K., and Finn, C. (2019) · 2019
Later among the works it cites.
A survey of reinforcement learning informed by natural language
Luketina, J., Nardelli, N., Farquhar, G., Foerster, J. N., Andreas, J., Grefenstette, E., Whiteson, S., and Rocktäschel, T. (2019) · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Mohamed, S., and Rezende, D. J. (2015) · 2015
Cited alongside, same era.
Illuminating search spaces by mapping elites
Mouret, J.-B., and Clune, J. (2015) · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Cited alongside, same era.
Neural module networks
Andreas, J., Rohrbach, M., Darrell, T., and Klein, D. (2016) · 2016
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Cited alongside, same era.
Modular active curiosity-driven discovery of tool use
Forestier, S., and Oudeyer, P.-Y. (2016) · 2016
Cited alongside, same era.
Later among the works it cites.
Planning with goal-conditioned policies
Nasiriany, S., Pong, V., Lin, S., and Levine, S. (2019) · 2019
Later among the works it cites.
Successor options: An option discovery framework for reinforcement learning
Ramesh, R., Tomar, M., and Ravindran, B. (2019) · 2019
Later among the works it cites.
Why Open-Endedness Matters
Stanley, K. O. (2019) · 2019
Later among the works it cites.
Self-supervised learning of distance functions for goal-conditioned reinforcement learning.
Venkattaramanujam, S., Crawford, E., Doan, T., and Precup, D. (2019) · 2019
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
Warde-Farley, D., de Wiele, T. V., Kulkarni, T. D., Ionescu, C., Hansen, S., and Mnih, V. (2019) · 2019
Later among the works it cites.
The laplacian in RL: learning representations with efficient approximations
Wu, Y., Tucker, G., and Nachum, O. (2019) · 2019
Later among the works it cites.
Interactive language learning by question answering
Yuan, X., Côté, M.-A., Fu, J., Lin, Z., Pal, C., Bengio, Y., and Trischler, A. (2019) · 2019
Later among the works it cites.
Imitating Interactive Intelligence.
Abramson, J., Ahuja, A., Brussee, A., Carnevale, F., Cassin, M., Clark, S., Dudzik, A., Georgiev, P., Guy, A., Harley, T., Hill, F., Hung, A., Kenton, Z., Landon, J., Lillicrap, T., Mathewson, K., Muldal, A., Santoro, A., Savinov, N., Varma, V., Wayne, G., Wong, N., Yan, C., and Zhu, R. (2020) · 2020
Closest in time.
Fast reinforcement learning with generalized policy updates
Barreto, A., Hou, S., Borsa, D., Silver, D., and Precup, D. (2020) · 2020
Closest in time.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z. (2020) · 2020
Closest in time.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2020
Closest in time.
Explore, discover and learn: Unsupervised discovery of state-covering skills
Campos, V., Trott, A., Xiong, C., Socher, R., Giró-i-Nieto, X., and Torres, J. (2020) · 2020
Closest in time.
Plangan: Model-based planning with sparse rewards and multiple goals
Charlesworth, H., and Montana, G. (2020) · 2020
Closest in time.
Exploratory play, rational action, and efficient search
Chu, J., and Schulz, L. (2020) · 2020
Closest in time.
Higher: Improving instruction following with hindsight generation for experience replay
Cideron, G., Seurin, M., Strub, F., and Pietquin, O. (2020) · 2020
Closest in time.
Hierarchically organized latent modules for exploratory search in morphogenetic systems
Etcheverry, M., Moulin-Frier, C., and Oudeyer, P. (2020) · 2020
Closest in time.
Rewriting history with inverse RL: hindsight inference for policy improvement
Eysenbach, B., Geng, X., Levine, S., and Salakhutdinov, R. R. (2020) · 2020
Closest in time.
Dynamical distance learning for semi-supervised and unsupervised skill discovery
Hartikainen, K., Geng, X., Haarnoja, T., and Levine, S. (2020) · 2020
Closest in time.
Active world model learning with progress curiosity
Kim, K., Sano, M., Freitas, J. D., Haber, N., and Yamins, D. (2020) · 2020
Closest in time.
One solution is not all you need: Few-shot extrapolation via structured maxent RL
Kumar, S., Kumar, A., Levine, S., and Finn, C. (2020) · 2020
Closest in time.
Towards practical multi-object manipulation using relational reinforcement learning
Li, R., Jabri, A., Darrell, T., and Agrawal, P. (2020) · 2020
Closest in time.
Adapting behavior via intrinsic reward: a survey and empirical study
Linke, C., Ady, N. M., White, M., Degris, T., and White, A. (2020) · 2020
Closest in time.
Working memory graphs
Loynd, R., Fernandez, R., Çelikyilmaz, A., Swaminathan, A., and Hausknecht, M. J. (2020) · 2020
Closest in time.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P. (2020) · 2020
Closest in time.
Contextual imagined goals for self-supervised robotic learning
Nair, A., Bahl, S., Khazatsky, A., Pong, V., Berseth, G., and Levine, S. (2020) · 2020
Closest in time.
Maximum entropy gain exploration for long horizon multi-goal reinforcement learning
Pitis, S., Chan, H., Zhao, S., Stadie, B. C., and Ba, J. (2020) · 2020
Closest in time.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S. (2020) · 2020
Closest in time.
RIDE: rewarding impact-driven exploration for procedurally-generated environments
Raileanu, R., and Rocktäschel, T. (2020) · 2020
Closest in time.
Curious hierarchical actor-critic reinforcement learning
Röder, F., Eppe, M., Nguyen, P. D., and Wermter, S. (2020) · 2020
Closest in time.
A benchmark for systematic generalization in grounded language understanding
Ruis, L., Andreas, J., Baroni, M., Bouchacourt, D., and Lake, B. M. (2020) · 2020
Closest in time.
Intrinsically motivated open-ended learning in autonomous robots
Santucci, V. G., Oudeyer, P.-Y., Barto, A., and Baldassarre, G. (2020) · 2020
Closest in time.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D. (2020) · 2020
Closest in time.
Planning to explore via self-supervised world models
Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., and Pathak, D. (2020) · 2020
Closest in time.
Dynamics-aware unsupervised discovery of skills
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K. (2020) · 2020
Closest in time.
A boolean task algebra for reinforcement learning
Tasse, G. N., James, S. D., and Rosman, B. (2020) · 2020
Closest in time.
Automatic curriculum learning through value disagreement
Zhang, Y., Abbeel, P., and Pinto, L. (2020) · 2020
Closest in time.
DECSTR: Learning goal-directed abstract behaviors using pre-verbal spatial predicates in intrinsically motivated agents
Akakzia, A., Colas, C., Oudeyer, P.-Y., Chetouani, M., and Sigaud, O. (2021) · 2021
Closest in time.
Learning with amigo: Adversarially motivated intrinsic goals
Campero, A., Raileanu, R., Küttler, H., Tenenbaum, J. B., Rocktäschel, T., and Grefenstette, E. (2021) · 2021
Closest in time.
Specializing Versatile Skill Libraries using Local Mixture of Experts
Celik, O., Zhou, D., Li, G., Becker, P., and Neumann, G. (2021) · 2021
Closest in time.
Glib: Efficient exploration for relational model-based reinforcement learning via goal-literal babbling
Chitnis, R., Silver, T., Tenenbaum, J., Kaelbling, L. P., and Lozano-Pérez, T. (2021) · 2021
Closest in time.
Variational Empowerment as Representation Learning for Goal-Based Reinforcement Learning
Choi, J., Sharma, A., Lee, H., Levine, S., and Gu, S. S. (2021) · 2021
Closest in time.
Epidemioptim: A toolbox for the optimization of control policies in epidemiological models
Colas, C., Hejblum, B., Rouillon, S., Thiébaut, R., Oudeyer, P.-Y., Moulin-Frier, C., and Prague, M. (2021) · 2021
Closest in time.
First return, then explore
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J. (2021) · 2021
Closest in time.
Discovering Generalizable Skills via Automated Generation of Diverse Tasks
Fang, K., Zhu, Y., Savarese, S., and Fei-Fei, L. (2021) · 2021
Closest in time.
Clic: Curriculum learning and imitation for object control in nonrewarding environments
Fournier, P., Colas, C., Chetouani, M., and Sigaud, O. (2021) · 2021
Closest in time.
Recurrent independent mechanisms
Goyal, A., Lamb, A., Hoffmann, J., Sodhani, S., Levine, S., Bengio, Y., and Schölkopf, B. (2021) · 2021
Closest in time.
On the role of planning in model-based deep reinforcement learning
Hamrick, J. B., Friesen, A. L., Behbahani, F., Guez, A., Viola, F., Witherspoon, S., Anthony, T., Buesing, L. H., Velickovic, P., and Weber, T. (2021) · 2021
Closest in time.
Grounded language learning fast and slow
Hill, F., Tieleman, O., von Glehn, T., Wong, N., Merzic, H., and Clark, S. (2021) · 2021
Closest in time.
Grounding spatio-temporal language with transformers
Karch, T., Teodorescu, L., Hofmann, K., Moulin-Frier, C., and Oudeyer, P.-Y. (2021) · 2021
Closest in time.
The Intersection of Planning and Learning
Moerland, T. M. (2021) · 2021
Closest in time.
Discovering Diverse Solutions in Deep Reinforcement Learning
Osa, T., Tangkaratt, V., and Sugiyama, M. (2021) · 2021
Closest in time.
General value function networks
Schlegel, M., Jacobsen, A., Abbas, Z., Patterson, A., White, A., and White, M. (2021) · 2021
Closest in time.
Towards teachable autonomous agents
Sigaud, O., Caselles-Dupré, H., Colas, C., Akakzia, A., Oudeyer, P.-Y., and Chetouani, M. (2021) · 2021
Closest in time.
Open-ended learning leads to generally capable agents
Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., et al. (2021) · 2021
Closest in time.