Fetching the paper…
Reading the bibliography…
Open-ended learning agents must efficiently prioritize goals in vast possibility spaces, focusing on those that maximize learning progress (LP).
A theory of human curiosity
Berlyne, D. E · 1954
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V · 2006
Earlier work this paper cites.
In search of the neural circuits of intrinsic motivation
Kaplan, F. and Oudeyer, P.-Y · 2007
Earlier work this paper cites.
Visualizing data using t-sne
van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
R-IAC: Robust intrinsically motivated exploration and active learning
Baranes, A. and Oudeyer, P.-Y · 2009
Earlier work this paper cites.
Competence progress intrinsic motivation
Stout, A. and Barto, A. G · 2010
Earlier work this paper cites.
Intrinsically motivated learning systems: an overview
Baldassarre, G. and Mirolli, M · 2012
Earlier work this paper cites.
Active learning of inverse models with intrinsically motivated goal exploration in robots
Baranes, A. and Oudeyer, P.-Y · 2012
Earlier work this paper cites.
The strategic student approach for life-long exploration and learning
Lopes, M. and Oudeyer, P.-Y · 2012
Earlier work this paper cites.
The strategic student approach for life-long exploration and learning
Lopes, M. and Oudeyer, P.-Y · 2012
Earlier work this paper cites.
Exploration strategies in developmental robotics: A unified probabilistic framework
Moulin-Frier, C. and Oudeyer, P.-Y · 2013
Earlier work this paper cites.
Self-organization of early vocal development in infants and machines: the role of intrinsic motivation
Moulin-Frier, C., Nguyen, S. M., and Oudeyer, P.-Y · 2013
Earlier work this paper cites.
PowerPlay: Training an increasingly general problem solver by continually searching for the simplest still unsolvable problem
Schmidhuber, J · 2013
Earlier work this paper cites.
Multi-armed bandits for intelligent tutoring systems
Clement, B., Roy, D., Oudeyer, P.-Y., and Lopes, M · 2015
Earlier work this paper cites.
The psychology and neuroscience of curiosity
Kidd, C. and Hayden, B. Y · 2015
Earlier work this paper cites.
Modular active curiosity-driven discovery of tool use
Forestier, S. and Oudeyer, P.-Y · 2016
Earlier work this paper cites.
The malmo platform for artificial intelligence experimentation
Johnson, M., Hofmann, K., Hutton, T., and Bignell, D · 2016
Earlier work this paper cites.
How evolution may work through curiosity-driven developmental process
Oudeyer, P.-Y. and Smith, L. B · 2016
Earlier work this paper cites.
Automatic goal generation for reinforcement learning agents
Held, D., Geng, X., Florensa, C., and Abbeel, P · 2017
Earlier work this paper cites.
Accuracy-based curriculum learning in deep reinforcement learning, 2018
Fournier, P., Sigaud, O., Chetouani, M., and Oudeyer, P.-Y · 2018
Earlier work this paper cites.
Towards a neuroscience of active sampling and curiosity
Gottlieb, J. and Oudeyer, P.-Y · 2018
Earlier work this paper cites.
Curiosity driven exploration of learned disentangled goal spaces
Laversanne-Finot, A., Pere, A., and Oudeyer, P.-Y · 2018
Earlier work this paper cites.
Control what you can: Intrinsically motivated task-planning agent
Blaes, S., Vlastelica Pogančić, M., Zhu, J., and Martius, G · 2019
Cited alongside, same era.
Babyai: A platform to study the sample efficiency of grounded language learning
Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., and Bengio, Y · 2019
Cited alongside, same era.
CURIOUS: intrinsically motivated modular multi-goal reinforcement learning
Colas, C., Fournier, P., Chetouani, M., Sigaud, O., and Oudeyer, P.-Y · 2019
Cited alongside, same era.
Where’s the reward? a review of reinforcement learning for instructional sequencing
Doroudi, S., Aleven, V., and Brunskill, E · 2019
Cited alongside, same era.
Skew-fit: State-covering self-supervised reinforcement learning
Pong, V. H., Dalal, M., Lin, S., Nair, A., Bahl, S., and Levine, S · 2019
Cited alongside, same era.
Unsupervised control through non-parametric discriminative rewards
Grimgep: Learning progress for robust goal sampling in visual deep reinforcement learning
Kovač, G., Laversanne-Finot, A., and Oudeyer, P.-Y · 2022
Later among the works it cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2022
Later among the works it cites.
Grounding large language models in interactive environments with online reinforcement learning
Carta, T., Romac, C., Wolf, T., Lamprier, S., Sigaud, O., and Oudeyer, P.-Y · 2023
Later among the works it cites.
Stein variational goal generation for adaptive exploration in multi-goal reinforcement learning
Castanet, N., Sigaud, O., and Lamprier, S · 2023
Later among the works it cites.
QLoRA: Efficient finetuning of quantized LLMs
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Reasoning with language model is planning with world model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Warde-Farley, D., de Wiele, T. V., Kulkarni, T., Ionescu, C., Hansen, S., and Mnih, V · 2019
Cited alongside, same era.
Language as a cognitive tool to imagine goals in curiosity-driven exploration
Colas, C., Karch, T., Lair, N., Dussoux, J.-M., Moulin-Frier, C., Dominey, P. F., and Oudeyer, P.-Y · 2020
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A., Russell, S., Critch, A., and Levine, S · 2020
Cited alongside, same era.
Wordcraft: An environment for benchmarking commonsense agents
Jiang, M., Luketina, J., Nardelli, N., Minervini, P., Torr, P., Whiteson, S., and Rocktäschel, T · 2020
Cited alongside, same era.
Teacher–student curriculum learning, 2020
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J · 2020
Cited alongside, same era.
Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement Learning
Pitis, S., Chan, H., Zhao, S., Stadie, B., and Ba, J · 2020
Cited alongside, same era.
Automated curriculum generation through setter-solver interactions
Racaniere, S., Lampinen, A., Santoro, A., Reichert, D., Firoiu, V., and Lillicrap, T · 2020
Cited alongside, same era.
Hao, S., Gu, Y., Ma, H., Hong, J., Wang, Z., Wang, D., and Hu, Z · 2023
Later among the works it cites.
General intelligence requires rethinking exploration
Jiang, M., Rocktäschel, T., and Grefenstette, E · 2023
Later among the works it cites.
Young children calibrate effort based on the trajectory of their performance
Leonard, J. A., Cordrey, S. R., Liu, H. Z., and Mackey, A. P · 2023
Later among the works it cites.
Learning progress mediates the link between cognitive effort and task engagement
Sayalı, C., Heling, E., and Cools, R · 2023
Later among the works it cites.
Reflexion: language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2023
Later among the works it cites.
A definition of open-ended learning problems for goal-conditioned agents
Sigaud, O., Baldassarre, G., Colas, C., Doncieux, S., Duro, R. J., Perrin-Gilbert, N., and Santucci, V. G · 2023
Later among the works it cites.
Sac-glam: Improving online rl for llm agents with soft actor-critic and hindsight relabeling, 2024
Gaven, L., Romac, C., Carta, T., Lamprier, S., Sigaud, O., and Oudeyer, P.-Y · 2024
Later among the works it cites.
Practice makes perfect: Planning to learn skill parameter policies
Kumar, N., Silver, T., McClinton, W., Zhao, L., Proulx, S., Lozano-Pérez, T., Kaelbling, L. P., and Barry, J · 2024
Later among the works it cites.
Kinetix: Investigating the training of general agents through open-ended physics-based control tasks, 2024
Matthews, M., Beukman, M., Lu, C., and Foerster, J · 2024
Later among the works it cites.
ACES: Generating a diversity of challenging programming puzzles with autotelic generative models
Pourcel, J., Colas, C., Molinaro, G., Oudeyer, P.-Y., and Teodorescu, L · 2024
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models, 2024
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2024
Later among the works it cites.
OMNI: Open-endedness via models of human notions of interestingness
Zhang, J., Lehman, J., Stanley, K., and Clune, J · 2024
Later among the works it cites.
ArCHer: Training language model agents via hierarchical multi-turn RL
Zhou, Y., Zanette, A., Pan, J., Levine, S., and Kumar, A · 2024
Later among the works it cites.
Open r1: A fully open reproduction of deepseek-r1, January 2025
Face, H · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Tang, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Wan, Y., Liu, Y., Cui, Z., Zhang, Z., and Qiu, Z · 2025
Closest in time.
Humans monitor learning progress in curiosity-driven exploration
Ten, A., Kaushik, P., Oudeyer, P.-Y., and Gottlieb, J · 2041
Closest in time.