Fetching the paper…
Reading the bibliography…
Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set by existing data.
Go-explore: a new approach for hard-exploration problems, 2021
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O. Stanley, and Jeff Clune · 1901
Earlier work this paper cites.
A formal framework for robot construction problems: A hybrid planning approach, 2019
Faseeh Ahmad, Esra Erdem, and Volkan Patoglu · 1903
Earlier work this paper cites.
Efficient exploration via state marginal matching, 2020
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 1906
Earlier work this paper cites.
Minerl: A large-scale dataset of minecraft demonstrations, 2019
William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov · 1907
Earlier work this paper cites.
A baseline for few-shot image classification, 2020
Guneet S. Dhillon, Pratik Chaudhari, Avinash Ravichandran, and Stefano Soatto · 1909
Earlier work this paper cites.
Rlbench: The robot learning benchmark and learning environment, 2019
Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J. Davison · 1909
Earlier work this paper cites.
On the measure of intelligence, 2019
François Chollet · 1911
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning, 2020
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 1912
Earlier work this paper cites.
Reinforcement learning upside down: Don’t predict rewards – just map them to actions, 2020
Juergen Schmidhuber · 1912
Earlier work this paper cites.
Block construction: Children’s developmental landmarks in representation of space
Stuart Reifel · 1984
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: The meta-meta-. hook, 1987
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
On the complexity of blocks-world planning
Naresh Gupta and Dana S. Nau · 1992
Earlier work this paper cites.
Learning to achieve goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Multitask learning , page 95–133
Rich Caruana · 1998
Earlier work this paper cites.
Motor processes in mental rotation
Mark Wexler, Stephen M Kosslyn, and Alain Berthoz · 1998
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning, 2021
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2004
Earlier work this paper cites.
The nethack learning environment, 2020
Heinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel · 2006
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V. Hafner · 2006
Earlier work this paper cites.
Play = Learning: How Play Motivates and Enhances Children’s Cognitive and Social-Emotional Growth
Dorothy G. Singer, Roberta Michnick Golinkoff, and Kathy Hirsh-Pasek · 2006
Earlier work this paper cites.
Mike Paterson, Yuval Peres, Mikkel Thorup, Peter Winkler, and Uri Zwick · 2007
Earlier work this paper cites.
Core knowledge
Elizabeth S. Spelke and Katherine D. Kinzler · 2007
Earlier work this paper cites.
The development of spatial skills through interventions involving block building activities
Beth Casey, Nicole Andrews, Holly Schindle, Joanne Kersh, Alexandra Samper, and Juanita Copley · 2008
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Stuart Russell and Peter Norvig · 2010
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research [best of the web]
Li Deng · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: an evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning, 2013
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling, 2014
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2014
Earlier work this paper cites.
Spatial training improves children’s mathematics ability
Yi-Ling Cheng and Kelly S Mix · 2014
Earlier work this paper cites.
Deconstructing building blocks: Preschoolers’ spatial assembly performance relates to early mathematical skills
Brian N. Verdine, Roberta M. Golinkoff, Kathryn Hirsh-Pasek, Nora S. Newcombe, Andrew T. Filipowicz, and Alicia Chang · 2014
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
Brenden M. Lake, Ruslan Salakhutdinov, and Joshua B. Tenenbaum · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge, 2015
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn, 2016
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Koray Kavukcuoglu, and Daan Wierstra · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks, 2017
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Johan Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Venkatesh Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra · 2022
Later among the works it cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale, 2023
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Kuang-Huei Lee, Sergey Levine, Yao Lu, Utsav Malla, Deeksha Manjunath, Igor Mordatch, Ofir Nachum, Carolina Parada, Jodilyn Peralta, Emily Perez, Karl Pertsch, Jornell Quiambao, Kanishka Rao, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Kevin Sayed, Jaspiar Singh, Sumedh Sontakke, Austin Stone, Clayton Tan, Huong Tran, Vincent Vanhoucke, Steve Vega, Quan Vuong, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Open-endedness: The last grand challenge you’ve never heard of
Kenneth Stanley · 2017
Cited alongside, same era.
Exploration: A study of count-based exploration for deep reinforcement learning, 2017
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2017
Cited alongside, same era.
Hindsight experience replay, 2018
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Exploration by random network distillation, 2018
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents, 2018
Carlos Florensa, David Held, Xinyang Geng, and Pieter Abbeel · 2018
Cited alongside, same era.
Later among the works it cites.
Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation, 2023
Minho Heo, Youngwoon Lee, Doohyun Lee, and Joseph J. Lim · 2023
Later among the works it cites.
A survey of zero-shot generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel · 2023
Later among the works it cites.
Pgx: Hardware-accelerated parallel game simulators for reinforcement learning
Sotetsu Koyamada, Shinri Okano, Soichiro Nishimori, Yu Murata, Keigo Habara, Haruka Kita, and Shin Ishii · 2023
Later among the works it cites.
Mastering the unsupervised reinforcement learning benchmark from pixels
Sai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Piché, Bart Dhoedt, Aaron Courville, and Alexandre Lacoste · 2023
Later among the works it cites.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao · 2023
Later among the works it cites.
Human-timescale adaptation in an open-ended task space, 2023
Adaptive Agent Team · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Karthik Valmeekam, Matthew Marquez, Alberto Olmo, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding, 2023
Hang Zhang, Xin Li, and Lidong Bing · 2023
Later among the works it cites.
Large language models for mathematical reasoning: Progresses and challenges, 2024
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin · 2024
Later among the works it cites.
Jumanji: a diverse suite of scalable reinforcement learning environments in JAX
Clément Bonnet, Daniel Luo, Donal John Byrne, Shikha Surana, Sasha Abramowitz, Paul Duckworth, Vincent Coyette, Laurence Illing Midgley, Elshadai Tegegn, Tristan Kalloniatis, Omayma Mahjoub, Matthew Macfarlane, Andries Petrus Smit, Nathan Grinsztajn, Raphael Boige, Cemlyn Neil Waters, Mohamed Ali Ali Mimouni, Ulrich Armel Mbou Sob, Ruan John de Kock, Siddarth Singh, Daniel Furelos-Blanco, Victor Le, Arnu Pretorius, and Alexandre Laterre · 2024
Later among the works it cites.
Mastering diverse domains through world models, 2024
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2024
Later among the works it cites.
Position: Open-endedness is essential for artificial superhuman intelligence
Edward Hughes, Michael D Dennis, Jack Parker-Holder, Feryal Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, and Tim Rocktäschel · 2024
Later among the works it cites.
Swe-bench: Can language models resolve real-world github issues?, 2024
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2024
Later among the works it cites.
Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control, 2024
Michal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Miłoś, and Marek Cygan · 2024
Later among the works it cites.
OpenAI · 2024
Later among the works it cites.
No regrets: Investigating and improving regret approximations for curriculum discovery
Alexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda, Nick Hawes, and Jakob Nicolaus Foerster · 2024
Later among the works it cites.
A definition of open-ended learning problems for goal-conditioned agents, 2024
Olivier Sigaud, Gianluca Baldassarre, Cedric Colas, Stephane Doncieux, Richard Duro, Pierre-Yves Oudeyer, Nicolas Perrin-Gilbert, and Vieri Giuliano Santucci · 2024
Later among the works it cites.
See and think: Embodied agent in virtual environment, 2024
Zhonghan Zhao, Wenhao Chai, Xuan Wang, Li Boyi, Shengyu Hao, Shidong Cao, Tian Ye, and Gaoang Wang · 2024
Later among the works it cites.
Physbench: Benchmarking and enhancing vision-language models for physical world understanding
Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita, Vitor Campagnolo Guizilini, and Yue Wang · 2025
Closest in time.
Deepseek-v3 technical report, 2025
DeepSeek-AI · 2025
Closest in time.
Leon Guertler, Bobby Cheng, Simon Yu, Bo Liu, Leshem Choshen, and Cheston Tan · 2025
Closest in time.
π 0.6 ∗ \pi^{*}_{0.6} : a vla that learns from experience, 2025
Physical Intelligence, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Kevin Black, Ken Conley, Grace Connors, James Darpinian, Karan Dhabalia, Jared DiCarlo, Danny Driess, Michael Equi, Adnan Esmail, Yunhao Fang, Chelsea Finn, Catherine Glossop, Thomas Godden, Ivan Goryachev, Lachy Groom, Hunter Hancock, Karol Hausman, Gashon Hussein, Brian Ichter, Szymon Jakubczak, Rowan Jen, Tim Jones, Ben Katz, Liyiming Ke, Chandra Kuchi, Marinda Lamb, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Yao Lu, Vishnu Mano, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Charvi Sharma, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, Will Stoeckle, Alex Swerdlow, James Tanner, Marcel Torne, Quan Vuong, Anna Walling, Haohuan Wang, Blake Williams, Sukwon Yoo, Lili Yu, Ury Zhilinsky, and Zhiyuan Zhou · 2025
Closest in time.
Kinetix: Investigating the training of general agents through open-ended physics-based control tasks
Michael Matthews, Michael Beukman, Chris Lu, and Jakob Nicolaus Foerster · 2025
Closest in time.
Balrog: Benchmarking agentic llm and vlm reasoning on games, 2025
Davide Paglieri, Bartłomiej Cupiał, Samuel Coward, Ulyana Piterbarg, Maciej Wolczyk, Akbir Khan, Eduardo Pignatelli, Łukasz Kuciński, Lerrel Pinto, Rob Fergus, Jakob Nicolaus Foerster, Jack Parker-Holder, and Tim Rocktäschel · 2025
Closest in time.
Ogbench: Benchmarking offline goal-conditioned rl, 2025
Seohong Park, Kevin Frans, Benjamin Eysenbach, and Sergey Levine · 2025
Closest in time.
Welcome to the era of experience
David Silver and Richard S Sutton · 2025
Closest in time.
Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025
Zafir Stojanovski, Oliver Stanley, Joe Sharratt, Richard Jones, Abdulhakeem Adefioye, Jean Kaddour, and Andreas Köpf · 2025
Closest in time.
Gemini: A family of highly capable multimodal models, 2025
Gemini Team · 2025
Closest in time.
Assessing adaptive world models in machines with novel games, 2025
Lance Ying, Katherine M. Collins, Prafull Sharma, Cedric Colas, Kaiya Ivy Zhao, Adrian Weller, Zenna Tavares, Phillip Isola, Samuel J. Gershman, Jacob D. Andreas, Thomas L. Griffiths, Francois Chollet, Kelsey R. Allen, and Joshua B. Tenenbaum · 2025
Closest in time.