Fetching the paper…
Reading the bibliography…
Language model (LM) agents are increasingly used as autonomous decision-makers which need to actively gather information to guide their decisions.
Expected information as expected utility
José M Bernardo · 1979
Earlier work this paper cites.
Causal learning mechanisms in very young children: Two-, three-, and four-year-olds infer causal relations from patterns of variation and covariation
Alison Gopnik, David M. Sobel, Laura E. Schulz, and Clark Glymour · 2001
Earlier work this paper cites.
Causal reasoning in rats
Aaron P Blaisdell, Kosuke Sawa, Kenneth J Leising, and Michael R Waldmann · 2006
Earlier work this paper cites.
Bayes nets and babies: Infants’ developing representations of causal knowledge
David M. Sobel and Natasha Z. Kirkham · 2007
Earlier work this paper cites.
Intuitive theories as grammars for causal inference
Joshua B Tenenbaum, Thomas L Griffiths, and Sourabh Niyogi · 2007
Earlier work this paper cites.
Theory-based causal induction
Thomas L Griffiths and Joshua B Tenenbaum · 2009
Earlier work this paper cites.
Just do it? investigating the gap between prediction and action in toddlers’ causal inferences
Elizabeth Baraff Bonawitz, David Ferranti, Rebecca Saxe, Alison Gopnik, Andrew N. Meltzoff, and James Woodward · 2010
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Children balance theories and evidence in exploration, explanation, and learning
Elizabeth B Bonawitz et al · 2012
Earlier work this paper cites.
Toddlers infer higher-order relational principles in causal learning
Caren M. Walker and Alison Gopnik · 2014
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences, 2017
Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Changes in cognitive flexibility and hypothesis search across human life history from childhood to adolescence to adulthood
Alison Gopnik, Shaun O’Grady, Christopher G Lucas, Thomas L Griffiths, Adrienne Wente, Sophie Bridgers, Rosie Aboody, Hoki Fung, and Ronald E Dahl · 2017
Cited alongside, same era.
Language models are few-shot learners
Tom Brown et al · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, et al · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck et al · 2023
Later among the works it cites.
Bridging the data gap between children and large language models
Michael C Frank · 2023
Later among the works it cites.
Thilo Hagendorff · 2023
Later among the works it cites.
Agentbench: Evaluating llms as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al · 2023
Later among the works it cites.
OpenAI · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamile Ndousse, Amanda Askell, Anna Chen, et al · 2022
Cited alongside, same era.
Language models show human-like content effects on reasoning tasks
Ishita Dasgupta, Andrew K Lampinen, Stephanie CY Chan, Hannah R Sheahan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Xi Ouyang, Ye Wu, and (etc.) · 2022
Cited alongside, same era.
Self-instruct: Aligning language model with self generated instructions, 2022
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei et al · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao et al · 2022
Cited alongside, same era.
Emergent autonomous scientific research capabilities of large language models
Daniil A. Boiko, Robert MacKnight, and Gabe Gomes · 2023
Cited alongside, same era.
Later among the works it cites.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Noah Shinn et al · 2023
Later among the works it cites.
The ai scientist: Towards fully automated open-ended scientific discovery
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha · 2024
Later among the works it cites.
Doing experiments and revising rules with natural language and probabilistic reasoning
Top Piriyakulkij, Cassidy Langenfeld, Tuan Anh Le, and Kevin Ellis · 2024
Later among the works it cites.
Modern bayesian experimental design
Tom Rainforth, Adam Foster, Desi R Ivanova, and Freddie Bickford Smith · 2024
Later among the works it cites.
The virtual lab: Ai agents design new sars-cov-2 nanobodies with experimental validation
Kyle Swanson, Wesley Wu, Nash L Bulaong, John E Pak, and James Zou · 2024
Later among the works it cites.
Ollama: Get up and running with llama 3.3, deepseek-r1, phi-4, gemma 3, and other large language models
Ollama · 2025
Closest in time.
Beyond semantics: The unreasonable effectiveness of reasonless intermediate tokens
Kaya Stechly, Karthik Valmeekam, Atharva Gundawar, Vardhan Palod, and Subbarao Kambhampati · 2025
Closest in time.