Fetching the paper…
Reading the bibliography…
Using AI to create autonomous researchers has the potential to accelerate scientific discovery.
On a measure of the information provided by an experiment
Dennis V Lindley · 1956
Earlier work this paper cites.
A Markovian Decision Process
Richard Bellman · 1957
Earlier work this paper cites.
Uncertainty, Information, and Sequential Experiments
MH DeGroot · 1962
Earlier work this paper cites.
Inductive Inference: Theory and Methods
Dana Angluin and Carl H Smith · 1983
Earlier work this paper cites.
Diversity-Based Inference of Finite Automata
Ronald L Rivest and Robert E Schapire · 1987
Earlier work this paper cites.
Queries and Concept Learning
Dana Angluin · 1988
Earlier work this paper cites.
Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm
Nick Littlestone · 1988
Earlier work this paper cites.
Inference of Finite Automata Using Homing Sequences
Ronald L Rivest and Robert E Schapire · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Active exploration in dynamic environments
Sebastian B Thrun and Knut Möller · 1991
Earlier work this paper cites.
ANOVA: Repeated measures
Ellen R Girden · 1992
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Bayesian Experimental Design: A Review
Kathryn Chaloner and Isabella Verdinelli · 1995
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
A Bayesian framework for reinforcement learning
Malcolm JA Strens · 2000
Earlier work this paper cites.
R-MAX – A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
A robust class of context-sensitive languages
Salvatore La Torre, Parthasarathy Madhusudan, and Gennaro Parlato · 2007
Earlier work this paper cites.
Knows What It Knows: A Framework for Self-Aware Learning
Lihong Li, Michael L Littman, and Thomas J Walsh · 2008
Earlier work this paper cites.
An Analysis of Model-based interval estimation for Markov Decision Processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Aleatory or Epistemic? Does it Matter?
Armen Der Kiureghian and Ove Ditlevsen · 2009
Earlier work this paper cites.
Active Learning Literature Survey
Burr Settles · 2009
Earlier work this paper cites.
Exploring Compact Reinforcement-Learning Representations with Linear Regression
Thomas J Walsh, István Szita, Carlos Diuk, and Michael L Littman · 2009
Earlier work this paper cites.
Reducing Reinforcement Learning to KWIK Online Regression
Lihong Li and Michael L Littman · 2010
Earlier work this paper cites.
Category learning through active sampling
Doug Markant and Todd Gureckis · 2010
Earlier work this paper cites.
Trading off Mistakes and Don’t-Know Predictions
Amin Sayedi, Morteza Zadimoghaddam, and Avrim Blum · 2010
Earlier work this paper cites.
PILCO: A Model-Based and Data-Efficient Approach to Policy Search
Marc Deisenroth and Carl E Rasmussen · 2011
Cited alongside, same era.
Agnostic KWIK learning and Efficient Approximate Reinforcement Learning
István Szita and Csaba Szepesvári · 2011
Cited alongside, same era.
Large-Scale Bandit Problems and KWIK Learning
Jacob Abernethy, Kareem Amin, Michael Kearns, and Moez Draief · 2013
Cited alongside, same era.
(More) Efficient Reinforcement Learning via Posterior Sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Amplify scientific discovery with artificial intelligence
Yolanda Gil, Mark Greaves, James Hendler, and Haym Hirsh · 2014
Cited alongside, same era.
Is it better to select or to receive? learning via active and passive hypothesis testing
Douglas B Markant and Todd M Gureckis · 2014
The ai scientist: Towards fully automated open-ended scientific discovery
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha · 2024
Later among the works it cites.
Embers of autoregression show how large language models are shaped by the problem they are trained to solve
R Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D Hardy, and Thomas L Griffiths · 2024
Later among the works it cites.
Matpilot: an llm-enabled ai materials scientist under the framework of human-machine collaboration
Ziqi Ni, Yahao Li, Kaijia Hu, Kunyuan Han, Ming Xu, Xingyu Chen, Fengqi Liu, Yicong Ye, and Shuxin Bai · 2024
Later among the works it cites.
Symbolic metaprogram search improves learning efficiency and explains rule learning in humans
Joshua S Rule, Steven T Piantadosi, Andrew Cropper, Kevin Ellis, Maxwell Nye, and Joshua B Tenenbaum · 2024
Later among the works it cites.
Can LLMs generate novel research ideas? A large-scale human study with 100+ NLP researchers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Variational inference: A review for statisticians
David M Blei, Alp Kucukelbir, and Jon D McAuliffe · 2017
Cited alongside, same era.
The first crank of the cultural ratchet: Learning and transmitting concepts through language
Sahil Chopra, Michael Henry Tessler, and Noah D Goodman · 2019
Cited alongside, same era.
Variational Bayesian optimal experimental design
Adam Foster, Martin Jankowiak, Elias Bingham, Paul Horsfall, Yee Whye Teh, Thomas Rainforth, and Noah Goodman · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Decoupling Exploration and Exploitation for Meta-Reinforcement Learning Without Sacrifices
Evan Z Liu, Aditi Raghunathan, Percy Liang, and Chelsea Finn · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit Bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2021
Cited alongside, same era.
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto · 2024
Later among the works it cites.
To CoT or not to CoT? chain-of-thought helps mainly on math and symbolic reasoning
Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Durrett · 2024
Later among the works it cites.
People use fast, goal-directed simulation to reason about novel games
Cedegao E Zhang, Katherine M Collins, Lionel Wong, Mauricio Barba, Adrian Weller, and Joshua B Tenenbaum · 2024
Later among the works it cites.
Toward Efficient Exploration by Large Language Model Agents
Dilip Arumugam and Thomas L Griffiths · 2025
Closest in time.
Why do multi-agent llm systems fail?
Mert Cemri, Melissa Z Pan, Shuyi Yang, Lakshya A Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, et al · 2025
Closest in time.
The danger of overthinking: Examining the reasoning-action dilemma in agentic tasks
Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, et al · 2025
Closest in time.
BoxingGym: Benchmarking progress in automated experimental design and model discovery
Kanishk Gandhi, Michael Y Li, Lyle Goodyear, Louise Li, Aditi Bhaskar, Mohammed Zaman, and Noah D Goodman · 2025
Closest in time.
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Kohli, Yossi Matias, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R D Costa, José R Penadés, Gary Peltz, Yunhan Xu, Annalisa Pawlosky, Alan Karthikesalingam, and Vivek Natarajan · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Can large language models detect errors in long chain-of-thought reasoning?
Yancheng He, Shilong Li, Jiaheng Liu, Weixun Wang, Xingyuan Bu, Ge Zhang, Zhongyuan Peng, Zhaoxiang Zhang, Wenbo Su, and Bo Zheng · 2025
Closest in time.
Alexander Ku, Declan Campbell, Xuechunzi Bai, Jiayi Geng, Ryan Liu, Raja Marjieh, R Thomas McCoy, Andrew Nam, Ilia Sucholutsky, Veniamin Veselovsky, et al · 2025
Closest in time.
Structured chain-of-thought prompting for code generation
Jia Li, Ge Li, Yongmin Li, and Zhi Jin · 2025
Closest in time.
Yijia Luo, Yulin Song, Xingyao Zhang, Jiaheng Liu, Weixun Wang, GengRu Chen, Wenbo Su, and Bo Zheng · 2025
Closest in time.
Sparks of science: Hypothesis generation using structured paper data
Charles O’Neill, Tirthankar Ghosal, Roberta Răileanu, Mike Walmsley, Thang Bui, Kevin Schawinski, and Ioana Ciucă · 2025
Closest in time.
Towards scientific discovery with generative ai: Progress, opportunities, and challenges
Chandan K Reddy and Parshin Shojaee · 2025
Closest in time.
Towards automation of cognitive modeling using large language models
Milena Rmus, Akshay K. Jagadish, Marvin Mathony, Tobias Ludwig, and Eric Schulz · 2025
Closest in time.
Agent laboratory: Using llm agents as research assistants
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum · 2025
Closest in time.
Llm-sr: Scientific equation discovery via programming with large language models
Parshin Shojaee, Kazem Meidani, Shashank Gupta, Amir Barati Farimani, and Chandan K Reddy · 2025
Closest in time.
Paperbench: Evaluating ai’s ability to replicate ai research
Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, et al · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Hanjie Chen, Xia Hu, et al · 2025
Closest in time.
Thoughts are all over the place: On the underthinking of o1-like llms
Yue Wang, Qiuzhi Liu, Jiahao Xu, Tian Liang, Xingyu Chen, Zhiwei He, Linfeng Song, Dian Yu, Juntao Li, Zhuosheng Zhang, et al · 2025
Closest in time.
On benchmarking human-like intelligence in machines
Lance Ying, Katherine M Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L Griffiths, and Joshua B Tenenbaum · 2025
Closest in time.