Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are catalyzing a paradigm shift in scientific discovery, evolving from task-specific automation tools into increasingly autonomous agents and fundamentally redefining research processes and human-AI collaboration.
Ai feynman: a physics-inspired method for symbolic regression
Silviu-Marian Udrescu and Max Tegmark. 2020 · 1905
Earlier work this paper cites.
The Logic of Scientific Discovery
Karl R. Popper. 1935 · 1935
Earlier work this paper cites.
The Structure of Scientific Revolutions
Thomas Samuel Kuhn. 1962 · 1962
Earlier work this paper cites.
End-to-end symbolic regression with transformers
Pierre-Alexandre Kamienny, Stéphane d’Ascoli, Guillaume Lample, and François Charton. 2022 · 2022
Earlier work this paper cites.
Ds-1000: A natural and reliable benchmark for data science code generation
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Scott Wen tau Yih, Daniel Fried, Sida Wang, and Tao Yu. 2022 · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022 · 2022
Earlier work this paper cites.
Natural language to code generation in interactive data science notebooks
Pengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Alex Polozov, and Charles Sutton. 2022 · 2022
Earlier work this paper cites.
Autonomous chemical research with large language models
Daniil A. Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. 2023 · 2023
Earlier work this paper cites.
Grounding large language models in interactive environments with online reinforcement learning
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer. 2023 · 2023
Earlier work this paper cites.
Teaching large language models to self-debug
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Earlier work this paper cites.
Ioana Ciucă, Yuan-Sen Ting, Sandor Kruk, and Kartheik Iyer. 2023 · 2023
Earlier work this paper cites.
Control risk for potential misuse of artificial intelligence in science
Jiyan He, Weitao Feng, Yaosen Min, Jingwei Yi, Kunsheng Tang, Shuai Li, Jie Zhang, Kejiang Chen, Wenbo Zhou, Xing Xie, Weiming Zhang, Nenghai Yu, and Shuxin Zheng. 2023 · 2023
Earlier work this paper cites.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang. 2023 · 2023
Earlier work this paper cites.
Paperqa: Retrieval-augmented generative agent for scientific research
Jakub Lála, Odhran O’Donoghue, Aleksandar Shtedritski, Sam Cox, Samuel G Rodriques, and Andrew D White. 2023 · 2023
Earlier work this paper cites.
Reviewergpt? an exploratory study on using large language models for paper reviewing
Ryan Liu and Nihar B. Shah. 2023 · 2023
Earlier work this paper cites.
Clin: A continually learning language agent for rapid task adaptation and generalization
Bodhisattwa Prasad Majumder, Bhavana Dalvi Mishra, Peter Jansen, Oyvind Tafjord, Niket Tandon, Li Zhang, Chris Callison-Burch, and Peter Clark. 2023 · 2023
Earlier work this paper cites.
Bioplanner: Automatic evaluation of llms on protocol planning in biology
Odhran O’Donoghue, Aleksandar Shtedritski, John Ginger, Ralph Abboud, Ali Essa Ghareeb, Justin Booth, and Samuel G Rodriques. 2023 · 2023
Earlier work this paper cites.
Large language models are zero shot hypothesis proposers
Biqing Qi, Kaiyan Zhang, Haoxiang Li, Kai Tian, Sihang Zeng, Zhang-Ren Chen, and Bowen Zhou. 2023 · 2023
Earlier work this paper cites.
Towards autonomous hypothesis verification via language models with minimal guidance
Shiro Takagi, Ryutaro Yamauchi, and Wataru Kumagai. 2023 · 2023
Earlier work this paper cites.
Large language models for chemistry robotics
Naruki Yoshikawa, Marta Skreta, Kourosh Darvish, Sebastian Arellano-Rubach, Zhi Ji, Lasse Bjørn Kristensen, Andrew Zou Li, Yuchi Zhao, Haoping Xu, Artur Kuramshin, Alán Aspuru-Guzik, Florian Shkurti, and Animesh Garg. 2023 · 2023
Earlier work this paper cites.
Large language models for scientific synthesis, inference and explanation
Yizhen Zheng, Huan Yee Koh, Jiaxin Ju, Anh T. N. Nguyen, Lauren T. May, Geoffrey I. Webb, and Shirui Pan. 2023 · 2023
Earlier work this paper cites.
Tejumade Afonja, Ivaxi Sheth, Ruta Binkyte, Waqar Hanif, Thomas Ulas, Matthias Becker, and Mario Fritz. 2024 · 2024
Earlier work this paper cites.
Litllm: A toolkit for scientific literature review
Shubham Agarwal, Gaurav Sahu, Abhay Puri, Issam H. Laradji, Krishnamurthy DJ Dvijotham, Jason Stanley, Laurent Charlin, and Christopher Pal. 2024 · 2024
Earlier work this paper cites.
Ai trustworthy challenges in drug discovery
Pegah Ahadian and Qiang Guan. 2024 · 2024
Earlier work this paper cites.
Automatikz: Text-guided synthesis of scientific vector graphics with tikz
Jonas Belouadi, Anne Lauscher, and Steffen Eger. 2024 · 2024
Earlier work this paper cites.
Generative adversarial reviews: When llms become the critic
Nicolas Bougie and Narimasa Watanabe. 2024 · 2024
Earlier work this paper cites.
Markus J. Buehler. 2024 · 2024
Earlier work this paper cites.
Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain
Fabio Dennstädt, Johannes Zink, Paul Martin Putora, Janna Hastings, and Nikola Cihoric. 2024 · 2024
Earlier work this paper cites.
Llms assist nlp researchers: Critique paper (meta-)reviewing
Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, and 21 others. 2024 · 2024
Earlier work this paper cites.
Empowering biomedical discovery with ai agents
Shanghua Gao, Ada Fang, Yepeng Huang, Valentina Giunchiglia, Ayush Noori, Jonathan Richard Schwarz, Yasha Ektefaie, Jovana Kondic, and Marinka Zitnik. 2024 · 2024
Earlier work this paper cites.
Blade: Benchmarking language model agents for data-driven science
Ken Gu, Ruoxi Shang, Ruien Jiang, Keying Kuang, Richard-John Lin, Donghe Lyu, Yue Mao, Youran Pan, Teng Wu, Jiaqian Yu, Yikun Zhang, Tianmai M. Zhang, Lanyi Zhu, Mike A. Merrill, Jeffrey Heer, and Tim Althoff. 2024 · 2024
Earlier work this paper cites.
Webvoyager: Building an end-to-end web agent with large multimodal models
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024 · 2024
Earlier work this paper cites.
Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark. 2024 · 2024
Earlier work this paper cites.
TKGT: Redefinition and a new way of text-to-table tasks based on real world demands and knowledge graphs augmented LLMs
Peiwen Jiang, Xinbo Lin, Zibo Zhao, Ruhui Ma, Yvonne Jie Chen, and Jinhua Cheng. 2024b · 2024
Earlier work this paper cites.
Dsbench: How far are data science agents from becoming data science experts?
Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao, Wenhao Yu, Kaixin Ma, Hongming Zhang, Xinya Du, and Dong Yu. 2024 · 2024
Cited alongside, same era.
Online continual learning for interactive instruction following agents
Byeonghwi Kim, Minhyuk Seo, and Jonghyun Choi. 2024 · 2024
Cited alongside, same era.
The ai scientist: Towards fully automated open-ended scientific discovery
Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024 · 2024
Cited alongside, same era.
From intention to implementation: Automating biomedical research via llms
Yi Luo, Linghang Shi, Yihao Li, Aobo Zhuang, Yeyun Gong, Ling Liu, and Chen Lin. 2024 · 2024
Cited alongside, same era.
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Kohli, and 15 others. 2025 · 2025
Closest in time.
Zochi technical report: The first artificial scientist
IntologyAI. 2025 · 2025
Closest in time.
Codescientist: End-to-end semi-automated scientific discovery with code-based experimentation
Peter Jansen, Oyvind Tafjord, Marissa Radensky, Pao Siangliulue, Tom Hope, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Daniel S. Weld, and Peter Clark. 2025 · 2025
Closest in time.
Aide: Ai-driven exploration in the space of code
Zhengyao Jiang, Dominik Schmidt, Dhruv Srikanth, Dixing Xu, Ian Kaplan, Deniss Jacenko, and Yuxiang Wu. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, Bhavana Dalvi Mishra, Abhijeetsingh Meena, Aryan Prakhar, Tirth Vora, Tushar Khot, Ashish Sabharwal, and Peter Clark. 2024 · 2024
Cited alongside, same era.
Arxivdigestables: Synthesizing scientific literature into tables using language models
Benjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue, Raymond Fok, Juho Kim, Daniel S. Weld, Joseph Chee Chang, and Kyle Lo. 2024 · 2024
Cited alongside, same era.
Machine learning for hypothesis generation in biology and medicine: exploring the latent space of neuroscience and developmental bioelectricity
Thomas O’Brien, Joel Stremmel, Léo Pio-Lopez, Patrick McMillen, Cody Rasmussen-Ivey, and Michael Levin. 2024 · 2024
Cited alongside, same era.
Introducing openai o1 preview
OpenAI. 2024 · 2024
Cited alongside, same era.
Large language models as biomedical hypothesis generators: A comprehensive evaluation
Biqing Qi, Kaiyan Zhang, Kai Tian, Haoxiang Li, Zhang-Ren Chen, Sihang Zeng, Ermo Hua, Hu Jinfang, and Bowen Zhou. 2024 · 2024
Cited alongside, same era.
Infobench: Evaluating instruction following ability in large language models
Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu. 2024 · 2024
Cited alongside, same era.
Verification and refinement of natural language explanations through llm-symbolic theorem proving
Xin Quan, Marco Valentino, Louise A. Dennis, and André Freitas. 2024 · 2024
Cited alongside, same era.
Automated machine learning: From principles to practices
Zhenqian Shen, Yongqi Zhang, Lanning Wei, Huan Zhao, and Quanming Yao. 2024 · 2024
Cited alongside, same era.
Nolan Koblischke, Hyunseok Jang, Kristen Menou, and Mohamad Ali-Dib. 2025 · 2025
Closest in time.
Curie: Toward rigorous and automated scientific experimentation with ai agents
Patrick Tser Jern Kon, Jiachen Liu, Qingyun Ding, Yuxin Qiu, Zhaoning Yang, Yufan Huang, Jer-Shannassa, Moontae Lee, Muntasir Chowdhury, and Aonan Chen. 2025 · 2025
Closest in time.
Can large language models help experimental design for causal discovery?
Junyi Li, Yongqiang Chen, Chenxi Liu, Qianyi Cai, Tongliang Liu, Bo Han, Kun Zhang, and Hui Xiong. 2025 · 2025
Closest in time.
Evosld: Automated neural scaling law discovery with large language models
Haowei Lin, Yacine Jernite, H. T. Kung, and Andrew Gordon Wilson. 2025 · 2025
Closest in time.
Llm4sr: A survey on large language models for scientific research
Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du. 2025 · 2025
Closest in time.
Mlgym: A new framework and benchmark for advancing ai research agents
Deepak Nathani, Lovish Madaan, Nicholas Roberts, Nikolay Bashlykov, Ajay Menon, Vincent Moens, Amar Budhiraja, Despoina Magka, Vladislav Vorotilov, Gaurav Chaurasia, Dieuwke Hupkes, Ricardo Silveira Cabral, Tatiana Shavrina, Jakob Foerster, Yoram Bachrach, William Yang Wang, and Roberta Raileanu. 2025 · 2025
Closest in time.
Sparks of science: Hypothesis generation using structured paper data
Charles O’Neill, Tirthankar Ghosal, Roberta Răileanu, Mike Walmsley, Thang Bui, Kevin Schawinski, and Ioana Ciucă. 2025 · 2025
Closest in time.
Introducing deep research
OpenAI. 2025 · 2025
Closest in time.
Claimcheck: How grounded are llm critiques of scientific papers?
Jiefu Ou, William Gantt Walden, Kate Sanders, Zhengping Jiang, Kaiser Sun, Jeffrey Cheng, William Jurayj, Miriam Wanner, Shaobo Liang, Candice Morgan, Seunghoon Han, Weiqi Wang, Chandler May, Hannah Recknor, Daniel Khashabi, and Benjamin Van Durme. 2025 · 2025
Closest in time.
Ai idea bench 2025: Ai research idea generation benchmark
Yansheng Qiu, Haoquan Zhang, Zhaopan Xu, Ming Li, Diping Song, Zheng Wang, and Kaipeng Zhang. 2025 · 2025
Closest in time.
Tool learning with large language models: a survey
Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-rong Wen. 2025 · 2025
Closest in time.
Gollam Rabby, Diyana Muhammed, Prasenjit Mitra, and Sören Auer. 2025 · 2025
Closest in time.
Scideator: Human-llm scientific idea generation grounded in research-paper facet recombination
Marissa Radensky, Simra Shahid, Raymond Fok, Pao Siangliulue, Tom Hope, and Daniel S. Weld. 2025 · 2025
Closest in time.
Towards scientific discovery with generative ai: Progress, opportunities, and challenges
Chandan K Reddy and Parshin Shojaee. 2025 · 2025
Closest in time.
Kai Ruan, Xuan Wang, Jixiang Hong, Peng Wang, Yang Liu, and Hao Sun. 2025 · 2025
Closest in time.
Agent laboratory: Using llm agents as research assistants
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum. 2025 · 2025
Closest in time.
Hierarchically encapsulated representation for protocol design in self-driving labs
Yu-Zhe Shi, Mingchen Liu, Fanxu Meng, Qiao Xu, Zhangqian Bi, Kun He, Lecheng Ruan, and Qining Wang. 2025 · 2025
Closest in time.
Futurehouse platform: Superintelligent ai agents for scientific discovery
Michael Skarlinski, Tyler Nadolski, James Braza, Remo Storni, Mayk Caldas, Ludovico Mitchener, Michaela Hinks, Andrew White, and Sam Rodriques. 2025 · 2025
Closest in time.
Paperbench: Evaluating ai’s ability to replicate ai research
Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, and Tejal Patwardhan. 2025 · 2025
Closest in time.
HypER: Literature-grounded hypothesis generation and distillation with provenance
Rosni Vasu, Chandrayee Basu, Bhavana Dalvi Mishra, Cristina Sarasua, Peter Clark, and Abraham Bernstein. 2025 · 2025
Closest in time.
Predicting empirical ai research outcomes with language models
Yutai Wen, Yifan Jiang, Zitao Li, Whytnee Wade, J. D. Zamfirescu-Pereira, Xinyun Chen, Yuxin Wen, Yiming Yang, Graham Neubig, Jacob Andreas, and Daniel S. Weld. 2025 · 2025
Closest in time.
Cycleresearcher: Improving automated research via automated review
Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. 2025 · 2025
Closest in time.
Tablebench: A comprehensive and complex benchmark for table question answering
Xianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang, Jiaheng Liu, Xinrun Du, Di Liang, Daixin Shu, Xianfu Cheng, Tianzhen Sun, Guanglin Niu, Tongliang Li, and Zhoujun Li. 2025 · 2025
Closest in time.
Grok 3 beta – the age of reasoning agents
xAI. 2025 · 2025
Closest in time.
Chartx & chartvlm: A versatile benchmark and foundation model for complicated chart reasoning
Renqiu Xia, Bo Zhang, Hancheng Ye, Xiangchao Yan, Qi Liu, Hongbin Zhou, Zijun Chen, Peng Ye, Min Dou, Botian Shi, Junchi Yan, and Yu Qiao. 2025 · 2025
Closest in time.
Scireplicate-bench: Benchmarking llms in agent-driven algorithmic reproduction from research papers
Yanzheng Xiang, Hanqi Yan, Shuyin Ouyang, Lin Gui, and Yulan He. 2025 · 2025
Closest in time.
Improve: Iterative model pipeline refinement and optimization leveraging llm agents
Eric Xue, Zeyi Huang, Yuyang Ji, and Haohan Wang. 2025 · 2025
Closest in time.
The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search
Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. 2025 · 2025
Closest in time.
Moose-chem: Large language models for rediscovering unseen chemistry scientific hypotheses
Zonglin Yang, Wanhao Liu, Ben Gao, Tong Xie, Yuqiang Li, Wanli Ouyang, Soujanya Poria, Erik Cambria, and Dongzhan Zhou. 2025 · 2025
Closest in time.
Text2chart31: Instruction tuning for chart generation with automatic feedback
Fatemeh Pesaran Zadeh, Juyeon Kim, Jin-Hwa Kim, and Gunhee Kim. 2025 · 2025
Closest in time.