Fetching the paper…
Reading the bibliography…
The emergence of large language models (LLMs) is propelling automated scientific discovery to the next level, with LLM-based Artificial Intelligence (AI) Scientist systems now taking the lead in scientific research.
The general theory of relativity
Albert Einstein · 1922
Earlier work this paper cites.
Undiscovered public knowledge
Don R Swanson · 1986
Earlier work this paper cites.
Scientific discovery: Computational explorations of the creative processes
P Langley · 1987
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Tldr: Extreme summarization of scientific documents
Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel S Weld · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Why most published research findings are false
John PA Ioannidis · 2005
Earlier work this paper cites.
The logic of scientific discovery
Karl Popper · 2005
Earlier work this paper cites.
Gödel machines: Fully self-referential optimal universal self-improvers
Jürgen Schmidhuber · 2007
Earlier work this paper cites.
The automation of science
Ross D King, Jem Rowland, Stephen G Oliver, Michael Young, Wayne Aubrey, Emma Byrne, Maria Liakata, Magdalena Markham, Pinar Pir, Larisa N Soldatova, et al · 2009
Earlier work this paper cites.
The toronto paper matching system: an automated paper-reviewer assignment system
Laurent Charlin and Richard Zemel · 2013
Earlier work this paper cites.
The history of science
Thomas Kuhn · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
1,500 scientists lift the lid on reproducibility, 2016
Monya Baker · 2016
Earlier work this paper cites.
Scientific literature: Information overload
Esther Landhuis · 2016
Earlier work this paper cites.
Figureqa: An annotated figure dataset for visual reasoning
Samira Ebrahimi Kahou, Adam Atkinson, Vincent Michalski, Ákos Kádár, Adam Trischler, and Yoshua Bengio · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Semantic scholar
Suzanne Fricke · 2018
Earlier work this paper cites.
A dataset of peer reviews (peerread): Collection, insights and nlp applications
Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine Van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz · 2018
Earlier work this paper cites.
Scibert: Pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan · 2019
Earlier work this paper cites.
Identification of tasks, datasets, evaluation metrics, and numeric scores for scientific leaderboards construction
Yufang Hou, Charles Jochim, Martin Gleize, Francesca Bonin, and Debasis Ganguly · 2019
Earlier work this paper cites.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath · 2019
Earlier work this paper cites.
Probing biomedical embeddings from language models
Qiao Jin, Bhuwan Dhingra, William Cohen, and Xinghua Lu · 2019
Earlier work this paper cites.
The global landscape of ai ethics guidelines
Anna Jobin, Marcello Ienca, and Effy Vayena · 2019
Earlier work this paper cites.
Does my rebuttal matter? insights from a major nlp conference
Yusuke Miyao · 2019
Earlier work this paper cites.
Bigpatent: A large-scale dataset for abstractive and coherent summarization
Eva Sharma, Chen Li, and Lu Wang · 2019
Earlier work this paper cites.
Peerreview4all: Fair and accurate reviewer assignment in peer review
Ivan Stelmakh, Nihar B Shah, and Aarti Singh · 2019
Earlier work this paper cites.
Unsupervised word embeddings capture latent knowledge from materials science literature
Vahe Tshitoyan, John Dagdelen, Leigh Weston, Alexander Dunn, Ziqin Rong, Olga Kononova, Kristin A Persson, Gerbrand Ceder, and Anubhav Jain · 2019
Earlier work this paper cites.
Paperrobot: Incremental draft generation of scientific ideas
Qingyun Wang, Lifu Huang, Zhiying Jiang, Kevin Knight, Heng Ji, Mohit Bansal, and Yi Luan · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
TLDR: Extreme summarization of scientific documents
Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel Weld · 2020
Earlier work this paper cites.
Domain-specific language model pretraining for biomedical natural language processing, 2020
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon · 2020
Earlier work this paper cites.
AxCell: Automatic extraction of results from machine learning papers
Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel, Ross Taylor, and Robert Stojnic · 2020
Earlier work this paper cites.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
BioMegatron: Larger biomedical domain language model
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, and Raghav Mani · 2020
Earlier work this paper cites.
Prior and prejudice
Ivan Stelmakh, Nihar B. Shah, Aarti Singh, and Hal Daum’e · 2020
Earlier work this paper cites.
Agatha: automatic graph mining and transformer based hypothesis generation approach
Justin Sybrandt, Ilya Tyagin, Michael Shtutman, and Ilya Safro · 2020
Earlier work this paper cites.
Reviewrobot: Explainable paper review generation based on knowledge synthesis
Qingyun Wang, Qi Zeng, Lifu Huang, Kevin Knight, Heng Ji, and Nazneen Fatema Rajani · 2020
Earlier work this paper cites.
BioM-transformers: Building large biomedical language models with BERT, ALBERT and ELECTRA
Sultan Alrowili and Vijay Shanker · 2021
Earlier work this paper cites.
React: A re view comment dataset for act ionability (and more)
Gautam Choudhary, Natwar Modani, and Nitish Maurya · 2021
Earlier work this paper cites.
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, et al · 2021
Earlier work this paper cites.
The perils of using mechanical turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer · 2021
Earlier work this paper cites.
Scigen: a dataset for reasoning-aware text generation from scientific tables
Nafise Sadat Moosavi, Andreas Rücklé, Dan Roth, and Iryna Gurevych · 2021
Earlier work this paper cites.
Literature-based discovery beyond the abc paradigm: a contrastive approach
Erwan Moreau, Orla Hardiman, Mark Heverin, and Declan O’sullivan · 2021
Earlier work this paper cites.
Towards fair, equitable, and efficient peer review
Ivan Stelmakh · 2021
Earlier work this paper cites.
Towards table-to-text generation with numerical reasoning
Lya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura, and Hiroya Takamura · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
An overview of artificial intelligence ethics
Changwu Huang, Zeqi Zhang, Bifei Mao, and Xin Yao · 2022
Earlier work this paper cites.
From who you know to what you read: Augmenting scientific recommendations with implicit social networks
Hyeonsu B Kang, Rafal Kocielnik, Andrew Head, Jiangjiang Yang, Matt Latzke, Aniket Kittur, Daniel S Weld, Doug Downey, and Jonathan Bragg · 2022
Earlier work this paper cites.
Ethics of ai: A systematic literature review of principles and challenges
Arif Ali Khan, Sher Badshah, Peng Liang, Muhammad Waseem, Bilal Khan, Aakash Ahmad, Mahdi Fahmideh, Mahmood Niazi, and Muhammad Azeem Akbar · 2022
Earlier work this paper cites.
Peersum: a peer review dataset for abstractive multi-document summarization
Miao Li, Jianzhong Qi, and Jey Han Lau · 2022
Earlier work this paper cites.
Moprd: A multidisciplinary open peer review dataset
Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen, and X. Shi · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
X-scitldr: cross-lingual extreme summarization of scholarly documents
Sotaro Takeshita, Tommaso Green, Niklas Friedrich, Kai Eckert, and Simone Paolo Ponzetto · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Large language models are better reasoners with self-verification
Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao · 2022
Earlier work this paper cites.
TELIN: Table entity LINker for extracting leaderboards from machine learning publications
Sean Yang, Chris Tensmeyer, and Curtis Wigington · 2022
Earlier work this paper cites.
Knowledge integration and decision support for accelerated discovery of antibiotic resistance genes
Jason Youn, Navneet Rai, and Ilias Tagkopoulos · 2022
Earlier work this paper cites.
Can we automate scientific reviewing?
Weizhe Yuan, Pengfei Liu, and Graham Neubig · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Earlier work this paper cites.
Investigating fairness disparities in peer review: A language model enhanced approach
Jiayao Zhang, Hongming Zhang, Zhun Deng, and Dan Roth · 2022
Earlier work this paper cites.
Alina Beygelzimer, Yann N Dauphin, Percy Liang, and Jennifer Wortman Vaughan · 2023
Earlier work this paper cites.
Chemcrow: Augmenting large-language models with chemistry tools
Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller · 2023
Cited alongside, same era.
Double-blind peer review affects reviewer ratings and editor decisions at an ecology journal
Charles W. Fox, Jennifer Meyer, and Emilie Aimé · 2023
Cited alongside, same era.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang · 2023
Cited alongside, same era.
Automatic analysis of substantiation in scientific peer reviews
Yanzhu Guo, Guokan Shang, Virgile Rennard, Michalis Vazirgiannis, and Chloé Clavel · 2023
Cited alongside, same era.
The diminishing returns of masked language models to science, 2023
Zhi Hong, Aswathy Ajith, Gregory Pauloski, Eamon Duede, Kyle Chard, and Ian Foster · 2023
Language agents achieve superhuman synthesis of scientific knowledge
Michael D. Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela Hinks, Michael J. Hammerling, Manvitha Ponnapati, Samuel G. Rodriques, and Andrew D. White · 2024
Later among the works it cites.
Metawriter: Exploring the potential and perils of ai writing support in scientific peer review
Lu Sun, Stone Tao, Junjie Hu, and Steven P. Dow · 2024
Later among the works it cites.
Peer review as a multi-turn and long-context dialogue with role-based interactions
Cheng Tan, Dongxin Lyu, Siyuan Li, Zhangyang Gao, Jingxuan Wei, Siqi Ma, Zicheng Liu, and Stan Z Li · 2024
Later among the works it cites.
Scicode: A research coding benchmark curated by scientists
Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, Shengyan Liu, Di Luo, Yutao Ma, Hao Tong, Kha Trinh, Chenyu Tian, Zihan Wang, Bohao Wu, Yanyu Xiong, Shengzhu Yin, Min Zhu, Kilian Adriano Lieret, Yanxin Lu, Genglin Liu, Yufeng Du, Tianhua Tao, Ofir Press, Jamie Callan, E. A. Huerta, and Hao Peng · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Cited alongside, same era.
Orkg-leaderboards: a systematic workflow for mining leaderboards as a knowledge graph
Salomon Kabongo KABENAMUALU, Jennifer D’Souza, and S. Auer · 2023
Cited alongside, same era.
Comlittee: Literature discovery with personal elected author committees
Hyeonsu B Kang, Nouran Soliman, Matt Latzke, Joseph Chee Chang, and Jonathan Bragg · 2023
Cited alongside, same era.
Longeval: Guidelines for human evaluation of faithfulness in long-form summarization
Kalpesh Krishna, Erin Bransom, Bailey Kuehl, Mohit Iyyer, Pradeep Dasigi, Arman Cohan, and Kyle Lo · 2023
Cited alongside, same era.
Paperqa: Retrieval-augmented generative agent for scientific research
Jakub Lála, Odhran O’Donoghue, Aleksandar Shtedritski, Sam Cox, Samuel G Rodriques, and Andrew D White · 2023
Cited alongside, same era.
All data on the table: Novel dataset and benchmark for cross-modality scientific information extraction
Yuhan Li, Jian Wu, Zhiwei Yu, Börje F. Karlsson, Wei Shen, Manabu Okumura, and Chin-Yew Lin · 2023
Cited alongside, same era.
Moprd: A multidisciplinary open peer review dataset
Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen, and Xiaodong Shi · 2023
Cited alongside, same era.
Later among the works it cites.
A new sociology of humans and machines
Milena Tsvetkova, Taha Yasseri, Niccolo Pescetelli, and Tobias Werner · 2024
Later among the works it cites.
Openreviewer: Mitigating challenges in LLM reviewing, 2024
Keith Tyser, Jason Lee, Avi Shporer, Madeleine Udell, Dov Te’eni, and Iddo Drori · 2024
Later among the works it cites.
SciMON: Scientific inspiration machines optimized for novelty
Qingyun Wang, Doug Downey, Heng Ji, and Tom Hope · 2024
Later among the works it cites.
Automated review generation method based on large language models
Shican Wu, Xiao Ma, Dehui Luo, Lulu Li, Xiangcheng Shi, Xin Chang, Xiaoyun Lin, Ran Luo, Chunlei Pei, Changying Du, et al · 2024
Later among the works it cites.
Large language models for automated open-domain scientific hypotheses discovery
Zonglin Yang, Xinya Du, Junxian Li, Jie Zheng, Soujanya Poria, and Erik Cambria · 2024
Later among the works it cites.
Are we there yet? revealing the risks of utilizing large language models in scholarly peer review
Rui Ye, Xianghe Pang, Jingyi Chai, Jiaao Chen, Zhenfei Yin, Zhen Xiang, Xiaowen Dong, Jing Shao, and Siheng Chen · 2024
Later among the works it cites.
Mcx-llm: an experiment in bridging natural language problem descriptions with quantitative scientific simulations
Fan-Yu Yen and Qianqian Fang · 2024
Later among the works it cites.
Researchtown: Simulator of human research community
Haofei Yu, Zhaochen Hong, Zirui Cheng, Kunlun Zhu, Keyang Xuan, Jinwei Yao, Tao Feng, and Jiaxuan You · 2024
Later among the works it cites.
Is LLM a reliable reviewer? a comprehensive evaluation of LLM on automatic paper reviewing tasks
Ruiyang Zhou, Lu Chen, and Kai Yu · 2024
Later among the works it cites.
Peerqa: A scientific question answering dataset from peer reviews
Tim Baumgärtner, Ted Briscoe, and Iryna Gurevych · 2025
Closest in time.
Joeran Beel, Min-Yen Kan, and Moritz Baumgart · 2025
Closest in time.
Superintelligent agents pose catastrophic risks: Can scientist ai offer a safer path?
Yoshua Bengio, Michael Cohen, Damiano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott, Sören Mindermann, Adam Oberman, Jesse Richardson, Oliver Richardson, et al · 2025
Closest in time.
Reviewagents: Bridging the gap between human and ai-generated paper reviews
Xian Gao, Jiacheng Ruan, Jingsheng Gao, Ting Liu, and Yuzhuo Fu · 2025
Closest in time.
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al · 2025
Closest in time.
A survey on the rise of the ai scientists: Accelerating discovery and confronting ethical frontiers
Muskaan Goyal · 2025
Closest in time.
Agentic ai for scientific discovery: A survey of progress, challenges, and future directions
Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack · 2025
Closest in time.
Pasa: An llm agent for comprehensive academic paper search, 2025
Yichen He, Guanhua Huang, Peiyuan Feng, Yuan Lin, Yuchen Zhang, Hang Li, and Weinan E · 2025
Closest in time.
Automatic evaluation metrics for artificially generated scientific research
Niklas Höpner, Leon Eshuijs, Dimitrios Alivanistos, Giacomo Zamprogno, and Ilaria Tiddi · 2025
Closest in time.
Biomni: A general-purpose biomedical ai agent
Kexin Huang, Serena Zhang, Hanchen Wang, Yuanhao Qu, Yingzhou Lu, Yusuf Roohani, Ryan Li, Lin Qiu, Junze Zhang, Yin Di, et al · 2025
Closest in time.
Zochi technical report
Intology · 2025
Closest in time.
Codescientist: End-to-end semi-automated scientific discovery with code-based experimentation
Peter Jansen, Oyvind Tafjord, Marissa Radensky, Pao Siangliulue, Tom Hope, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Daniel S Weld, and Peter Clark · 2025
Closest in time.
Lgar: Zero-shot llm-guided neural ranking for abstract screening in systematic literature reviews
Christian Jaumann, Andreas Wiedholz, and Annemarie Friedrich · 2025
Closest in time.
Ai-researcher: Autonomous scientific innovation, 2025
Tang Jiabin, Xia Lianghao, Li Zhonghang, and Huang Chao · 2025
Closest in time.
Aide: Ai-driven exploration in the space of code
Zhengyao Jiang, Dominik Schmidt, Dhruv Srikanth, Dixing Xu, Ian Kaplan, Deniss Jacenko, and Yuxiang Wu · 2025
Closest in time.
Dsbench: How far are data science agents from becoming data science experts?, 2025
Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao, Wenhao Yu, Kaixin Ma, Hongming Zhang, Xinya Du, and Dong Yu · 2025
Closest in time.
PlagBench: Exploring the duality of large language models in plagiarism generation and detection
Jooyoung Lee, Toshini Agrawal, Adaku Uchendu, Thai Le, Jinghui Chen, and Dongwon Lee · 2025
Closest in time.
MMSci: A dataset for graduate-level multi-discipline multimodal scientific understanding, 2025
Zekun Li, Xianjun Yang, Kyuri Choi, Wanrong Zhu, Ryan Hsieh, HyeonJung Kim, Jin Hyuk Lim, Sungyoung Ji, Byungju Lee, Xifeng Yan, Linda Ruth Petzold, Stephen D. Wilson, Woosang Lim, and William Yang Wang · 2025
Closest in time.
Surveyx: Academic survey automation via large language models
Xun Liang, Jiawei Yang, Yezhaohui Wang, Chen Tang, Zifan Zheng, Shichao Song, Zehao Lin, Yebin Yang, Simin Niu, Hanyu Wang, et al · 2025
Closest in time.
Zijie Lin, Yiqing Shen, Qilin Cai, He Sun, Jinrui Zhou, and Mingjun Xiao · 2025
Closest in time.
Aaar-1.0: Assessing ai’s potential to assist research, 2025
Renze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Jian Xie, Yuxuan Sun, Yusen Zhang, Jihyun Janice Ahn, Hongchao Fang, Zhuoyang Zou, Wenchao Ma, Xi Li, Kai Zhang, Congying Xia, Lifu Huang, and Wenpeng Yin · 2025
Closest in time.
Llm4sr: A survey on large language models for scientific research
Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du · 2025
Closest in time.
Dora ai scientist: Multi-agent virtual research team for scientific exploration discovery and automated report generation
Vladimir Naumov, Diana Zagirova, Sha Lin, Yupeng Xie, Wenhao Gou, Anatoly Urban, Nina Tikhonova, Khadija Alawi, Mike Durymanov, Fedor Galkin, et al · 2025
Closest in time.
Alphaevolve: A coding agent for scientific and algorithmic discovery
Alexander Novikov, Ngân Vu, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al · 2025
Closest in time.
Ml-dev-bench: Comparative analysis of ai agents on ml development workflows, 2025
Harshith Padigela, Chintan Shah, and Dinkar Juyal · 2025
Closest in time.
Jungsoo Park, Junmo Kang, Gabriel Stanovsky, and Alan Ritter · 2025
Closest in time.
Automated research review support using machine learning, large language models, and natural language processing
Vishnu S Pendyala, Karnavee Kamdar, and Kapil Mulchandani · 2025
Closest in time.
Piflow: Principle-aware scientific discovery with multi-agent collaboration, 2025
Yingming Pu, Tao Lin, and Hongyu Chen · 2025
Closest in time.
Towards scientific intelligence: A survey of llm-based scientific agents
Shuo Ren, Pu Jian, Zhenjiang Ren, Chunlin Leng, Can Xie, and Jiajun Zhang · 2025
Closest in time.
Agentrxiv: Towards collaborative autonomous research
Samuel Schmidgall and Michael Moor · 2025
Closest in time.
Agent laboratory: Using llm agents as research assistants
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum · 2025
Closest in time.
Shortcutsbench: A large-scale real-world benchmark for api-based agents, 2025
Haiyang Shen, Yue Li, Desong Meng, Dongqi Cai, Sheng Qi, Li Zhang, Mengwei Xu, and Yun Ma · 2025
Closest in time.
The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity
Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar · 2025
Closest in time.
Paperbench: Evaluating ai’s ability to replicate ai research
Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, et al · 2025
Closest in time.
Reviewriter: Ai-generated instructions for peer review writing
Xiaotian Su, Thiemo Wambsganss, Roman Rietsche, Seyed Parsa Neshaei, and Tanja Käser · 2025
Closest in time.
Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents
Peisong Wang, Ruotian Ma, Bang Zhang, Xingyu Chen, Zhiwei He, Kang Luo, Qingsong Lv, Qingxuan Jiang, Zheng Xie, Shanyi Wang, et al · 2025
Closest in time.
Browsecomp: A simple yet challenging benchmark for browsing agents
Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese · 2025
Closest in time.
Cycleresearcher: Improving automated research via automated review
Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang · 2025
Closest in time.
Lag: Llm agents for leaderboard auto generation on demanding
Jian Wu, Jiayu Zhang, Dongyuan Li, Linyi Yang, Aoxiao Zhong, Renhe Jiang, Qingsong Wen, and Yue Zhang · 2025
Closest in time.
Scireplicate-bench: Benchmarking llms in agent-driven algorithmic reproduction from research papers
Yanzheng Xiang, Hanqi Yan, Shuyin Ouyang, Lin Gui, and Yulan He · 2025
Closest in time.
An empirical analysis of uncertainty in large language model evaluations
Qiujie Xie, Qingqiu Li, Zhuohao Yu, Yuejie Zhang, Yue Zhang, and Linyi Yang · 2025
Closest in time.
The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search
Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha · 2025
Closest in time.
Extracting knowledge from scientific texts on patient-derived cancer models using large language models: algorithm development and validation
Jiarui Yao, Zinaida Perova, Tushar Mandloi, Elizabeth Lewis, Helen Parkinson, and Guergana Savova · 2025
Closest in time.
Sungduk Yu, Man Luo, Avinash Madusu, Vasudev Lal, and Phillip Howard · 2025
Closest in time.
Dolphin: Closed-loop open-ended auto-research through thinking, practice, and feedback
Jiakang Yuan, Xiangchao Yan, Botian Shi, Tao Chen, Wanli Ouyang, Bo Zhang, Lei Bai, Yu Qiao, and Bowen Zhou · 2025
Closest in time.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al · 2025
Closest in time.
From automation to autonomy: A survey on large language models in scientific discovery
Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, and Yangqiu Song · 2025
Closest in time.
Deepreview: Improving llm-based paper review with human-like deep thinking process
Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang · 2025
Closest in time.
Large language models for automated scholarly paper review: A survey
Zhenzhen Zhuang, Jiandong Chen, Hongfeng Xu, Yuwen Jiang, and Jialiang Lin · 2025
Closest in time.
MatSciBERT: A materials domain language model for text mining and information extraction
Tanishq Gupta, Mohd Zaki, N. M. Anoop Krishnan, and Mausam · 2057
Closest in time.