Fetching the paper…
Reading the bibliography…
Recent advances in the intrinsic reasoning capabilities of large language models (LLMs) have given rise to LLM-based agent systems that exhibit near-human performance on a variety of automated tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Society of mind
Marvin Minsky. 1986 · 1986
Earlier work this paper cites.
Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings
Christopher A Lipinski, Franco Lombardo, Beryl W Dominy, and Paul J Feeney. 1997 · 1997
Earlier work this paper cites.
Pitfalls of agent-oriented development. In Proceedings of the second international conference on Autonomous agents . 385–391
Michael Wooldridge and Nicholas R Jennings. 1998 · 1998
Earlier work this paper cites.
The protein data bank
Helen M Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N Bhat, Helge Weissig, Ilya N Shindyalov, and Philip E Bourne. 2000 · 2000
Earlier work this paper cites.
Manifesto for Agile Software Development
Agile Manifesto. 2001 · 2001
Earlier work this paper cites.
Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions
Peter Ertl and Ansgar Schuffenhauer. 2009 · 2009
Earlier work this paper cites.
DECIPHER: database of chromosomal imbalance and phenotype in humans using ensembl resources
Helen V Firth, Shola M Richards, A Paul Bevan, Stephen Clayton, Manuel Corpas, Diana Rajan, Steven Van Vooren, Yves Moreau, Roger M Pettett, and Nigel P Carter. 2009 · 2009
Earlier work this paper cites.
Computer-aided prediction of rodent carcinogenicity by PASS and CISOC-PSCT
Alexey Lagunin, Dmitrii Filimonov, Alexey Zakharov, Wei Xie, Ying Huang, Fucheng Zhu, Tianxiang Shen, Jianhua Yao, and Vladimir Poroikov. 2009 · 2009
Earlier work this paper cites.
The role of play in human development
Anthony D Pellegrini. 2009 · 2009
Earlier work this paper cites.
AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading
Oleg Trott and Arthur J Olson. 2010 · 2010
Earlier work this paper cites.
Comprehensive assay of kinase catalytic activity reveals features of kinase inhibitor selectivity
Theonie Anastassiadis, Sean W Deacon, Karthik Devarajan, Haiching Ma, and Jeffrey R Peterson. 2011 · 2011
Earlier work this paper cites.
Does reflection lead to wise choices?
Lisa Bortolotti. 2011 · 2011
Earlier work this paper cites.
Team roles at work
RM Belbin and V Brown. 2012 · 2012
Earlier work this paper cites.
Quantifying the chemical beauty of drugs
G Richard Bickerton, Gaia V Paolini, Jérémy Besnard, Sorel Muresan, and Andrew L Hopkins. 2012 · 2012
Earlier work this paper cites.
An integrated map of genetic variation from 1,092 human genomes
1000 Genomes Project Consortium et al · 2012
Earlier work this paper cites.
ZINC: a free tool to discover chemistry for biology
John J Irwin, Teague Sterling, Michael M Mysinger, Erin S Bolstad, and Ryan G Coleman. 2012 · 2012
Earlier work this paper cites.
Peopleware: productive projects and teams
Tom DeMarco and Tim Lister. 2013 · 2013
Earlier work this paper cites.
RYBP and Cbx7 define specific biological functions of polycomb complexes in mouse embryonic stem cells
Lluis Morey, Luigi Aloia, Luca Cozzuto, Salvador Aznar Benitah, and Luciano Di Croce. 2013 · 2013
Earlier work this paper cites.
The construction of reality in the child
Jean Piaget. 2013 · 2013
Earlier work this paper cites.
The ChEMBL bioactivity database: an update
A Patrícia Bento, Anna Gaulton, Anne Hersey, Louisa J Bellis, Jon Chambers, Mark Davies, Felix A Krüger, Yvonne Light, Lora Mak, Shaun McGlinchey, et al · 2014
Earlier work this paper cites.
Defects4J: A database of existing faults to enable controlled testing studies for Java programs. In Proceedings of the 2014 international symposium on software testing and analysis . 437–440
René Just, Darioush Jalali, and Michael D Ernst. 2014 · 2014
Earlier work this paper cites.
A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping
Suhas SP Rao, Miriam H Huntley, Neva C Durand, Elena K Stamenova, Ivan D Bochkov, James T Robinson, Adrian L Sanborn, Ido Machol, Arina D Omer, Eric S Lander, et al · 2014
Earlier work this paper cites.
Predicting chemically-induced skin reactions. Part I: QSAR models of skin sensitization and their application to identify potentially hazardous compounds
Vinicius M Alves, Eugene Muratov, Denis Fourches, Judy Strickland, Nicole Kleinstreuer, Carolina H Andrade, and Alexander Tropsha. 2015 · 2015
Earlier work this paper cites.
LMC: large model collaboration with cross-assessment for training-free open-set object recognition. In Proceedings of the 37th International Conference on Neural Information Processing Systems . Red Hook, NY, USA, Article 2016, 14 pages
Haoxuan Qu, Xiaofei Hui, Yujun Cai, and Jun Liu. 2023 · 2016
Earlier work this paper cites.
Artificial intelligence: a modern approach
Stuart J Russell and Peter Norvig. 2016 · 2016
Earlier work this paper cites.
STITCH 5: augmenting protein–chemical interaction networks with tissue and affinity data
Damian Szklarczyk, Alberto Santos, Christian Von Mering, Lars Juhl Jensen, Peer Bork, and Michael Kuhn. 2016 · 2016
Earlier work this paper cites.
ADMET evaluation in drug discovery. 16. Predicting hERG blockers by combining multiple pharmacophores and machine learning approaches
Shuangquan Wang, Huiyong Sun, Hui Liu, Dan Li, Youyong Li, and Tingjun Hou. 2016 · 2016
Earlier work this paper cites.
The Drug Repurposing Hub: a next-generation drug library and information resource
Steven M Corsello, Joshua A Bittker, Zihan Liu, Joshua Gould, Patrick McCarren, Jodi E Hirschman, Stephen E Johnston, Anita Vrcic, Bang Wong, Mariya Khan, et al · 2017
Earlier work this paper cites.
Systematic integration of biomedical knowledge prioritizes drugs for repurposing
Daniel Scott Himmelstein, Antoine Lizee, Christine Hessler, Leo Brueggeman, Sabrina L Chen, Dexter Hadley, Ari Green, Pouya Khankhanian, and Sergio E Baranzini. 2017 · 2017
Earlier work this paper cites.
Predicting organic reaction outcomes with weisfeiler-lehman network
Wengong Jin, Connor Coley, Regina Barzilay, and Tommi Jaakkola. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1601–1611
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Metacognition and Reflection by Interdisciplinary Experts: Insights from Cognitive Science and Philosophy
M Keestra et al · 2017
Earlier work this paper cites.
Mapping the genetic landscape of human cells
Max A Horlbeck, Albert Xu, Min Wang, Neal K Bennett, Chong Y Park, Derek Bogdanoff, Britt Adamson, Eric D Chow, Martin Kampmann, Tim R Peterson, et al · 2018
Earlier work this paper cites.
A dataset of clinically generated visual questions and answers about radiology images
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. 2018 · 2018
Earlier work this paper cites.
DrugBank 5.0: a major update to the DrugBank database for 2018
David S Wishart, Yannick D Feunang, An C Guo, Elvis J Lo, Ana Marcu, Jason R Grant, Tanvir Sajed, Daniel Johnson, Carin Li, Zinat Sayeeda, et al · 2018
Earlier work this paper cites.
MoleculeNet: a benchmark for molecular machine learning
Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. 2018 · 2018
Earlier work this paper cites.
Compendium of China’s First List of Rare Disease
S Zhang. 2018 · 2018
Earlier work this paper cites.
PubMedQA: A Dataset for Biomedical Research Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . 2567–2577
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019 · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Evaluating protein transfer learning with TAPE
Roshan Rao, Nicholas Bhattacharya, Neil Thomas, Yan Duan, Peter Chen, John Canny, Pieter Abbeel, and Yun Song. 2019 · 2019
Earlier work this paper cites.
Three-dimensional convolutional neural networks and a cross-docked data set for structure-based drug design
Paul G Francoeur, Tomohide Masuda, Jocelyn Sunseri, Andrew Jia, Richard B Iovanisci, Ian Snyder, and David R Koes. 2020 · 2020
Earlier work this paper cites.
Pathvqa: 30000+ questions for medical visual question answering
Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. 2020 · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
Unsupervised Evaluation of Interactive Dialog with DialoGPT. In Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue . 225–235
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Earlier work this paper cites.
MHCflurry 2.0: improved pan-allele prediction of MHC class I-presented peptides by incorporating antigen processing
Timothy J O’Donnell, Alex Rubinsteyn, and Uri Laserson. 2020 · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021 · 2021
Earlier work this paper cites.
Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 2021 IEEE 18th international symposium on biomedical imaging (ISBI) . IEEE, 1650–1654
Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. 2021 · 2021
Earlier work this paper cites.
Genome-wide CRISPR screen identifies protein pathways modulating tau protein levels in neurons
Carlos G Sanchez, Christopher M Acker, Audrey Gray, Malini Varadarajan, Cheng Song, Nadire R Cochran, Steven Paula, Alicia Lindeman, Shaojian An, Gregory McAllister, et al · 2021
Earlier work this paper cites.
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning. In International Conference on Learning Representations
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Cote, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2021 · 2021
Earlier work this paper cites.
RASA2 ablation in T cells boosts antigen sensitivity and long-term function
Julia Carnevale, Eric Shifrut, Nupura Kale, William A Nyberg, Franziska Blaeschke, Yan Yi Chen, Zhongmei Li, Sagar P Bapat, Morgan E Diolaiti, Patrick O’Leary, et al · 2022
Earlier work this paper cites.
Huihui Fang, Fei Li, Junde Wu, Huazhu Fu, Xu Sun, Jaemin Son, Shuang Yu, Menglu Zhang, Chenglang Yuan, Cheng Bian, et al · 2022
Earlier work this paper cites.
Ddxplus: A new dataset for automatic medical diagnosis
Arsene Fansi Tchango, Rishab Goel, Zhi Wen, Julien Martel, and Joumana Ghosn. 2022 · 2022
Earlier work this paper cites.
Multi-agent deep reinforcement learning: a survey
Sven Gronauer and Klaus Diepold. 2022 · 2022
Earlier work this paper cites.
Eliciting thinking hierarchy without a prior
Yuqing Kong, Yunqi Li, Yubo Zhang, Zhihuan Huang, and Jinzhao Wu. 2022 · 2022
Earlier work this paper cites.
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 11048–11064
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning . PMLR, 248–260
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022 · 2022
Earlier work this paper cites.
CRISPR activation and interference screens decode stimulation responses in primary human T cells
Ralf Schmidt, Zachary Steinhart, Madeline Layeghi, Jacob W Freimer, Raymund Bueno, Vinh Q Nguyen, Franziska Blaeschke, Chun Jimmie Ye, and Alexander Marson. 2022 · 2022
Earlier work this paper cites.
Perception and navigation in autonomous systems in the era of learning: A survey
Yang Tang, Chaoqiang Zhao, Jianrui Wang, Chongzhen Zhang, Qiyu Sun, Wei Xing Zheng, Wenli Du, Feng Qian, and Juergen Kurths. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Nlice: Synthetic medical record generation for effective primary healthcare differential diagnosis. In 2023 IEEE 23rd International Conference on Bioinformatics and Bioengineering (BIBE) . IEEE, 397–402
Zaid Al-Ars, Obinna Agba, Zhuoran Guo, Christiaan Boerkamp, Ziyaad Jaber, and Tareq Jaber. 2023 · 2023
Earlier work this paper cites.
Ehrxqa: A multi-modal question answering dataset for electronic health records with chest x-ray images
Seongsu Bae, Daeun Kyung, Jaehee Ryu, Eunbyeol Cho, Gyubok Lee, Sunjun Kweon, Jungwoo Oh, Lei Ji, Eric Chang, Tackeun Kim, et al · 2023
Earlier work this paper cites.
Chemcrow: Augmenting large-language models with chemistry tools
Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. 2023 · 2023
Earlier work this paper cites.
Kexin Chen, Junyou Li, Kunyi Wang, Yuyang Du, Jiahui Yu, Jiamin Lu, Lanqing Li, Jiezhong Qiu, Jianzhang Pan, Yi Huang, et al · 2023
Earlier work this paper cites.
S3: Social-network Simulation System with Large Language Model-Empowered Agents
Chen Gao, Xiaochong Lan, Zhihong Lu, Jinzhu Mao, Jinghua Piao, Huandong Wang, Depeng Jin, and Yong Li. 2023a · 2023
Earlier work this paper cites.
Enabling Large Language Models to Generate Text with Citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 6465–6488
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023c · 2023
Earlier work this paper cites.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023b · 2023
Earlier work this paper cites.
What can large language models do in chemistry? a comprehensive benchmark on eight tasks
Taicheng Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh Chawla, Olaf Wiest, Xiangliang Zhang, et al · 2023
Earlier work this paper cites.
A dataset for medical instructional video classification and question answering
Deepak Gupta, Kush Attal, and Dina Demner-Fushman. 2023 · 2023
Earlier work this paper cites.
Fid-light: Efficient and effective retrieval-augmented text generation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1437–1447
Sebastian Hofstätter, Jiecao Chen, Karthik Raman, and Hamed Zamani. 2023 · 2023
Earlier work this paper cites.
diseaseGPS: auxiliary diagnostic system for genetic disorders based on genotype and phenotype
Daoyi Huang, Jianping Jiang, Tingting Zhao, Shengnan Wu, Pin Li, Yongfen Lyu, Jincai Feng, Mingyue Wei, Zhixing Zhu, Jianlei Gu, et al · 2023
Earlier work this paper cites.
Agentcoder: Multi-agent-based code generation with iterative testing and optimisation
Dong Huang, Jie M Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui. 2023b · 2023
Earlier work this paper cites.
MIMIC-IV, a freely accessible electronic health record dataset
Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al · 2023
Earlier work this paper cites.
DS-1000: A natural and reliable benchmark for data science code generation. In International Conference on Machine Learning . PMLR, 18319–18345
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-tau Yih, Daniel Fried, Sida Wang, and Tao Yu. 2023 · 2023
Earlier work this paper cites.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2023b · 2023
Earlier work this paper cites.
Camel: Communicative agents for" mind" exploration of large language model society
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023a · 2023
Earlier work this paper cites.
Finding Support Examples for In-Context Learning. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 6219–6235
Xiaonan Li and Xipeng Qiu. 2023 · 2023
Earlier work this paper cites.
Yuan Li, Yixuan Zhang, and Lichao Sun. 2023c · 2023
Earlier work this paper cites.
Autonomous GIS: the next-generation AI-powered GIS
Zhenlong Li and Huan Ning. 2023 · 2023
Earlier work this paper cites.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023a · 2023
Earlier work this paper cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023b · 2023
Earlier work this paper cites.
GitAgent: Facilitating Autonomous Agent with GitHub by Tool Extension
Bohan Lyu, Xin Cong, Heyang Yu, Pan Yang, Yujia Qin, Yining Ye, Yaxi Lu, Zhong Zhang, Yukun Yan, Yankai Lin, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Earlier work this paper cites.
A comprehensive overview of large language models
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. 2023 · 2023
Earlier work this paper cites.
The next-generation Open Targets Platform: reimagined, redesigned, rebuilt
David Ochoa, Andrew Hercules, Miguel Carmona, Daniel Suveges, Jarrod Baker, Cinzia Malangone, Irene Lopez, Alfredo Miranda, Carlos Cruz-Castillo, Luca Fumis, et al · 2023
Earlier work this paper cites.
LogicLM: Empowering Large Language Models with Tool-Enhanced Logic-Evolving Reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 8500–8518
Aosong Pan, Sameen Al-Azani, Yifei An, Zhipeng Jiang, Wen-Bin Wang, Xipeng Wan, and Man Lan. 2023 · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology . 1–22
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023 · 2023
Earlier work this paper cites.
An SPNS1-dependent lysosomal lipid transport pathway that enables cell survival under choline limitation
Samantha G Scharenberg, Wentao Dong, Ali Ghoochani, Kwamina Nyame, Roni Levin-Konigsberg, Aswini R Krishnan, Eshaan S Rawat, Kaitlyn Spees, Michael C Bassik, and Monther Abu-Remaileh. 2023 · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Earlier work this paper cites.
Role play with large language models
Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023 · 2023
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al · 2023
Earlier work this paper cites.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023 · 2023
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adri Garriga-Alonso, et al · 2023
Earlier work this paper cites.
Voyager: An Open-Ended Embodied Agent with Large Language Models
Yuqi Xie Yunfan Jiang Ajay Mandlekar Chaowei Xiao Yuke Zhu Linxi Fan Wang, Guanzhi and Anima Anandkumar. 2023 · 2023
Earlier work this paper cites.
Gpt-4v in wonderland: Large multimodal models for zero-shot smartphone gui navigation
An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, et al · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR)
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023 · 2023
Earlier work this paper cites.
Pmc-vqa: Visual instruction tuning for medical visual question answering
Xiaoman Zhang, Chaoyi Wu, Ziheng Zhao, Weixiong Lin, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023 · 2023
Earlier work this paper cites.
MITEA: A dataset for machine learning segmentation of the left ventricle in 3D echocardiography using subject-specific labels from cardiac magnetic resonance imaging
Debbie Zhao, Edward Ferdian, Gonzalo D Maso Talou, Gina M Quill, Kathleen Gilbert, Vicky Y Wang, Thiranja P Babarenda Gamage, João Pedrosa, Jan D’hooge, Timothy M Sutton, et al · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models. In International Conference on Learning Representations
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al · 2023
Earlier work this paper cites.
Accurate structure prediction of biomolecular interactions with AlphaFold 3
Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J Ballard, Joshua Bambrick, et al · 2024
Cited alongside, same era.
Evaluating correctness and faithfulness of instruction-following models for question answering
Vaibhav Adlakha, Parishad BehnamGhader, Xing Han Lu, Nicholas Meade, and Siva Reddy. 2024 · 2024
Cited alongside, same era.
OptiMUS: scalable optimization modeling with (MI) LP solvers and large language models. In Proceedings of the 41st International Conference on Machine Learning . 577–596
Ali AhmadiTeshnizi, Wenzhi Gao, and Madeleine Udell. 2024 · 2024
Cited alongside, same era.
Guided code generation with llms: A multi-agent framework for complex code tasks. In 2024 12th International Japan-Africa Conference on Electronics, Communications, and Computations (JAC-ECC) . IEEE, 215–218
Amr Almorsi, Mohanned Ahmed, and Walid Gomaa. 2024 · 2024
Cited alongside, same era.
A Survey of Cooperative Multi-Agent Reinforcement Learning for Multi-Task Scenarios
Jiajun Chai, Zijie Zhao, Yuanheng Zhu, and Dongbin Zhao. 2025b · 2025
Closest in time.
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. In The Thirteenth International Conference on Learning Representations
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al · 2025
Closest in time.
Benchmarking large language models on answering and explaining challenging medical questions. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) . 3563–3599
Hanjie Chen, Zhouxiang Fang, Yash Singla, and Mark Dredze. 2025a · 2025
Closest in time.
Kai Chen, Xinfeng Li, Tianpei Yang, Hewei Wang, Wei Dong, and Yang Gao. 2025b · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D Armstrong. 2024 · 2024
Cited alongside, same era.
Nestful: A benchmark for evaluating llms on nested sequences of api calls
Kinjal Basu, Ibrahim Abdelaziz, Kiran Kate, Mayank Agarwal, Maxwell Crouse, Yara Rizk, Kelsey Bradford, Asim Munawar, Sadhana Kumaravel, Saurabh Goyal, et al · 2024
Cited alongside, same era.
ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate. In The Twelfth International Conference on Learning Representations
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2024 · 2024
Cited alongside, same era.
Red teaming large language models in medicine: real-world insights on model behavior
Crystal T Chang, Hodan Farah, Haiwen Gui, Shawheen Justin Rezaei, Charbel Bou-Khalil, Ye-Jean Park, Akshay Swaminathan, Jesutofunmi A Omiye, Akaash Kolluri, Akash Chaurasia, et al · 2024
Cited alongside, same era.
RareAgents: Autonomous Multi-disciplinary Team for Rare Disease Diagnosis and Treatment
Xuanzhong Chen, Ye Jin, Xiaohao Mao, Lun Wang, Shuyang Zhang, and Ting Chen. 2024b · 2024
Cited alongside, same era.
MetaOpenFOAM: an LLM-based multi-agent framework for CFD
Yuxuan Chen, Xu Zhu, Hua Zhou, and Zhuyin Ren. 2024f · 2024
Cited alongside, same era.
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) . 942–966
Debrup Das, Debopriyo Banerjee, Somak Aditya, and Ashish Kulkarni. 2024 · 2024
Cited alongside, same era.
Large language model agent in financial trading: A survey
Han Ding, Yinheng Li, Junhao Wang, and Hang Chen. 2024 · 2024
Cited alongside, same era.
Enhancing diagnostic capability with multi-agents conversational large language models
Xi Chen, Huahui Yi, Mingke You, WeiZhi Liu, Li Wang, Hairui Li, Xue Zhang, Yingman Guo, Lei Fan, Gang Chen, et al · 2025
Closest in time.
Yuxuan Chen, Xu Zhu, Hua Zhou, and Zhuyin Ren. 2025g · 2025
Closest in time.
Locagent: Graph-guided llm agents for code localization
Zhaoling Chen, Xiangru Tang, Gangda Deng, Fang Wu, Jialong Wu, Zhiwei Jiang, Viktor Prasanna, Arman Cohan, and Xingyao Wang. 2025d · 2025
Closest in time.
LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval. In AI for Accelerated Materials Design-ICLR
Yuan Chiang, Elvis Hsieh, Chia-Hong Chou, and Janosh Riebesell. 2025 · 2025
Closest in time.
Codescore: Evaluating code generation by learning code execution
Yihong Dong, Jiazheng Ding, Xue Jiang, Ge Li, Zhuo Li, and Zhi Jin. 2025 · 2025
Closest in time.
Abul Ehtesham, Aditi Singh, Gaurav Kumar Gupta, and Saket Kumar. 2025 · 2025
Closest in time.
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator. In Proceedings of the 31st International Conference on Computational Linguistics . 10183–10213
Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. 2025 · 2025
Closest in time.
MCP-Zero: Proactive Toolchain Construction for LLM Agents from Scratch
Xiang Fei, Xiawu Zheng, and Hao Feng. 2025 · 2025
Closest in time.
Pharmagents: Building a virtual pharma with large language model agents
Bowen Gao, Yanwen Huang, Yiqiao Liu, Wenxuan Xie, Wei-Ying Ma, Ya-Qin Zhang, and Yanyan Lan. 2025b · 2025
Closest in time.
A survey of self-evolving agents: On path to artificial super intelligence
Huan-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Yiran Wu, et al · 2025
Closest in time.
TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools
Shanghua Gao, Richard Zhu, Zhenglun Kong, Ayush Noori, Xiao-Rui Su, Curtis Ginder, Theodoros Tsiligkaridis, and Marinka Zitnik. 2025c · 2025
Closest in time.
Automating alloy design and discovery with physics-aware multimodal multiagent AI
Alireza Ghafarollahi and Markus J Buehler. 2025a · 2025
Closest in time.
SciAgents: automating scientific discovery through bioinspired multi-agent intelligent graph reasoning
Alireza Ghafarollahi and Markus J Buehler. 2025b · 2025
Closest in time.
Robin: A multi-agent system for automating scientific discovery
Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J Szostkiewicz, Jon M Laurent, Muhammed T Razzak, Andrew D White, Michaela M Hinks, and Samuel G Rodriques. 2025 · 2025
Closest in time.
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al · 2025
Closest in time.
Agentic ai for scientific discovery: A survey of progress, challenges, and future directions
Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, and Christina Mack. 2025 · 2025
Closest in time.
MIRROR: Multi-agent Intra-and Inter-Reflection for Optimized Reasoning in Tool Learning
Zikang Guo, Benfeng Xu, Xiaorui Wang, and Zhendong Mao. 2025b · 2025
Closest in time.
Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments. In The Thirteenth International Conference on Learning Representations
SU Hongjin, Ruoxi Sun, Jinsung Yoon, Pengcheng Yin, Tao Yu, and Sercan O Arik. 2025 · 2025
Closest in time.
Model context protocol (mcp): Landscape, security threats, and future research directions
Xinyi Hou, Yanjie Zhao, Shenao Wang, and Haoyu Wang. 2025 · 2025
Closest in time.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al · 2025
Closest in time.
DrugAgent: Multi-Agent Large Language Model-Based Reasoning for Drug-Target Interaction Prediction. In ICLR Workshop on Machine Learning for Genomics Explorations
Yoshitaka Inoue, Tianci Song, Xinling Wang, Augustin Luna, and Tianfan Fu. 2025 · 2025
Closest in time.
Large language models open new way of AI-assisted molecule design for chemists
Shoichi Ishida, Tomohiro Sato, Teruki Honma, and Kei Terayama. 2025 · 2025
Closest in time.
A survey of frontiers in llm reasoning: Inference scaling, learning to reason, and agentic systems
Zixuan Ke, Fangkai Jiao, Yifei Ming, Xuan-Phi Nguyen, Austin Xu, Do Xuan Long, Minzhi Li, Chengwei Qin, Peifeng Wang, Silvio Savarese, et al · 2025
Closest in time.
Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) . 9497–9516
Jaechang Kim, Jinmin Goh, Inseok Hwang, Jaewoong Cho, and Jungseul Ok. 2025a · 2025
Closest in time.
Tiered Agentic Oversight: A Hierarchical Multi-Agent System for AI Safety in Healthcare
Yubin Kim, Hyewon Jeong, Chanwoo Park, Eugene Park, Haipeng Zhang, Xin Liu, Hyeonhoon Lee, Daniel McDuff, Marzyeh Ghassemi, Cynthia Breazeal, et al · 2025
Closest in time.
LeanAgent: Lifelong Learning for Formal Theorem Proving. In The Thirteenth International Conference on Learning Representations
Adarsh Kumarappan, Mo Tiwari, Peiyang Song, Robert Joseph George, Chaowei Xiao, and Anima Anandkumar. 2025 · 2025
Closest in time.
Multi-Agent Geospatial Copilots for Remote Sensing Workflows
Chaehong Lee, Varatheepan Paramanayakam, Andreas Karatzas, Yanan Jian, Michael Fore, Heming Liao, Fuxun Yu, Ruopu Li, Iraklis Anagnostopoulos, and Dimitrios Stamoulis. 2025b · 2025
Closest in time.
A review of prominent paradigms for llm-based agents: Tool use, planning (including rag), and feedback learning. In Proceedings of the 31st International Conference on Computational Linguistics . 9760–9779
Xinzhe Li. 2025 · 2025
Closest in time.
Giscience in the era of artificial intelligence: A research agenda towards autonomous gis
Zhenlong Li, Huan Ning, Song Gao, Krzysztof Janowicz, Wenwen Li, Samantha T Arundel, Chaowei Yang, Budhendra Bhaduri, Shaowen Wang, A Zhu, et al · 2025
Closest in time.
ChemHTS: Hierarchical Tool Stacking for Enhancing Chemical Agents
Zhucong Li, Jin Xiao, Bowei Zhang, Zhijian Zhou, Qianyu He, Fenglei Cao, Jiaqing Liang, and Yuan Qi. 2025b · 2025
Closest in time.
Surveyx: Academic survey automation via large language models
Xun Liang, Jiawei Yang, Yezhaohui Wang, Chen Tang, Zifan Zheng, Shichao Song, Zehao Lin, Yebin Yang, Simin Niu, Hanyu Wang, et al · 2025
Closest in time.
Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, et al · 2025
Closest in time.
A quantitative analysis of knowledge-learning preferences in large language models in molecular science
Pengfei Liu, Jun Tao, and Zhixiang Ren. 2025e · 2025
Closest in time.
DrBioRight 2.0: an LLM-powered bioinformatics chatbot for large-scale cancer functional proteomics analysis
Wei Liu, Jun Li, Yitao Tang, Yining Zhao, Chaozhong Liu, Meiyi Song, Zhenlin Ju, Shwetha V Kumar, Yiling Lu, Rehan Akbani, et al · 2025
Closest in time.
Aligning cyber space with physical world: A comprehensive survey on embodied ai
Yang Liu, Weixing Chen, Yongjie Bai, Xiaodan Liang, Guanbin Li, Wen Gao, and Liang Lin. 2025a · 2025
Closest in time.
Agent That Debugs: Dynamic State-Guided Vulnerability Repair
Zhengyao Liu, Yunlong Ma, Jingxuan Xu, Junchen Ai, Xiang Gao, Hailong Sun, and Abhik Roychoudhury. 2025d · 2025
Closest in time.
From intention to implementation: automating biomedical research via LLMs
Yi Luo, Linghang Shi, Yihao Li, Aobo Zhuang, Yeyun Gong, Ling Liu, and Chen Lin. 2025 · 2025
Closest in time.
BioAgents: Democratizing bioinformatics analysis with multi-agent systems
Nikita Mehandru, Amanda K Hall, Olesya Melnichenko, Yulia Dubinina, Daniel Tsirulnikov, David Bamman, Ahmed Alaa, Scott Saponas, and Venkat S Malladi. 2025 · 2025
Closest in time.
A Survey of Context Engineering for Large Language Models
Lingrui Mei, Jiayu Yao, Yuyao Ge, Yiwei Wang, Baolong Bi, Yujun Cai, Jiazhi Liu, Mingyu Li, Zhong-Zhi Li, Duzhen Zhang, et al · 2025
Closest in time.
The AI Cosmologist I: An Agentic System for Automated Data Analysis
Adam Moss. 2025 · 2025
Closest in time.
Responsibility-aware Strategic Reasoning in Probabilistic Multi-Agent Systems. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 23258–23266
Chunyan Mu, Muhammad Najib, and Nir Oren. 2025 · 2025
Closest in time.
Human-in-the-loop or AI-in-the-loop? Automate or Collaborate?. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 28594–28600
Sriraam Natarajan, Saurabh Mathur, Sahil Sidheekh, Wolfgang Stammer, and Kristian Kersting. 2025 · 2025
Closest in time.
Towards a hipaa compliant agentic ai system in healthcare
Subash Neupane, Sudip Mittal, and Shahram Rahimi. 2025 · 2025
Closest in time.
An autonomous GIS agent framework for geospatial data retrieval
Huan Ning, Zhenlong Li, Temitope Akinboyewa, and M Naser Lessani. 2025 · 2025
Closest in time.
Why do multiagent systems fail?. In ICLR 2025 Workshop on Building Trust in Language Models and Applications
Melissa Z Pan, Mert Cemri, Lakshya A Agrawal, Shuyi Yang, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Kannan Ramchandran, Dan Klein, et al · 2025
Closest in time.
Accelerating Earth Science Discovery via Multi-Agent LLM Systems
Dmitrii Pantiukhin, Boris Shapkin, Ivan Kuznetsov, Antonia Anna Jost, and Nikolay Koldunov. 2025 · 2025
Closest in time.
Ideasynth: Iterative research idea development through evolving and composing idea facets with literature-grounded feedback. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–31
Kevin Pu, KJ Kevin Feng, Tovi Grossman, Tom Hope, Bhavana Dalvi Mishra, Matt Latzke, Jonathan Bragg, Joseph Chee Chang, and Pao Siangliulue. 2025 · 2025
Closest in time.
BotSim: LLM-Powered Malicious Social Botnet Simulation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 39. 14377–14385
Boyu Qiao, Kun Li, Wei Zhou, Shilong Li, Qianqian Lu, and Songlin Hu. 2025 · 2025
Closest in time.
CRISPR-GPT for agentic automation of gene-editing experiments
Yuanhao Qu, Kaixuan Huang, Ming Yin, Kanghong Zhan, Dyllan Liu, Di Yin, Henry C Cousins, William A Johnson, Xiaotong Wang, Mihir Shah, et al · 2025
Closest in time.
Towards scientific intelligence: A survey of llm-based scientific agents
Shuo Ren, Pu Jian, Zhenjiang Ren, Chunlin Leng, Can Xie, and Jiajun Zhang. 2025 · 2025
Closest in time.
Evaluating Agent-based Program Repair at Google
Pat Rondon, Renyao Wei, José Cambronero, Jürgen Cito, Aaron Sun, Siddhant Sanyam, Michele Tufano, and Satish Chandra. 2025 · 2025
Closest in time.
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments. In The Thirteenth International Conference on Learning Representations
Yusuf H Roohani, Andrew H Lee, Qian Huang, Jian Vora, Zachary Steinhart, Kexin Huang, Alexander Marson, Percy Liang, and Jure Leskovec. 2025 · 2025
Closest in time.
AstroAgents: A Multi-Agent AI for Hypothesis Generation from Mass Spectrometry Data. In Towards Agentic AI for Science: Hypothesis Generation, Comprehension, Quantification, and Validation
Daniel Saeedi, Denise K Buckner, Jose C Aponte, and Amirali Aghazadeh. 2025 · 2025
Closest in time.
Agentrxiv: Towards collaborative autonomous research
Samuel Schmidgall and Michael Moor. 2025 · 2025
Closest in time.
Agent laboratory: Using llm agents as research assistants
Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. 2025 · 2025
Closest in time.
Paper2code: Automating code generation from scientific papers in machine learning
Minju Seo, Jinheon Baek, Seongyun Lee, and Sung Ju Hwang. 2025 · 2025
Closest in time.
Molecular analysis and design using generative artificial intelligence via multi-agent modeling
Isabella Stewart and Markus J Buehler. 2025 · 2025
Closest in time.
The ICML 2023 ranking experiment: Examining author self-assessment in ML/AI peer review
Buxin Su, Jiayao Zhang, Natalie Collina, Yuling Yan, Didong Li, Kyunghyun Cho, Jianqing Fan, Aaron Roth, and Weijie Su. 2025c · 2025
Closest in time.
BioMaster: Multi-agent System for Automated Bioinformatics Analysis Workflow
Houcheng Su, Weicai Long, and Yanlin Zhang. 2025b · 2025
Closest in time.
A survey of reasoning with foundation models: Concepts, methodologies, and outlook
Jiankai Sun, Chuanyang Zheng, Enze Xie, Zhengying Liu, Ruihang Chu, Jianing Qiu, Jiaqi Xu, Mingyu Ding, Hongyang Li, Mengzhe Geng, et al · 2025
Closest in time.
GenSim: A General Social Simulation Platform with Large Language Model based Agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations) . 143–150
Jiakai Tang, Heyang Gao, Xuchen Pan, Lei Wang, Haoran Tan, Dawei Gao, Yushuo Chen, Xu Chen, Yankai Lin, Yaliang Li, et al · 2025
Closest in time.
AI-Researcher: Autonomous Scientific Innovation
Jiabin Tang, Lianghao Xia, Zhonghang Li, and Chao Huang. 2025b · 2025
Closest in time.
OptimAI: Optimization from Natural Language Using LLM-Powered AI Agents
Raghav Thind, Youran Sun, Ling Liang, and Haizhao Yang. 2025 · 2025
Closest in time.
Multi-agent collaboration mechanisms: A survey of llms
Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry O’Sullivan, and Hoang D Nguyen. 2025 · 2025
Closest in time.
Towards conversational diagnostic artificial intelligence
Tao Tu, Mike Schaekermann, Anil Palepu, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Yong Cheng, et al · 2025
Closest in time.
A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment
Kun Wang, Guibin Zhang, Zhenhong Zhou, Jiahao Wu, Miao Yu, Shiqian Zhao, Chenlong Yin, Jinhu Fu, Yibo Yan, Hanjun Luo, et al · 2025
Closest in time.
User behavior simulation with large language model-based agents
Lei Wang, Jingsen Zhang, Hao Yang, Zhi-Yuan Chen, Jiakai Tang, Zeyu Zhang, Xu Chen, Yankai Lin, Hao Sun, Ruihua Song, et al · 2025
Closest in time.
MA-LoT: Multi-Agent Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving
Ruida Wang, Rui Pan, Yuxin Li, Jipeng Zhang, Yizhen Jia, Shizhe Diao, Renjie Pi, Junjie Hu, and Tong Zhang. 2025e · 2025
Closest in time.
Empowering Medical Multi-Agents with Clinical Consultation Flow for Dynamic Diagnosis
Sihan Wang, Suiyang Jiang, Yibo Gao, Boming Wang, Shangqi Gao, and Xiahai Zhuang. 2025a · 2025
Closest in time.
Can’t See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang, Youliang Yuan, Pinjia He, Shuai Wang, and Zhaopeng Tu. 2025c · 2025
Closest in time.
A survey of llm-based agents in medicine: How far are we from baymax?
Wenxuan Wang, Zizhan Ma, Zheng Wang, Chenghan Wu, Jiaming Ji, Wenting Chen, Xiang Li, and Yixuan Yuan. 2025d · 2025
Closest in time.
MedAgent-Pro: Towards Evidence-Based Multi-Modal Medical Diagnosis via Reasoning Agentic Workflow
Ziyue Wang, Junde Wu, Linghan Cai, Chang Han Low, Xihong Yang, Qiaxuan Li, and Yueming Jin. 2025f · 2025
Closest in time.
Mobile-agent-e: Self-evolving mobile assistant for complex tasks
Zhenhailong Wang, Haiyang Xu, Junyang Wang, Xi Zhang, Ming Yan, Ji Zhang, Fei Huang, and Heng Ji. 2025g · 2025
Closest in time.
CycleResearcher: Improving Automated Research via Automated Review. In The Thirteenth International Conference on Learning Representations
Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. 2025 · 2025
Closest in time.
TradingAgents: Multi-Agents LLM Financial Trading Framework. In The First MARW: Multi-Agent AI in the Real World Workshop at AAAI
Yijia Xiao, Edward Sun, Di Luo, and Wei Wang. 2025 · 2025
Closest in time.
Towards large reasoning models: A survey of reinforced reasoning with large language models
Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, et al · 2025
Closest in time.
The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search
Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. 2025 · 2025
Closest in time.
Xiangchao Yan, Shiyang Feng, Jiakang Yuan, Renqiu Xia, Bin Wang, Bo Zhang, and Lei Bai. 2025 · 2025
Closest in time.
DocAgent: A Multi-Agent System for Automated Code Documentation Generation
Dayu Yang, Antoine Simoulin, Xin Qian, Xiaoyi Liu, Yuwei Cao, Zhaopu Teng, and Grey Yang. 2025c · 2025
Closest in time.
AI-HOPE: An AI-Driven conversational agent for enhanced clinical and genomic data integration in precision medicine research
Ei-Wen Yang and Enrique Velazquez-Villarreal. 2025 · 2025
Closest in time.
Agentnet: Decentralized evolutionary coordination for llm-based multi-agent systems
Yingxuan Yang, Huacan Chai, Shuai Shao, Yuanyi Song, Siyuan Qi, Renting Rui, and Weinan Zhang. 2025a · 2025
Closest in time.
Survey on evaluation of llm-based agents
Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, and Michal Shmueli-Scheuer. 2025 · 2025
Closest in time.
OrcaLoca: An LLM Agent Framework for Software Issue Localization. In Forty-second International Conference on Machine Learning
Zhongming Yu, Hejia Zhang, Yujie Zhao, Hanxian Huang, Matrix Yao, Ke Ding, and Jishen Zhao. 2025 · 2025
Closest in time.
A survey of large language model agents for question answering
Murong Yue. 2025 · 2025
Closest in time.
Api agents vs. gui agents: Divergence and convergence
Chaoyun Zhang, Shilin He, Liqun Li, Si Qin, Yu Kang, Qingwei Lin, Saravan Rajmohan, and Dongmei Zhang. 2025a · 2025
Closest in time.
Appagent: Multimodal agents as smartphone users. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–20
Chi Zhang, Zhao Yang, Jiaxuan Liu, Yanda Li, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2025e · 2025
Closest in time.
Xinnong Zhang, Jiayu Lin, Xinyi Mou, Shiyue Yang, Xiawei Liu, Libo Sun, Hanjia Lyu, Yihang Yang, Weihong Qi, Yue Chen, et al · 2025
Closest in time.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al · 2025
Closest in time.
Igniting language intelligence: The hitchhiker’s guide from chain-of-thought reasoning to language agents
Zhuosheng Zhang, Yao Yao, Aston Zhang, Xiangru Tang, Xinbei Ma, Zhiwei He, Yiming Wang, Mark Gerstein, Rui Wang, Gongshen Liu, et al · 2025
Closest in time.
Lifelong learning of large language model based agents: A roadmap
Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu, and Qianli Ma. 2025 · 2025
Closest in time.
An LLM-Driven Multi-Agent Debate System for Mendelian Diseases
Xinyang Zhou, Yongyong Ren, Qianqian Zhao, Daoyi Huang, Xinbo Wang, Tingting Zhao, Zhixing Zhu, Wenyuan He, Shuyuan Li, Yan Xu, et al · 2025
Closest in time.
SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations) . 350–360
Xuhui Zhou, Zhe Su, Sophie Feng, Jiaxu Zhou, Jen-tse Huang, Hsien-Te Kao, Spencer Lynch, Svitlana Volkova, Tongshuang Wu, Anita Woolley, et al · 2025
Closest in time.
Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation
Dongsheng Zhu, Weixian Shi, Zhengliang Shi, Zhaochun Ren, Shuaiqiang Wang, Lingyong Yan, and Dawei Yin. 2025 · 2025
Closest in time.
El Agente: An autonomous agent for quantum chemistry
Yunheng Zou, Austin H Cheng, Abdulrahman Aldossary, Jiaru Bai, Shi Xuan Leong, Jorge Arturo Campos-Gonzalez-Angulo, Changhyeok Choi, Cher Tian Ser, Gary Tom, Andrew Wang, et al · 2025
Closest in time.
Kg4diagnosis: A hierarchical multi-agent llm framework with knowledge graph enhancement for medical diagnosis. In AAAI Bridge Program on AI for Medicine and Healthcare . PMLR, 195–204
Kaiwen Zuo, Yirui Jiang, Fan Mo, and Pietro Lio. 2025 · 2025
Closest in time.
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair. In Findings of the Association for Computational Linguistics ACL 2024 . 2081–2107
Hanbin Wang, Zhenghao Liu, Shuo Wang, Ganqu Cui, Ning Ding, Zhiyuan Liu, and Ge Yu. 2024e · 2081
Closest in time.
Isr-llm: Iterative self-refined large language model for long-horizon sequential task planning. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2081–2088
Zhehua Zhou, Jiayang Song, Kunpeng Yao, Zhan Shu, and Lei Ma. 2024a · 2088
Closest in time.
Deep learning for drug-induced liver injury
Youjun Xu, Ziwei Dai, Fangjin Chen, Shuaishi Gao, Jianfeng Pei, and Luhua Lai. 2015 · 2093
Closest in time.