Fetching the paper…
Reading the bibliography…
Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored.
History: A Very Short Introduction
John H. Arnold · 2000
Earlier work this paper cites.
A Global History of History
Daniel Woolf · 2011
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W Cohen, and Xinghua Lu · 2019
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Chemcrow: Augmenting large-language models with chemistry tools
Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller · 2023
Earlier work this paper cites.
Gaia: a benchmark for general ai assistants
Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom · 2023
Earlier work this paper cites.
War and peace (waragent): Large language model-based multi-agent simulation of world wars
Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang · 2023
Earlier work this paper cites.
Scibench: Evaluating college-level scientific problem-solving abilities of large language models
Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang · 2023
Earlier work this paper cites.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, et al · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Earlier work this paper cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Earlier work this paper cites.
Crispr-gpt: An llm agent for automated design of gene-editing experiments
Kaixuan Huang, Yuanhao Qu, Henry Cousins, William A Johnson, Di Yin, Mihir Shah, Denny Zhou, Russ Altman, Mengdi Wang, and Le Cong · 2024
Cited alongside, same era.
Long term memory: The foundation of ai self-evolution
Xun Jiang, Feng Li, Han Zhao, Jiaying Wang, Jun Shao, Shihao Xu, Shu Zhang, Weiling Chen, Xavier Tang, Yize Chen, et al · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo · 2024
Cited alongside, same era.
Ds-agent: Automated data science by empowering large language models with case-based reasoning, 2024
Siyuan Guo, Cheng Deng, Ying Wen, Hechang Chen, Yi Chang, and Jun Wang · 2024
Cited alongside, same era.
Battleagent: Multi-modal dynamic emulation on historical battles to complement historical analysis
Autoagent: A fully-automated and zero-code framework for llm agents
Jiabin Tang, Tianyu Fan, and Chao Huang · 2025
Closest in time.
Owl: Optimized workforce learning for general multi-agent assistance in real-world task automation, 2025
Mengkang Hu, Yuhang Zhou, Wendong Fan, Yuzhou Nie, Bowei Xia, Tao Sun, Ziyu Ye, Zhaoxuan Jin, Yingru Li, Zeyu Zhang, Yifeng Wang, Qianshuo Ye, Ping Luo, and Guohao Li · 2025
Closest in time.
‘smolagents‘: a smol library to build great agentic systems
Aymeric Roucher, Albert Villanova del Moral, Thomas Wolf, Leandro von Werra, and Erik Kaunismäki · 2025
Closest in time.
Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation
Cheng-Yi Li, Kao-Jung Chang, Cheng-Fu Yang, Hsin-Yu Wu, Wenting Chen, Hritik Bansal, Ling Chen, Yi-Ping Yang, Yu-Chun Chen, Shih-Pin Chen, et al · 2025
Closest in time.
Medical large language model benchmarks should prioritize construct validity
Ahmed Alaa, Thomas Hartvigsen, Niloufar Golchini, Shiladitya Dutta, Frances Dean, Inioluwa Deborah Raji, and Travis Zack · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shuhang Lin, Wenyue Hua, Lingyao Li, Che-Jui Chang, Lizhou Fan, Jianchao Ji, Hang Hua, Mingyu Jin, Jiebo Luo, and Yongfeng Zhang · 2024
Cited alongside, same era.
Legalbench-rag: A benchmark for retrieval-augmented generation in the legal domain
Nicholas Pipitone and Ghita Houir Alami · 2024
Cited alongside, same era.
Large language models’ expert-level global history knowledge benchmark (hist-llm)
Jakob Hauser, Daniel Kondor, Jenny Reddish, Majid Benam, Enrico Cioni, Federica Villa, James Bennett, Daniel Hoyer, Pieter Francois, Peter Turchin, et al · 2024
Cited alongside, same era.
Browser use: Enable ai to control your browser, 2024
Magnus Müller and Gregor Žunič · 2024
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Cited alongside, same era.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, et al · 2025
Cited alongside, same era.
Phybench: Holistic evaluation of physical perception and reasoning in large language models
Shi Qiu, Shaoyang Guo, Zhuo-Yang Song, Yunbo Sun, Zeyu Cai, Jiashen Wei, Tianyu Luo, Yixuan Yin, Haoxu Zhang, Yi Hu, et al · 2025
Cited alongside, same era.
Introducing deep research
OpenAI
Cited in the paper.
Closest in time.
Benchmarking large language models via random variables
Zijin Hong, Hao Wu, Su Dong, Junnan Dong, Yilin Xiao, Yujing Zhang, Zhu Wang, Feiran Huang, Linyi Li, Hongxia Yang, et al · 2025
Closest in time.
Unveiling challenges for llms in enterprise data engineering
Jan-Micha Bodensohn, Ulf Brackmann, Liane Vogel, Anupam Sanghi, and Carsten Binnig · 2025
Closest in time.
Financeqa: A benchmark for evaluating financial analysis capabilities of large language models
Spencer Mateega, Carlos Georgescu, and Danny Tang · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Introducing openai o3 and o4-mini, 2025
OpenAI · 2025
Closest in time.
Grok 3 beta — the age of reasoning agents, 2025
xAI · 2025
Closest in time.