Fetching the paper…
Reading the bibliography…
We introduce a benchmark to evaluate the capability of AI to solve problems in theoretical physics, focusing on high-energy theory and cosmology.
A Mathematician’s Apology
G. H. Hardy · 1940
Earlier work this paper cites.
How to Solve It: A New Aspect of Mathematical Method
George Pólya · 1945
Earlier work this paper cites.
Spinor structure of space-times in general relativity. i
Robert P. Geroch · 1968
Earlier work this paper cites.
Proof of the positive mass theorem. 2
Richard Schon and Shing-Tung Yau · 1981
Earlier work this paper cites.
On Witten’s Proof of the Positive Energy Theorem
Thomas Parker and Clifford Henry Taubes · 1982
Earlier work this paper cites.
The mathematical theory of black holes
Subrahmanyan Chandrasekhar · 1985
Earlier work this paper cites.
Intersection theory on the moduli space of curves and the matrix Airy function
M. Kontsevich · 1992
Earlier work this paper cites.
Entropy and area
Mark Srednicki · 1993
Earlier work this paper cites.
Some properties of Noether charge and a proposal for dynamical black hole entropy
Vivek Iyer and Robert M. Wald · 1994
Earlier work this paper cites.
Supersymmetric Gauge Field Theory and String Theory
D. Bailin and Alexander Love · 1994
Earlier work this paper cites.
An Introduction to quantum field theory
Michael E. Peskin and Daniel V. Schroeder · 1995
Earlier work this paper cites.
Ten lessons i wish i had been taught
Gian-Carlo Rota · 1997
Earlier work this paper cites.
Mirror symmetry, Langlands duality, and the Hitchin system
Tamas Hausel and Michael Thaddeus · 2003
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
TikZ-Feynman: Feynman diagrams with TikZ
Joshua Ellis · 2016
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Quantum interference in gravitational particle production
Edward Basso, Daniel J. H. Chung, Edward W. Kolb, and Andrew J. Long · 2022
Earlier work this paper cites.
Minif2f: a cross-system benchmark for formal olympiad-level mathematics, 2022
Kunhao Zheng, Jesse Michael Han, and Stanislas Polu · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Earlier work this paper cites.
Decomposed prompting: A modular approach for solving complex tasks
Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal · 2022
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi · 2022
Earlier work this paper cites.
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen · 2022
Earlier work this paper cites.
Have llms advanced enough? a challenging problem solving benchmark for large language models, 2023
Daman Arora, Himanshu Gaurav Singh, and Mausam · 2023
Earlier work this paper cites.
Mathprompter: Mathematical reasoning using large language models
Shima Imani, Liang Du, and Harsh Shrivastava · 2023
Earlier work this paper cites.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2023
Cited alongside, same era.
Do large language models know what they don’t know?
Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Xuanjing Huang · 2023
Cited alongside, same era.
Chain-of-verification reduces hallucination in large language models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston · 2023
Cited alongside, same era.
Jiaxin Zhang, Zhuohang Li, Kamalika Das, Bradley Malin, and Sricharan Kumar · 2023
Cited alongside, same era.
A tale of two fields: Neural network-enhanced non-gaussianity search with halos, 2024
Yurii Kvasiuk, Moritz Münchmeyer, and Kendrick Smith · 2024
Later among the works it cites.
Meta llama 3.1, 2024
Meta AI · 2024
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al · 2024
Later among the works it cites.
Putnambench: Evaluating neural theorem-provers on the putnam mathematical competition, 2024
George Tsoukalas, Jasper Lee, John Jennings, Jimmy Xin, Michelle Ding, Michael Jennings, Amitayush Thakur, and Swarat Chaudhuri · 2024
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tanuj Kumar and Mikhail A. Kats · 2023
Cited alongside, same era.
Mathchat: Converse to tackle challenging math problems with llm agents
Yiran Wu, Feiran Jia, Shaokun Zhang, Hangyu Li, Erkang Zhu, Yue Wang, Yin Tat Lee, Richard Peng, Qingyun Wu, and Chi Wang · 2023
Cited alongside, same era.
Fimo: A challenge formal dataset for automated theorem proving, 2023
Chengwu Liu, Jianhao Shen, Huajian Xin, Zhengying Liu, Ye Yuan, Haiming Wang, Wei Ju, Chuanyang Zheng, Yichun Yin, Lin Li, Ming Zhang, and Qun Liu · 2023
Cited alongside, same era.
Proofnet: Autoformalizing and formally proving undergraduate-level mathematics, 2023
Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W. Ayers, Dragomir Radev, and Jeremy Avigad · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Cited alongside, same era.
Take a step back: Evoking reasoning via abstraction in large language models
Huaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng, Ed H Chi, Quoc V Le, and Denny Zhou · 2023
Cited alongside, same era.
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar · 2024
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models, 2024
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu · 2024
Later among the works it cites.
Quiet-star: Language models can teach themselves to think before speaking
Eric Zelikman, Georges Harik, Yijia Shao, Varuna Jayasiri, Nick Haber, and Noah D Goodman · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
From decoding to meta-generation: Inference-time algorithms for large language models, 2024
Sean Welleck, Amanda Bertsch, Matthew Finlayson, Hailey Schoelkopf, Alex Xie, Graham Neubig, Ilia Kulikov, and Zaid Harchaoui · 2024
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2024
Later among the works it cites.
Analyzing the performance of self-refine on different large language models, 2024
Anton Forsman · 2024
Later among the works it cites.
Mutual reasoning makes smaller llms stronger problem-solvers
Zhenting Qi, Mingyuan Ma, Jiahang Xu, Li Lyna Zhang, Fan Yang, and Mao Yang · 2024
Later among the works it cites.
Llama-berry: Pairwise optimization for o1-like olympiad-level mathematical reasoning, 2024
Di Zhang, Jianbo Wu, Jingdi Lei, Tong Che, Jiatong Li, Tong Xie, Xiaoshui Huang, Shufei Zhang, Marco Pavone, Yuqiang Li, Wanli Ouyang, and Dongzhan Zhou · 2024
Later among the works it cites.
Mindstar: Enhancing math reasoning in pre-trained llms at inference time
Jikun Kang, Xin Zhe Li, Xi Chen, Amirreza Kazemi, Qianyi Sun, Boxing Chen, Dong Li, Xu He, Quan He, Feng Wen, et al · 2024
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2024
Later among the works it cites.
Archon: An architecture search framework for inference-time techniques
Jon Saad-Falcon, Adrian Gamarra Lafuente, Shlok Natarajan, Nahum Maru, Hristo Todorov, Etash Guha, E Kelly Buchanan, Mayee Chen, Neel Guha, Christopher Ré, et al · 2024
Later among the works it cites.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, et al · 2025
Closest in time.
o3-mini, 2025
OpenAI · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
Redstar: Does scaling long-cot data unlock better slow-reasoning systems?, 2025
Haotian Xu, Xing Wu, Weinong Wang, Zhongzhi Li, Da Zheng, Boyuan Chen, Yi Hu, Shijia Kang, Jiaming Ji, Yingying Zhang, Zhijiang Guo, Yaodong Yang, Muhan Zhang, and Debing Zhang · 2025
Closest in time.
Bespoke-stratos: The unreasonable effectiveness of reasoning distillation, 2025
Bespoke Labs · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.
Theoretical guarantees on the best-of-n alignment policy, 2025
Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D’Amour, Jacob Eisenstein, Chirag Nagpal, and Ananda Theertha Suresh · 2025
Closest in time.