Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are rapidly advancing across diverse domains, yet their application in theoretical physics remains inadequate.
Zur Theorie der Metalle
Bethe, H · 1931
Earlier work this paper cites.
The Theory of Positrons
Feynman, R. P · 1949
Earlier work this paper cites.
Question of Parity Conservation in Weak Interactions
Lee, T. D. and Yang, C. N · 1956
Earlier work this paper cites.
Fundamentals of Statistical and Thermal Physics
Reif, F · 1965
Earlier work this paper cites.
Confinement of quarks
Wilson, K. G · 1974
Earlier work this paper cites.
New Method for High-Accuracy Determination of the Fine-Structure Constant Based on Quantized Hall Resistance
v. Klitzing, K., Dorda, G., and Pepper, M · 1980
Earlier work this paper cites.
Quantized Hall Conductance in a Two-Dimensional Periodic Potential
Thouless, D. J., Kohmoto, M., Nightingale, M. P., and den Nijs, M · 1982
Earlier work this paper cites.
Kosterlitz-Thouless transition in the two-dimensional quantum XY model
Ding, H.-Q. and Makivić, M. S · 1990
Earlier work this paper cites.
Interacting Electrons and Quantum Magnetism
Auerbach, A · 1994
Earlier work this paper cites.
Principles of Quantum Mechanics
Shankar, R · 1994
Earlier work this paper cites.
An Introduction to Quantum Field Theory
Peskin, M. E. and Schroeder, D. V · 1995
Earlier work this paper cites.
The role of symmetry in fundamental physics
Gross, D. J · 1996
Earlier work this paper cites.
Theory of quantum error-correcting codes
Knill, E. and Laflamme, R · 1997
Earlier work this paper cites.
Five open problems in quantum information, February 2020
Horodecki, P., Rudnicki, Ł., and Życzkowski, K · 2002
Earlier work this paper cites.
Longformer: The long-document transformer, 2020
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Introduction to Solid State Physics
Kittel, C · 2004
Earlier work this paper cites.
The One-Dimensional Hubbard Model
Essler, F. H. L., Frahm, H., Göhmann, F., Klümper, A., and Korepin, V. E · 2005
Earlier work this paper cites.
Superadditivity of communication capacity using entangled inputs
Hastings, M. B · 2009
Earlier work this paper cites.
Fermi Questions
Weinstein, L · 2010
Earlier work this paper cites.
Symmetry-Protected Topological Orders in Interacting Bosonic Systems
Chen, X., Gu, Z.-C., Liu, Z.-X., and Wen, X.-G · 2012
Earlier work this paper cites.
Mathematical Methods for Physicists
Arfken, G. B., Weber, H. J., and Harris, F. E · 2013
Earlier work this paper cites.
Stability of Local Quantum Dissipative Systems
Cubitt, T. S., Lucia, A., Michalakis, S., and Perez-Garcia, D · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Earlier work this paper cites.
Quantum Monte Carlo Approaches for Correlated Systems
Becca, F. and Sorella, S · 2017
Earlier work this paper cites.
Hand-waving and interpretive dance: An introductory course on tensor networks
Bridgeman, J. C. and Chubb, C. T · 2017
Earlier work this paper cites.
Solving the quantum many-body problem with artificial neural networks
Carleo, G. and Troyer, M · 2017
Earlier work this paper cites.
Quantum simulations with ultracold atoms in optical lattices
Gross, C. and Bloch, I · 2017
Earlier work this paper cites.
Modern Quantum Mechanics
Sakurai, J. J. and Napolitano, J · 2017
Earlier work this paper cites.
Variational Study of Fermionic and Bosonic Systems with Non-Gaussian States: Theory and Applications
Shi, T., Demler, E., and Cirac, J. I · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D · 2017
Earlier work this paper cites.
Neural-Network Quantum States, String-Bond States, and Chiral Topological States
Glasser, I., Pancotti, N., August, M., Rodriguez, I. D., and Cirac, J. I · 2018
Earlier work this paper cites.
Introduction to Quantum Mechanics
Griffiths, D. J. and Schroeter, D. F · 2018
Earlier work this paper cites.
Deep Learning and Its Application to LHC Physics
Guest, D., Cranmer, K., and Whiteson, D · 2018
Earlier work this paper cites.
Mathematical picture language program
Jaffe, A., Liu, Z., Biamonte, J. D., Ewing, J., and Vdovina, A · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Combining tensor networks with Monte Carlo methods for lattice gauge theories
Zohar, E. and Cirac, J. I · 2018
Earlier work this paper cites.
SciBERT: A pretrained language model for scientific text
Beltagy, I., Lo, K., and Cohan, A · 2019
Earlier work this paper cites.
Exact correlation functions for dual-unitary lattice models in $1+1$ dimensions
Bertini, B., Kos, P., and Prosen, T · 2019
Earlier work this paper cites.
Machine learning and the physical sciences
Carleo, G., Cirac, I., Cranmer, K., Daudet, L., Schuld, M., Tishby, N., Vogt-Maranto, L., and Zdeborová, L · 2019
Earlier work this paper cites.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Bender, E. M. and Koller, A · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Lewis, P. S. H., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D · 2020
Earlier work this paper cites.
It happens to everyone…but it’s not fun
Vidick, T · 2020
Cited alongside, same era.
A General Language Assistant as a Laboratory for Alignment, December 2021
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., and Kaplan, J · 2021
Cited alongside, same era.
Matrix Product States and Projected Entangled Pair States: Concepts, Symmetries, and Theorems
Cirac, I., Perez-Garcia, D., Schuch, N., and Verstraete, F · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Physics-informed machine learning
Karniadakis, G. E., Kevrekidis, I. G., Lu, L., Perdikaris, P., Wang, S., and Yang, L · 2021
Cited alongside, same era.
Graph of Thoughts: Solving Elaborate Problems with Large Language Models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., and Hoefler, T · 2024
Later among the works it cites.
LLMPhy: Complex Physical Reasoning Using Large Language Models and World Models, December 2024
Cherian, A., Corcodel, R., Jain, S., and Romeres, D · 2024
Later among the works it cites.
Testing GPT-4-o1-preview on math and science problems: A follow-up study, October 2024
Davis, E · 2024
Later among the works it cites.
FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI, November 2024
Glazer, E., Erdil, E., Besiroglu, T., Chicharro, D., Chen, E., Gunning, A., Olsson, C. F., Denain, J.-S., Ho, A., Santos, E. d. O., Järviniemi, O., Barnett, M., Sandler, R., Sevilla, J., Ren, Q., Pratt, E., Levine, L., Barkley, G., Stewart, N., Grechuk, B., Grechuk, T., and Enugandla, S. V · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Correlations in Perturbed Dual-Unitary Circuits: Efficient Path-Integral Formula
Kos, P., Bertini, B., and Prosen, T · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., Showk, S. E., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T. B., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Cited alongside, same era.
Measuring progress on scalable oversight for large language models, 2022
Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., Lukosiute, K., Askell, A., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Olah, C., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Kernion, J., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lovitt, L., Elhage, N., Schiefer, N., Joseph, N., Mercado, N., DasSarma, N., Larson, R., McCandlish, S., Kundu, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatfield-Dodds, Z., Mann, B., and Kaplan, J · 2022
Cited alongside, same era.
Magnetic control of tokamak plasmas through deep reinforcement learning
Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de las Casas, D., Donner, C., Fritz, L., Galperti, C., Huber, A., Keeling, J., Tsimpoukelli, M., Kay, J., Merle, A., Moret, J.-M., Noury, S., Pesamosca, F., Pfau, D., Sauter, O., Sommariva, C., Coda, S., Duval, B., Fasoli, A., Kohli, P., Kavukcuoglu, K., Hassabis, D., and Riedmiller, M · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V · 2022
Cited alongside, same era.
Language models of code are few-shot commonsense learners, 2022
Madaan, A., Zhou, S., Alon, U., Yang, Y., and Neubig, G · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Hagendorff, T., Dasgupta, I., Binz, M., Chan, S. C. Y., Lampinen, A., Wang, J. X., Akata, Z., and Schulz, E · 2024
Later among the works it cites.
SWE-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Later among the works it cites.
Prover-Verifier Games improve legibility of LLM outputs, August 2024
Kirchner, J. H., Chen, Y., Edwards, H., Leike, J., McAleese, N., and Burda, Y · 2024
Later among the works it cites.
AgentBench: Evaluating LLMs as agents
Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., and Tang, J · 2024
Later among the works it cites.
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, August 2024
Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., and Ha, D · 2024
Later among the works it cites.
Augmenting large language models with chemistry tools
M. Bran, A., Cox, S., Schilter, O., Baldassari, C., White, A. D., and Schwaller, P · 2024
Later among the works it cites.
LLM Critics Help Catch LLM Bugs, June 2024
McAleese, N., Pokorny, R. M., Uribe, J. F. C., Nitishinskaya, E., Trebacz, M., and Leike, J · 2024
Later among the works it cites.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J. R., Ellenberg, J. S., Wang, P., Fawzi, O., Kohli, P., and Fawzi, A · 2024
Later among the works it cites.
Language agents achieve superhuman synthesis of scientific knowledge, 2024
Skarlinski, M. D., Cox, S., Laurent, J. M., Braza, J. D., Hinks, M., Hammerling, M. J., Ponnapati, M., Rodriques, S. G., and White, A. D · 2024
Later among the works it cites.
Tian, Y., Wang, L., and Wang, L · 2024
Later among the works it cites.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Later among the works it cites.
ReFT: Reasoning with Reinforced Fine-Tuning
Trung, L., Zhang, X., Jie, Z., Sun, P., Jin, X., and Li, H · 2024
Later among the works it cites.
SciBench: Evaluating college-level scientific problem-solving abilities of large language models
Wang, X., Hu, Z., Lu, P., Zhu, Y., Zhang, J., Subramaniam, S., Loomba, A. R., Zhang, S., Sun, Y., and Wang, W · 2024
Later among the works it cites.
Reasoning or reciting? Exploring the capabilities and limitations of language models through counterfactual tasks
Wu, Z., Qiu, L., Ross, A., Akyürek, E., Chen, B., Wang, B., Kim, N., Andreas, J., and Kim, Y · 2024
Later among the works it cites.
Zhou, A., Hawkins, L., and Gentine, P · 2024
Later among the works it cites.
How should the advancement of large language models affect the practice of science?
Binz, M., Alaniz, S., Roskies, A., Aczel, B., Bergstrom, C. T., Allen, C., Schad, D., Wulff, D., West, J. D., Zhang, Q., Shiffrin, R. M., Gershman, S. J., Popov, V., Bender, E. M., Marelli, M., Botvinick, M. M., Akata, Z., and Schulz, E · 2025
Closest in time.
The fiction machine
Bottou, L. and Schölkopf, B · 2025
Closest in time.
Chung, D. J. H., Gao, Z., Kvasiuk, Y., Li, T., Münchmeyer, M., Rudolph, M., Sala, F., and Tadepalli, S. C · 2025
Closest in time.
Davis, E. and Aaronson, S · 2025
Closest in time.
DeepSeek-V3 Technical Report, February 2025
DeepSeek-AI · 2025
Closest in time.
DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Xu, H., Ding, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Chen, J., Yuan, J., Tu, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., You, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Zhou, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z · 2025
Closest in time.
MATH-Perturb: Benchmarking LLMs’ Math Reasoning Abilities against Hard Perturbations, February 2025
Huang, K., Guo, J., Li, Z., Ji, X., Ge, J., Li, W., Guo, Y., Cai, T., Yuan, H., Wang, R., Wu, Y., Yin, M., Tang, S., Huang, Y., Jin, C., Chen, X., Zhang, C., and Wang, M · 2025
Closest in time.
Open quantum problems, 2025
IQOQI Vienna · 2025
Closest in time.
Large language models for human-machine collaborative particle accelerator tuning through natural language
Kaiser, J., Lauscher, A., and Eichler, A · 2025
Closest in time.
Introducing deep research, February 2025
OpenAI · 2025
Closest in time.
Quantum many-body physics calculations with large language models
Pan, H., Mudur, N., Taranto, W., Tikhanovskaya, M., Venugopalan, S., Bahri, Y., Brenner, M. P., and Kim, E.-A · 2025
Closest in time.
Humanity’s Last Exam, April 2025
Phan, L., Gatti, A., Han, Z., and al., e · 2025
Closest in time.
Qiu, S., Guo, S., Song, Z.-Y., Sun, Y., Cai, Z., Wei, J., Luo, T., Yin, Y., Zhang, H., Hu, Y., Wang, C., Tang, C., Chang, H., Liu, Q., Zhou, Z., Zhang, T., Zhang, J., Liu, Z., Li, M., Zhang, Y., Jing, B., Yin, X., Ren, Y., Fu, Z., Wang, W., Tian, X., Lv, A., Man, L., Li, J., Tao, F., Sun, Q., Liang, Z., Mu, Y., Li, Z., Zhang, J.-J., Zhang, S., Li, X., Xia, X., Lin, J., Shen, Z., Chen, J., Xiong, Q., Wang, B., Wang, F., Ni, Z., Zhang, B., Cui, F., Shao, C., Cao, Q.-H., Luo, M.-x., Zhang, M., and Zhu, H. X · 2025
Closest in time.
A benchmark for large language models in bioinformatics, April 2025
Sarwal, V., Andreoletti, G., Munteanu, V., Suhodolschi, A., Ciorba, D., Bostan, V., Dimian, M., Eskin, E., Wang, W., and Mangul, S · 2025
Closest in time.
PaperBench: Evaluating AI’s Ability to Replicate AI Research, April 2025
Starace, G., Jaffe, O., Sherburn, D., Aung, J., Chan, J. S., Maksin, L., Dias, R., Mays, E., Kinsella, B., Thompson, W., Heidecke, J., Glaese, A., and Patwardhan, T · 2025
Closest in time.
Kimi k1.5: Scaling Reinforcement Learning with LLMs, March 2025
Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., Tang, C., Wang, C., Zhang, D., Yuan, E., Lu, E., Tang, F., Sung, F., Wei, G., Lai, G., Guo, H., Zhu, H., Ding, H., Hu, H., Yang, H., Zhang, H., Yao, H., Zhao, H., Lu, H., Li, H., Yu, H., Gao, H., Zheng, H., Yuan, H., Chen, J., Guo, J., Su, J., Wang, J., Zhao, J., Zhang, J., Liu, J., Yan, J., Wu, J., Shi, L., Ye, L., Yu, L., Dong, M., Zhang, N., Ma, N., Pan, Q., Gong, Q., Liu, S., Ma, S., Wei, S., Cao, S., Huang, S., Jiang, T., Gao, W., Xiong, W., He, W., Huang, W., Wu, W., He, W., Wei, X., Jia, X., Wu, X., Xu, X., Zu, X., Zhou, X., Pan, X., Charles, Y., Li, Y., Hu, Y., Liu, Y., Chen, Y., Wang, Y., Liu, Y., Qin, Y., Liu, Y., Yang, Y., Bao, Y., Du, Y., Wu, Y., Wang, Y., Zhou, Z., Wang, Z., Li, Z., Zhu, Z., Zhang, Z., Wang, Z., Yang, Z., Huang, Z., Huang, Z., Xu, Z., and Yang, Z · 2025
Closest in time.
Seephys: Does seeing help thinking? – benchmarking vision-based physics reasoning, 2025
Xiang, K., Li, H., Zhang, T. J., Huang, Y., Liu, Z., Qu, P., He, J., Chen, J., Yuan, Y.-J., Han, J., Xu, H., Li, H., Sachan, M., and Liang, X · 2025
Closest in time.
Newtonbench: Benchmarking generalizable scientific law discovery in llm agents, 2025
Zheng, T., Tam, K. K.-W., Nguyen, N. H.-N. K., Xu, B., Wang, Z., Cheng, J., Tsang, H. T., Wang, W., Bai, J., Fang, T., Song, Y., Wong, G. Y., and See, S · 2025
Closest in time.
El agente: An autonomous agent for quantum chemistry
Zou, Y., Cheng, A. H., Aldossary, A., Bai, J., Leong, S. X., Campos-Gonzalez-Angulo, J. A., Choi, C., Ser, C. T., Tom, G., Wang, A., Zhang, Z., Yakavets, I., Hao, H., Crebolder, C., Bernales, V., and Aspuru-Guzik, A · 2025
Closest in time.
Position: Science is collaborative—llm for science should be too
Anonymous · 2026
Closest in time.
CMT-benchmark: A benchmark for condensed matter theory built by expert researchers
Pan, H., Roggeveen, J. V., Berg, E., Alvarez, J. F. C., Chowdhury, D., Ganguli, S., Ghimenti, F., Hasik, J., Hunt, H. S., Jiang, H.-C., Kamb, M., Kao, Y.-J., Khatami, E., Lawler, M. J., Luo, D., Neupert, T., Qi, X., Brenner, M., and Kim, E.-A · 2026
Closest in time.
CMPhysbench: A benchmark for evaluating large language models in condensed matter physics
Wang, W., Huang, D., LI, J., Yang, T., Zheng, Z., Peng, C., Zhang, D., Han, D., Chen, B., Luo, B., Liu, Z., kunling liu, Gao, Z., Shiqigeng, Ma, W., Su, J., Li, X., Pu, S., Shui, Y., Cheng, Q., Dou, Z., Cui, D., He, C., Zeng, J., Xie, Z., Su, M., Zhou, D., Li, Y., Ouyang, W., Cai, Y., Dai, X., Zhang, S., BAI, L., Cheng, J., Fang, Z., and Weng, H · 2026
Closest in time.