Fetching the paper…
Reading the bibliography…
Most of the existing Large Language Model (LLM) benchmarks on scientific problem reasoning focus on problems grounded in high-school subjects and are confined to elementary algebraic operations.
Probability and statistical inference , volume 993
Hogg, R. V., Tanis, E. A., and Zimmerman, D. L · 1977
Earlier work this paper cites.
Quantum chemistry
McQuarrie, D. A · 2008
Earlier work this paper cites.
Quantum chemistry , volume 6
Levine, I. N., Busch, D. H., and Shull, H · 2009
Earlier work this paper cites.
Thermodynamics, statistical thermodynamics, and kinetics
Engel, T. and Reid, P. J · 2010
Earlier work this paper cites.
Calculus: Early transcendentals, 8th
Stewart, J., Watson, S., and Clegg, D · 2012
Earlier work this paper cites.
Bigbench: Towards an industry standard benchmark for big data analytics
Ghazal, A., Rabl, T., Hu, M., Raab, F., Poess, M., Crolotte, A., and Jacobsen, H.-A · 2013
Earlier work this paper cites.
Fundamentals of physics
Halliday, D., Resnick, R., and Walker, J · 2013
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Elementary differential equations and boundary value problems
Boyce, W. E., DiPrima, R. C., and Meade, D. B · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning
Lu, P., Gong, R., Jiang, S., Qiu, L., Huang, S., Liang, X., and Zhu, S.-C · 2021
Earlier work this paper cites.
Classical dynamics of particles and systems
Thornton, S. T. and Marion, J. B · 2021
Earlier work this paper cites.
Naturalproofs: Mathematical theorem proving in natural language
Welleck, S., Liu, J., Bras, R. L., Hajishirzi, H., Choi, Y., and Cho, K · 2021
Cited alongside, same era.
PAL: Program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2022
Cited alongside, same era.
Large language models can self-improve
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Cited alongside, same era.
Easyocr: Ready-to-use ocr
JaidedAI · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Closest in time.
Mathematical capabilities of chatgpt
Frieder, S., Pinchetti, L., Griffiths, R.-R., Salvatori, T., Lukasiewicz, T., Petersen, P. C., Chevalier, A., and Berner, J · 2023
Closest in time.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Fu, Y., Ou, L., Chen, M., Wan, Y., Peng, H., and Khot, T · 2023
Closest in time.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., and Qiao, Y · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A · 2022
Cited alongside, same era.
Lila: A unified benchmark for mathematical reasoning
Mishra, S., Finlayson, M., Lu, P., Tang, L., Welleck, S., Baral, C., Rajpurohit, T., Tafjord, O., Sabharwal, A., Clark, P., et al · 2022
Cited alongside, same era.
Chatgpt: Optimizing language models for dialogue
OpenAI · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al · 2022
Cited alongside, same era.
Galactica: A large language model for science
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., and Zhou, D · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D · 2022
Cited alongside, same era.
Guo, T., Guo, K., Liang, Z., Guo, Z., Chawla, N. V., Wiest, O., Zhang, X., et al · 2023
Closest in time.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Closest in time.
Kabir, S., Udo-Imeh, D. N., Kou, B., and Zhang, T · 2023
Closest in time.
Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., et al · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Closest in time.
Scieval: A multi-level large language model evaluation benchmark for scientific research
Sun, L., Han, Y., Zhao, Z., Ma, D., Shen, Z., Chen, B., Chen, L., and Yu, K · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N · 2023
Closest in time.
Dong, X., Zhang, P., Zang, Y., Cao, Y., Wang, B., Ouyang, L., Wei, X., Zhang, S., Duan, H., Cao, M., Zhang, W., Li, Y., Yan, H., Gao, Y., Zhang, X., Li, W., Li, J., Chen, K., He, C., Zhang, X., Qiao, Y., Lin, D., and Wang, J · 2024
Closest in time.
Sciglm: Training scientific language models with self-reflective instruction annotation and tuning
Zhang, D., Hu, Z., Zhoubian, S., Du, Z., Yang, K., Wang, Z., Yue, Y., Dong, Y., and Tang, J · 2024
Closest in time.