Fetching the paper…
Reading the bibliography…
Chain-of-Thought (CoT) prompting enhances mathematical reasoning in large language models (LLMs) by enabling detailed step-by-step solutions.
Evaluation of text generation: A survey, 2021
Celikyilmaz, A., Clark, E., and Gao, J · 2006
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021a
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2009
Earlier work this paper cites.
Natural language generation challenges for explainable AI
Reiter, E · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
The lean 4 theorem prover and programming language
Moura, L. d. and Ullrich, S · 2021
Earlier work this paper cites.
Diagnosing the first-order logical reasoning ability through LogicNLI
Tian, J., Li, Y., Chen, W., Xiao, L., He, H., and Jin, Y · 2021
Earlier work this paper cites.
Naturalproofs: Mathematical theorem proving in natural language
Welleck, S., Liu, J., Bras, R. L., Hajishirzi, H., Choi, Y., and Cho, K · 2021
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback, 2022
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Autoformalization with large language models, 2022
Wu, Y., Jiang, A. Q., Li, W., Rabe, M. N., Staats, C., Jamnik, M., and Szegedy, C · 2022
Earlier work this paper cites.
Have LLMs advanced enough? a challenging problem solving benchmark for large language models
Arora, D., Singh, H., and Mausam · 2023
Earlier work this paper cites.
Alexa arena: A user-centric interactive platform for embodied ai
Gao, Q., Thattai, G., Shakiah, S., Gao, X., Pansare, S., Sharma, V., Sukhatme, G., Shi, H., Yang, B., Zhang, D., Hu, L., Arumugam, K., Hu, S., Wen, M., Guthy, D., Chung, S., Khanna, R., Ipek, O., Ball, L., Bland, K., Rocker, H., Johnston, M., Ghanadan, R., Hakkani-Tur, D., and Natarajan, P · 2023
Earlier work this paper cites.
Roscoe: A suite of metrics for scoring step-by-step reasoning, 2023
Golovneva, O., Chen, M., Poff, S., Corredor, M., Zettlemoyer, L., Fazel-Zarandi, M., and Celikyilmaz, A · 2023
Earlier work this paper cites.
Towards reasoning in large language models: A survey, 2023
Huang, J. and Chang, K. C.-C · 2023
Earlier work this paper cites.
MathPrompter: Mathematical reasoning using large language models
Imani, S., Du, L., and Shrivastava, H · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention, 2023
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Let’s verify step by step, 2023
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
Deductive verification of chain-of-thought reasoning, 2023
Ling, Z., Fang, Y., Li, X., Huang, Z., Lee, M., Memisevic, R., and Su, H · 2023
Cited alongside, same era.
Receval: Evaluating reasoning chains via correctness and informativeness, 2023
Prasad, A., Saha, S., Zhou, X., and Bansal, M · 2023
Cited alongside, same era.
Large language models can be easily distracted by irrelevant context
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., Chi, E., Schärli, N., and Zhou, D · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models, 2023
A survey on large language models for code generation, 2024
Jiang, J., Wang, F., Shen, J., Kim, S., and Kim, S · 2024
Later among the works it cites.
Orca-math: Unlocking the potential of slms in grade school math, 2024
Mitra, A., Khanpour, H., Rosset, C., and Awadallah, A · 2024
Later among the works it cites.
Autoformalizing euclidean geometry, 2024
Murphy, L., Yang, K., Sun, J., Li, Z., Anandkumar, A., and Si, X · 2024
Later among the works it cites.
Openai o1 system card
OpenAI · 2024
Later among the works it cites.
OpenAI, :, Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., Iftimie, A., Karpenko, A., Passos, A. T., Neitz, A., Prokofiev, A., Wei, A., Tam, A., Bennett, A., Kumar, A., Saraiva, A., Vallone, A., Duberstein, A., Kondrich, A., Mishchenko, A., Applebaum, A., Jiang, A., Nair, A., Zoph, B., Ghorbani, B., Rossen, B., Sokolowsky, B., Barak, B., McGrew, B., Minaiev, B., Hao, B., Baker, B., Houghton, B., McKinzie, B., Eastman, B., Lugaresi, C., Bassin, C., Hudson, C., Li, C. M., de Bourcy, C., Voss, C., Shen, C., Zhang, C., Koch, C., Orsinger, C., Hesse, C., Fischer, C., Chan, C., Roberts, D., Kappler, D., Levy, D., Selsam, D., Dohan, D., Farhi, D., Mely, D., Robinson, D., Tsipras, D., Li, D., Oprica, D., Freeman, E., Zhang, E., Wong, E., Proehl, E., Cheung, E., Mitchell, E., Wallace, E., Ritter, E., Mays, E., Wang, F., Such, F. P., Raso, F., Leoni, F., Tsimpourlas, F., Song, F., von Lohmann, F., Sulit, F., Salmon, G., Parascandolo, G., Chabot, G., Zhao, G., Brockman, G., Leclerc, G., Salman, H., Bao, H., Sheng, H., Andrin, H., Bagherinezhad, H., Ren, H., Lightman, H., Chung, H. W., Kivlichan, I., O’Connell, I., Osband, I., Gilaberte, I. C., Akkaya, I., Kostrikov, I., Sutskever, I., Kofman, I., Pachocki, J., Lennon, J., Wei, J., Harb, J., Twore, J., Feng, J., Yu, J., Weng, J., Tang, J., Yu, J., Candela, J. Q., Palermo, J., Parish, J., Heidecke, J., Hallman, J., Rizzo, J., Gordon, J., Uesato, J., Ward, J., Huizinga, J., Wang, J., Chen, K., Xiao, K., Singhal, K., Nguyen, K., Cobbe, K., Shi, K., Wood, K., Rimbach, K., Gu-Lemberg, K., Liu, K., Lu, K., Stone, K., Yu, K., Ahmad, L., Yang, L., Liu, L., Maksin, L., Ho, L., Fedus, L., Weng, L., Li, L., McCallum, L., Held, L., Kuhn, L., Kondraciuk, L., Kaiser, L., Metz, L., Boyd, M., Trebacz, M., Joglekar, M., Chen, M., Tintor, M., Meyer, M., Jones, M., Kaufer, M., Schwarzer, M., Shah, M., Yatbaz, M., Guan, M. Y., Xu, M., Yan, M., Glaese, M., Chen, M., Lampe, M., Malek, M., Wang, M., Fradin, M., McClay, M., Pavlov, M., Wang, M., Wang, M., Murati, M., Bavarian, M., Rohaninejad, M., McAleese, N., Chowdhury, N., Chowdhury, N., Ryder, N., Tezak, N., Brown, N., Nachum, O., Boiko, O., Murk, O., Watkins, O., Chao, P., Ashbourne, P., Izmailov, P., Zhokhov, P., Dias, R., Arora, R., Lin, R., Lopes, R. G., Gaon, R., Miyara, R., Leike, R., Hwang, R., Garg, R., Brown, R., James, R., Shu, R., Cheu, R., Greene, R., Jain, S., Altman, S., Toizer, S., Toyer, S., Miserendino, S., Agarwal, S., Hernandez, S., Baker, S., McKinney, S., Yan, S., Zhao, S., Hu, S., Santurkar, S., Chaudhuri, S. R., Zhang, S., Fu, S., Papay, S., Lin, S., Balaji, S., Sanjeev, S., Sidor, S., Broda, T., Clark, A., Wang, T., Gordon, T., Sanders, T., Patwardhan, T., Sottiaux, T., Degry, T., Dimson, T., Zheng, T., Garipov, T., Stasi, T., Bansal, T., Creech, T., Peterson, T., Eloundou, T., Qi, V., Kosaraju, V., Monaco, V., Pong, V., Fomenko, V., Zheng, W., Zhou, W., McCabe, W., Zaremba, W., Dubois, Y., Lu, Y., Chen, Y., Cha, Y., Bai, Y., He, Y., Zhang, Y., Wang, Y., Shao, Z., and Li, Z · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Cited alongside, same era.
Large language models are better reasoners with self-verification
Weng, Y., Zhu, M., Xia, F., Li, B., He, S., Liu, S., Sun, B., Liu, K., and Zhao, J · 2023
Cited alongside, same era.
Leandojo: Theorem proving with retrieval-augmented language models, 2023
Yang, K., Swope, A. M., Gu, A., Chalamala, R., Song, P., Yu, S., Godil, S., Prenger, R., and Anandkumar, A · 2023
Cited alongside, same era.
Tree of thoughts: Deliberate problem solving with large language models, 2023
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Cited alongside, same era.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W · 2023
Cited alongside, same era.
Large language models for mathematical reasoning: Progresses and challenges, 2024
Ahn, J., Verma, R., Lou, R., Liu, D., Zhang, R., and Yin, W · 2024
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., and Hoefler, T · 2024
Cited alongside, same era.
Daheim, N., Macina, J., Kapur, M., Gurevych, I., and Sachan, M · 2024
Cited alongside, same era.
Later among the works it cites.
Simulating user agents for embodied conversational-ai, 2024
Philipov, D., Dongre, V., Tur, G., and Hakkani-Tür, D · 2024
Later among the works it cites.
On the self-verification limitations of large language models on reasoning and planning tasks, 2024
Stechly, K., Valmeekam, K., and Kambhampati, S · 2024
Later among the works it cites.
LLMs cannot find reasoning errors, but can correct them given the error location
Tyen, G., Mansoor, H., Carbune, V., Chen, P., and Mak, T · 2024
Later among the works it cites.
Scibench: Evaluating college-level scientific problem-solving abilities of large language models, 2024
Wang, X., Hu, Z., Lu, P., Zhu, Y., Zhang, J., Subramaniam, S., Loomba, A. R., Zhang, S., Sun, Y., and Wang, W · 2024
Later among the works it cites.
Large language models can self-correct with key condition verification
Wu, Z., Zeng, Q., Zhang, Z., Tan, Z., Shen, C., and Jiang, M · 2024
Later among the works it cites.
Formal mathematical reasoning: A new frontier in ai, 2024
Yang, K., Poesia, G., He, J., Li, W., Lauter, K., Chaudhuri, S., and Song, D · 2024
Later among the works it cites.
Processbench: Identifying process errors in mathematical reasoning, 2024
Zheng, C., Zhang, Z., Zhang, B., Lin, R., Lu, K., Yu, B., Liu, D., Zhou, J., and Lin, J · 2024
Later among the works it cites.
Deductive beam search: Decoding deducible rationale for chain-of-thought reasoning, 2024
Zhu, T., Zhang, K., Xie, J., and Su, Y · 2024
Later among the works it cites.
Qwen2.5 technical report, 2025
Qwen, :, Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Tang, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Wan, Y., Liu, Y., Cui, Z., Zhang, Z., and Qiu, Z · 2025
Closest in time.