Fetching the paper…
Reading the bibliography…
Chain-of-Thought (CoT) prompting has shown promise in enhancing the reasoning capabilities of large language models (LLMs) by generating natural language (NL) rationales that lead to the final answer.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., Showk, S. E., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
Codet: Code generation with generated tests
Chen, B., Zhang, F., Nguyen, A., Zan, D., Lin, Z., Lou, J., and Chen, W · 2023
Earlier work this paper cites.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks
Chen, W., Ma, X., Wang, X., and Cohen, W. W · 2023
Earlier work this paper cites.
Automated repair of programs from large language models
Fan, Z., Gao, X., Mirchev, M., Roychoudhury, A., and Tan, S. H · 2023
Earlier work this paper cites.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Fu, Y., Ou, L., Chen, M., Wan, Y., Peng, H., and Khot, T · 2023
Earlier work this paper cites.
PAL: program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2023
Earlier work this paper cites.
Getting from generative AI to trustworthy AI: what llms might learn from cyc
Lenat, D. and Marcus, G · 2023
Earlier work this paper cites.
Deductive verification of chain-of-thought reasoning
Ling, Z., Fang, Y., Li, X., Huang, Z., Lee, M., Memisevic, R., and Su, H · 2023
Earlier work this paper cites.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Liu, J., Xia, C. S., Wang, Y., and Zhang, L · 2023
Earlier work this paper cites.
LEVER: learning to verify language-to-code generation with execution
Ni, A., Iyer, S., Radev, D., Stoyanov, V., Yih, W., Wang, S. I., and Lin, X. V · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D · 2023
Cited alongside, same era.
Keep the conversation going: Fixing 162 out of 337 bugs for $0.42 each using chatgpt
Xia, C. S. and Zhang, L · 2023
Cited alongside, same era.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2023
Cited alongside, same era.
Automatic chain of thought prompting in large language models
Zhang, Z., Zhang, A., Li, M., and Smola, A · 2023
Cited alongside, same era.
Least-to-most prompting enables complex reasoning in large language models
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q. V., and Chi, E. H · 2023
Cited alongside, same era.
Large language models have intrinsic self-correction ability
Liu, D., Nassereldine, A., Yang, Z., Xu, C., Hu, Y., Li, J., Kumar, U., Lee, C., and Xiong, J · 2024
Later among the works it cites.
At which training stage does code data help llms reasoning?
Ma, Y., Liu, Y., Yu, Y., Zhang, Y., Jiang, Y., Wang, C., and Li, S · 2024
Later among the works it cites.
Selfcheck: Using llms to zero-shot check their own step-by-step reasoning
Miao, N., Teh, Y. W., and Rainforth, T · 2024
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Mirzadeh, S., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., and Farajtabar, M · 2024
Later among the works it cites.
Beyond accuracy: Evaluating the reasoning behavior of large language models - A survey
Mondorf, P. and Plank, B · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Podstawski, M., Gianinazzi, L., Gajda, J., Lehmann, T., Niewiadomski, H., Nyczyk, P., and Hoefler, T · 2024
Cited alongside, same era.
Teaching large language models to self-debug
Chen, X., Lin, M., Schärli, N., and Zhou, D · 2024
Cited alongside, same era.
CYCLE: learning to self-refine the code generation
Ding, Y., Min, M. J., Kaiser, G. E., and Ray, B · 2024
Cited alongside, same era.
CRITIC: large language models can self-correct with tool-interactive critiquing
Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Duan, N., and Chen, W · 2024
Cited alongside, same era.
Metagpt: Meta programming for A multi-agent collaborative framework
Hong, S., Zhuge, M., Chen, J., Zheng, X., Cheng, Y., Wang, J., Zhang, C., Wang, Z., Yau, S. K. S., Lin, Z., Zhou, L., Ran, C., Xiao, L., Wu, C., and Schmidhuber, J · 2024
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D · 2024
Cited alongside, same era.
When can llms actually correct their own mistakes? A critical survey of self-correction of llms
Kamoi, R., Zhang, Y., Zhang, N., Han, J., and Zhang, R · 2024
Cited alongside, same era.
Later among the works it cites.
Self-refine instruction-tuning for aligning reasoning in language models
Ranaldi, L. and Freitas, A · 2024
Later among the works it cites.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges
Thakur, A. S., Choudhary, K., Ramayapally, V. S., Vaidyanathan, S., and Hupkes, D · 2024
Later among the works it cites.
Enhancing mathematical reasoning in llms by stepwise correction
Wu, Z., Zeng, Q., Zhang, Z., Tan, Z., Shen, C., and Jiang, M · 2024
Later among the works it cites.
Decompose, analyze and rethink: Solving intricate problems with human-like reasoning cycle
Xue, S., Huang, Z., Liu, J., Lin, X., Ning, Y., Jin, B., Li, X., and Liu, Q · 2024
Later among the works it cites.
Natural language reasoning, A survey
Yu, F., Zhang, H., Tiwari, P., and Wang, B · 2024
Later among the works it cites.
Opencodeinterpreter: Integrating code generation with execution and refinement
Zheng, T., Zhang, G., Shen, T., Liu, X., Lin, B. Y., Fu, J., Chen, W., and Yue, X · 2024
Later among the works it cites.
Evaluation of openai o1: Opportunities and challenges of AGI
Zhong, T., Liu, Z., Pan, Y., Zhang, Y., Zhou, Y., Liang, S., Wu, Z., Lyu, Y., Shu, P., Yu, X., Cao, C., Jiang, H., Chen, H., Li, Y., Chen, J., Hu, H., Liu, Y., Zhao, H., Xu, S., Dai, H., Zhao, L., Zhang, R., Zhao, W., Yang, Z., Chen, J., Wang, P., Ruan, W., Wang, H., Zhao, H., Zhang, J., Ren, Y., Qin, S., Chen, T., Li, J., Zidan, A. H., Jahin, A., Chen, M., Xia, S., Holmes, J., Zhuang, Y., Wang, J., Xu, B., Xia, W., Yu, J., Tang, K., Yang, Y., Sun, B., Yang, T., Lu, G., Wang, X., Chai, L., Li, H., Lu, J., Sun, L., Zhang, X., Ge, B., Hu, X., Zhang, L., Zhou, H., Zhang, L., Zhang, S., Liu, N., Jiang, B., Kong, L., Xiang, Z., Ren, Y., Liu, J., Jiang, X., Bao, Y., Zhang, W., Li, X., Li, G., Liu, W., Shen, D., Sikora, A., Zhai, X., Zhu, D., and Liu, T · 2024
Later among the works it cites.
Gsm8k, 2025
Paperwithcode · 2025
Closest in time.