Fetching the paper…
Reading the bibliography…
Recent advancements in large language models (LLMs) have demonstrated remarkable reasoning capabilities.
A statistical distribution function of wide applicability
Weibull, W · 1951
Earlier work this paper cites.
Evaluation metrics for language models
Chen, S. F., Beeferman, D., and Rosenfeld, R · 1998
Earlier work this paper cites.
The truncated mean of an asymmetric distribution
Marazzi, A. and Ruffieux, C · 1999
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D. M., Ng, A. Y., and Jordan, M. I · 2003
Earlier work this paper cites.
Towards open set deep networks
Bendale, A. and Boult, T. E · 2016
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Useful inequalities, 2021
Kozma, L · 2021
Earlier work this paper cites.
Automatically checking semantic equivalence between versions of large-scale c projects
Malík, V. and Vojnar, T · 2021
Earlier work this paper cites.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., Showk, S. E., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V. V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
Let’s sample step by step: Adaptive-consistency for efficient reasoning and coding with llms
Aggarwal, A. M. P., Yang, Y., and Mausam · 2023
Earlier work this paper cites.
Complexity-based prompting for multi-step reasoning
Fu, Y., Peng, H., Sabharwal, A., Clark, P., and Khot, T · 2023
Earlier work this paper cites.
Look before you leap: An exploratory study of uncertainty measurement for large language models
Huang, Y., Song, J., Wang, Z., Zhao, S., Chen, H., Juefei-Xu, F., and Ma, L · 2023
Cited alongside, same era.
Autoplan: Automatic planning of interactive decision-making tasks with large language models
Ouyang, S. and Li, L · 2023
Cited alongside, same era.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. D · 2023
Cited alongside, same era.
On the planning abilities of large language models - A critical investigation
Valmeekam, K., Marquez, M., Sreedharan, S., and Kambhampati, S · 2023
Cited alongside, same era.
Automatic chain of thought prompting in large language models
Zhang, Z., Zhang, A., Li, M., and Smola, A · 2023
Cited alongside, same era.
Ensembling large language models with process reward-guided tree search for better complex reasoning
Park, S., Liu, X., Gong, Y., and Choi, E · 2024
Later among the works it cites.
Integrating human expertise & automated methods for a dynamic and multi-parametric evaluation of large language models’ feasibility in clinical decision-making
Sblendorio, E., Dentamaro, V., Cascio, A. L., Germini, F., Piredda, M., and Cicolini, G · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., and Guo, D · 2024
Later among the works it cites.
Dynamic self-consistency: Leveraging reasoning paths for efficient LLM sampling
Wan, G., Wu, Y., Chen, J., and Li, S · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Least-to-most prompting enables complex reasoning in large language models
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q. V., and Chi, E. H · 2023
Cited alongside, same era.
To believe or not to believe your LLM
Abbasi-Yadkori, Y., Kuzborskij, I., György, A., and Szepesvári, C · 2024
Cited alongside, same era.
Cycles of thought: Measuring LLM confidence through stable explanations
Becker, E. and Soatto, S · 2024
Cited alongside, same era.
Steering large language models between code execution and textual reasoning
Chen, Y., Jhamtani, H., Sharma, S., Fan, C., and Wang, C · 2024
Cited alongside, same era.
Relic: Investigating large language model responses using self-consistency
Cheng, F., Zouhar, V., Arora, S., Sachan, M., Strobelt, H., and El-Assady, M · 2024
Cited alongside, same era.
Plug-and-play policy planner for large language model powered dialogue agents
Deng, Y., Zhang, W., Lam, W., Ng, S., and Chua, T · 2024
Cited alongside, same era.
Chain-of-verification reduces hallucination in large language models
Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., and Weston, J · 2024
Cited alongside, same era.
Wang, X., Feng, S., Li, Y., Yuan, P., Zhang, Y., Pan, B., Wang, H., Hu, Y., and Li, K · 2024
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B · 2024
Later among the works it cites.
Confidence calibration and rationalization for LLMs via multi-agent deliberation
Yang, R., Rajagopal, D., Hayati, S. A., Hu, B., and Kang, D · 2024
Later among the works it cites.
Learning from correctness without prompting makes LLM efficient reasoner
YAO, Y., Wu, H., Guo, Z., Biyan, Z., Gao, J., Luo, S., Hou, H., Fu, X., and Song, L · 2024
Later among the works it cites.
Internlm-math: Open math large language models toward verifiable reasoning, 2024
Ying, H., Zhang, S., Li, L., Zhou, Z., Shao, Y., Fei, Z., Ma, Y., Hong, J., Liu, K., Wang, Z., Wang, Y., Wu, Z., Li, S., Zhou, F., Liu, H., Zhang, S., Zhang, W., Yan, H., Qiu, X., Wang, J., Chen, K., and Lin, D · 2024
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W · 2024
Later among the works it cites.
Aime problems 1983 to 2024, 2024
Zamil, P. and Rabby, G · 2024
Later among the works it cites.
ReST-MCTS*: LLM self-training via process reward guided tree search
Zhang, D., Zhoubian, S., Hu, Z., Yue, Y., Dong, Y., and Tang, J · 2024
Later among the works it cites.
Fact-and-reflection (far) improves confidence calibration of large language models
Zhao, X., Zhang, H., Pan, X., Yao, W., Yu, D., Wu, T., and Chen, J · 2024
Later among the works it cites.
rstar-math: Small llms can master math reasoning with self-evolved deep thinking, 2025
Guan, X., Zhang, L. L., Liu, Y., Shang, N., Sun, Y., Zhu, Y., Yang, F., and Yang, M · 2025
Closest in time.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D · 2025
Closest in time.
Evaluating the evaluator: Measuring LLMs’ adherence to task evaluation instructions
Murugadoss, B., Poelitz, C., Drosos, I., Le, V., McKenna, N., Negreanu, C. S., Parnin, C., and Sarkar, A · 2025
Closest in time.