Fetching the paper…
Reading the bibliography…
Large Reasoning Models (LRMs) represent a breakthrough in AI problem-solving capabilities, but their effectiveness in interactive environments can be limited.
Artificial Intelligence: A Modern Approach
Russell, S. J. and Norvig, P · 1995
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2009
Earlier work this paper cites.
Binomial confidence intervals and contingency tests: Mathematical fundamentals and the evaluation of alternative methods
Wallis, S · 2013
Earlier work this paper cites.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks
Chen, W., Ma, X., Wang, X., and Cohen, W. W · 2023
Earlier work this paper cites.
Large language models are zero-shot reasoners, 2023
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2023
Earlier work this paper cites.
Deductive verification of chain-of-thought reasoning
Ling, Z., Fang, Y., Li, X., Huang, Z., Lee, M., Memisevic, R., and Su, H · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback, 2023
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., and Clark, P · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark, 2023
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models, 2023
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Earlier work this paper cites.
Claude 3.5: A sonnet of progress in ai
Anthropic · 2024
Earlier work this paper cites.
Openai o3 publication breakthrough, 2024
ARC · 2024
Earlier work this paper cites.
Reinventing the amazon q developer agent for software development
Blog, A. D · 2024
Earlier work this paper cites.
Reasoning paths optimization: Learning to reason and explore from diverse paths, 2024
Chia, Y. K., Chen, G., Xu, W., Tuan, L. A., Poria, S., and Bing, L · 2024
Cited alongside, same era.
A survey on in-context learning, 2024
Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., Liu, T., Chang, B., Sun, X., Li, L., and Sui, Z · 2024
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?, 2024
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2024
Cited alongside, same era.
Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., and Narayanan, A · 2024
Cited alongside, same era.
Evaluation of openai o1: Opportunities and challenges of agi
Zhong, T., Liu, Z., Pan, Y., Zhang, Y., Zhou, Y., Liang, S., Wu, Z., Lyu, Y., Shu, P., Yu, X., et al · 2024
Later among the works it cites.
2024 aime i
AoPS · 2025
Closest in time.
Reasoning Language Models: A Blueprint, January 2025
Besta, M., Barth, J., Schreiber, E., Kubicek, A., Catarino, A., Gerstenberger, R., Nyczyk, P., Iff, P., Li, Y., Houliston, S., Sternal, T., Copik, M., Kwaśniewski, G., Müller, J., Flis, l., Eberhard, H., Niewiadomski, H., and Hoefler, T · 2025
Closest in time.
Reasoning model guide
DeepSeek · 2025
Closest in time.
rstar-math: Small llms can master math reasoning with self-evolved deep thinking, 2025
Guan, X., Zhang, L. L., Liu, Y., Shang, N., Sun, Y., Zhu, Y., Yang, F., and Yang, M · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, Y., Gao, P., Wang, X., Liu, J., Shi, Y., Zhang, Z., and Peng, C · 2024
Cited alongside, same era.
Don’t sleep on single-agent systems
Neubig, G · 2024
Cited alongside, same era.
Openai function calling guide
OpenAI · 2024
Cited alongside, same era.
Openai o1-mini: Advancing cost-efficient reasoning
OpenAI · 2024
Cited alongside, same era.
Hyperagent: Generalist software engineering agents to solve coding tasks at scale, 2024
Phan, H. N., Nguyen, T. N., Nguyen, P. X., and Bui, N. D. Q · 2024
Cited alongside, same era.
Swe agents: Empowering software development with ai agents
Research, I · 2024
Cited alongside, same era.
Agentic ai whitepaper, 2024
Smeyatsky, A · 2024
Cited alongside, same era.
Agentless: Demystifying llm-based software engineering agents, 2024
Xia, C. S., Deng, Y., Dunn, S., and Zhang, L · 2024
Cited alongside, same era.
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Search-o1: Agentic search-enhanced large reasoning models, 2025
Li, X., Dong, G., Jin, J., Zhang, Y., Zhou, Y., Zhu, Y., Zhang, P., and Dou, Z · 2025
Closest in time.
Sky-t1: Train your own o1 preview model within $450
NovaSky · 2025
Closest in time.
Gpt-4o mini: Advancing cost-efficient intelligence
OpenAI · 2025
Closest in time.
Introducing swe bench verified
OpenAI · 2025
Closest in time.
Chat api reference
OpenAI · 2025
Closest in time.
Gpt-4o mini model documentation
OpenAI · 2025
Closest in time.
Openai o1
OpenAI · 2025
Closest in time.
Towards large reasoning models: A survey on scaling llm reasoning capabilities, 2025
Xu, F., Hao, Q., Zong, Z., Wang, J., Zhang, Y., Wang, J., Lan, X., Gong, J., Ouyang, T., Meng, F., Shao, C., Yan, Y., Yang, Q., Song, Y., Ren, S., Hu, X., Li, Y., Feng, J., Gao, C., and Li, Y · 2025
Closest in time.