Fetching the paper…
Reading the bibliography…
Advanced reasoning in large language models has achieved remarkable performance on challenging tasks, but the prevailing long-context reasoning paradigm faces critical limitations: quadratic computational scaling with sequence length, reasoning constrained by maximum context boundaries, and performance degradation beyond pre-training context windows.
Program synthesis with large language models, August 2021
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, July 2021
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, et al · 2021
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao · 2022
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, et al · 2022
Earlier work this paper cites.
Extending context window of large language models via positional interpolation, June 2023
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian · 2023
Earlier work this paper cites.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang · 2023
Earlier work this paper cites.
Things i’m learning while training superhot
kaiokendev · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, et al · 2023
Earlier work this paper cites.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, et al · 2023
Earlier work this paper cites.
Amc23 ⋅ \cdot datasets at hugging face
mathai · 2023
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding, November 2023
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2023
Earlier work this paper cites.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik R. Narasimhan · 2023
Earlier work this paper cites.
Metamath: Bootstrap your own mathematical questions for large language models
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, et al · 2023
Earlier work this paper cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, et al · 2023
Earlier work this paper cites.
Recurrentgpt: Interactive generation of (arbitrarily) long text, May 2023
Wangchunshu Zhou, Yuchen Eleanor Jiang, Peng Cui, Tiannan Wang, Zhenxin Xiao, et al · 2023
Earlier work this paper cites.
Large language monkeys: Scaling inference compute with repeated sampling, July 2024
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, and Azalia Mirhoseini · 2024
Earlier work this paper cites.
Alphamath almost zero: Process supervision without process
Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan · 2024
Earlier work this paper cites.
Compressed chain of thought: Efficient reasoning through dense representations, December 2024
Jeffrey Cheng and Benjamin Van Durme · 2024
Earlier work this paper cites.
The llama 3 herd of models, November 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, et al · 2024
Earlier work this paper cites.
Deepseek-coder: When the large language model meets programming – the rise of code intelligence, January 2024
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, et al · 2024
Earlier work this paper cites.
C3ot: Generating shorter chain-of-thought without compromising effectiveness, December 2024
Yu Kang, Xianghui Sun, Liangyu Chen, and Wei Zou · 2024
Earlier work this paper cites.
Babilong: Testing the limits of llms with long context reasoning-in-a-haystack
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Igorevich Sorokin, Artyom Sorokin, and Mikhail Burtsev · 2024
Earlier work this paper cites.
Can language models learn to skip steps?
Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang, Yue Zhang, Xipeng Qiu, and Zheng Zhang · 2024
Cited alongside, same era.
Introducing openai o1
OpenAI · 2024
Cited alongside, same era.
Learning to reason with llms
OpenAI · 2024
Cited alongside, same era.
Anchor-based large language models
Jianhui Pang, Fanghua Ye, Derek Wong, Xin He, Wanshun Chen, et al · 2024
Cited alongside, same era.
Qwq: Reflect deeply on the boundaries of the unknown
Qwen Team · 2024
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, et al · 2024
Cited alongside, same era.
Code llama: Open foundation models for code, January 2024
Token-budget-aware llm reasoning
Tingxu Han, Zhenting Wang, Chunrong Fang, Shiyu Zhao, Shiqing Ma, et al · 2025
Closest in time.
Training large language models to reason in a continuous latent space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason E. Weston, and Yuandong Tian · 2025
Closest in time.
Numinamath-1.5
Jia LI, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, et al · 2025
Closest in time.
Logicpro: Improving complex logical reasoning via program-guided learning
Jin Jiang, Yuchen Yan, Yang Liu, Jianing Wang, Shuai Peng, et al · 2025
Closest in time.
Kimi k1.5: Scaling reinforcement learning with llms, January 2025
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, et al · 2025
Closest in time.
Ada-r1: Hybrid-cot via bi-level adaptive reasoning optimization
Haotian Luo, Haiying He, Yibo Wang, Jinluan Yang, Rui Liu, et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, et al · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, February 2024
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, et al · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling parameters for reasoning
Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Cited alongside, same era.
A survey of reasoning with foundation models, January 2024
Jiankai Sun, Chuanyang Zheng, Enze Xie, Zhengying Liu, Ruihang Chu, et al · 2024
Cited alongside, same era.
Alphazero-like tree-search can guide large language model decoding and training
Ziyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer, Ying Wen, Weinan Zhang, and Jun Wang · 2024
Cited alongside, same era.
Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, September 2024
An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, et al · 2024
Cited alongside, same era.
Hierarchical budget policy optimization for adaptive reasoning, August 2025
Shangke Lyu, Linjuan Wu, Yuchen Yan, Xingyu Wu, Hao Li, et al · 2025
Closest in time.
Cot-valve: Length-compressible chain-of-thought tuning
Xinyin Ma, Guangnian Wan, Runpeng Yu, Gongfan Fang, Xinchao Wang, et al · 2025
Closest in time.
S1: Simple test-time scaling
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, et al · 2025
Closest in time.
Openr1-math-220k
open-r1 · 2025
Closest in time.
Open thoughts
Open Thoughts Team · 2025
Closest in time.
Qwen2.5 technical report, January 2025
Qwen, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, et al · 2025
Closest in time.
Efficient reasoning with hidden thinking, January 2025
Xuan Shen, Yizhou Wang, Xiangxi Shi, Yanzhi Wang, Pu Zhao, and Jiuxiang Gu · 2025
Closest in time.
Thoughts are all over the place: On the underthinking of long reasoning models
Yue Wang, Qiuzhi Liu, Jiahao Xu, Tian Liang, Xingyu Chen, et al · 2025
Closest in time.
Lapo: Internalizing reasoning efficiency via length-adaptive policy optimization, August 2025
Xingyu Wu, Yuchen Yan, Shangke Lyu, Linjuan Wu, Yiwen Qiu, et al · 2025
Closest in time.
Tokenskip: Controllable chain-of-thought compression in llms
Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, Wenjie Li, et al · 2025
Closest in time.
S^3cmath: Spontaneous step-level self-correction makes large language models better mathematical reasoners
Yuchen Yan, Jin Jiang, Yang Liu, Yixin Cao, Xin Xu, et al · 2025
Closest in time.
Limo: Less is more for reasoning
Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, and Pengfei Liu · 2025
Closest in time.
Lv-eval: A balanced long-context benchmark with 5 length levels up to 256k
Tao Yuan, Xuefei Ning, Dong Zhou, Zhijie Yang, Shiyao Li, et al · 2025
Closest in time.
Adaptthink: Reasoning models can learn when to think
Jiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng, Juanzi Li, et al · 2025
Closest in time.
Lightthinker: Thinking step-by-step compression
Jintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo, Shuofei Qiao, et al · 2025
Closest in time.
Inftythink+: Effective and efficient infinite-horizon reasoning via reinforcement learning, February 2026
Yuchen Yan, Liang Jiang, Jin Jiang, Shuaicheng Li, Zujie Wen, et al · 2026
Closest in time.
Mathodyssey: Benchmarking mathematical problem-solving skills in large language models using odyssey math data
Meng Fang, Xiangpeng Wan, Fei Lu, Fei Xing, and Kai Zou · 2052
Closest in time.