Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable capabilities in complex tasks.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Mirror descent policy optimization
Manan Tomar, Lior Shani, Yonathan Efroni, and Mohammad Ghavamzadeh · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
A survey of embodied ai: From simulators to research tasks
Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2023
Earlier work this paper cites.
Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengling Feng, and Erik Cambria · 2023
Earlier work this paper cites.
Tree-planner: Efficient close-loop task planning with large language models
Mengkang Hu, Yao Mu, Xinmiao Yu, Mingyu Ding, Shiguang Wu, Wenqi Shao, Qiguang Chen, Bin Wang, Yu Qiao, and Ping Luo · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Earlier work this paper cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Earlier work this paper cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2023
Earlier work this paper cites.
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al · 2024
Earlier work this paper cites.
Unlocking the capabilities of thought: A reasoning boundary framework to quantify and optimize chain-of-thought
Qiguang Chen, Libo Qin, Jiaqi Wang, Jingxuan Zhou, and Wanxiang Che · 2024
Earlier work this paper cites.
Distilling reasoning ability from large language models with adaptive thinking
Xiaoshu Chen, Sihang Zhou, Ke Liang, and Xinwang Liu · 2024
Earlier work this paper cites.
Do not think that much for 2+ 3=? on the overthinking of o1-like llms
Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, et al · 2024
Earlier work this paper cites.
Compressed chain of thought: Efficient reasoning through dense representations
Jeffrey Cheng and Benjamin Van Durme · 2024
Earlier work this paper cites.
Mixed distillation helps smaller language models reason better
Li Chenglin, Qianglong Chen, Liangyue Li, Caiyu Wang, Feng Tao, Yicheng Li, Zulong Chen, and Yin Zhang · 2024
Earlier work this paper cites.
A survey on multimodal large language models for autonomous driving
Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Yang Zhou, Kaizhao Liang, Jintai Chen, Juanwu Lu, Zichong Yang, Kuei-Da Liao, et al · 2024
Earlier work this paper cites.
From explicit cot to implicit cot: Learning to internalize cot step by step
Yuntian Deng, Yejin Choi, and Stuart Shieber · 2024
Earlier work this paper cites.
Teaching small language models reasoning through counterfactual distillation
Tao Feng, Yicheng Li, Li Chenglin, Hao Chen, Fei Yu, and Yin Zhang · 2024
Earlier work this paper cites.
Efficiently serving llm reasoning programs with certaindex
Yichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu, Zhongdongming Dai, Aurick Qiao, and Hao Zhang · 2024
Earlier work this paper cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Earlier work this paper cites.
Token-budget-aware llm reasoning
Tingxu Han, Chunrong Fang, Shiyu Zhao, Shiqing Ma, Zhenyu Chen, and Zhenting Wang · 2024
Earlier work this paper cites.
Training large language models to reason in a continuous latent space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian · 2024
Earlier work this paper cites.
Mengkang Hu, Tianxing Chen, Qiguang Chen, Yao Mu, Wenqi Shao, and Ping Luo · 2024
Earlier work this paper cites.
The impact of reasoning step length on large language models
Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, and Mengnan Du · 2024
Earlier work this paper cites.
C3ot: Generating shorter chain-of-thought without compromising effectiveness
Yu Kang, Xianghui Sun, Liangyu Chen, and Wei Zou · 2024
Earlier work this paper cites.
Escape sky-high cost: Early-stopping self-consistency for multi-step reasoning
Yiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan, Xinglin Wang, Bin Sun, Heda Wang, and Kan Li · 2024
Earlier work this paper cites.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han · 2024
Earlier work this paper cites.
Ryan Liu, Jiayi Geng, Addison J Wu, Ilia Sucholutsky, Tania Lombrozo, and Thomas L Griffiths · 2024
Earlier work this paper cites.
Can language models learn to skip steps?
Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang, Yue Zhang, Xipeng Qiu, and Zheng Zhang · 2024
Earlier work this paper cites.
Kivi: A tuning-free asymmetric 2bit quantization for kv cache
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu · 2024
Earlier work this paper cites.
Simpo: Simple preference optimization with a reference-free reward
Yu Meng, Mengzhou Xia, and Danqi Chen · 2024
Earlier work this paper cites.
Routellm: Learning to route llms with preference data
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E Gonzalez, M Waleed Kadous, and Ion Stoica · 2024
Earlier work this paper cites.
Chain-of-action: Faithful and multimodal question answering through large language models
Zhenyu Pan, Haozheng Luo, Manling Li, and Han Liu · 2024
Earlier work this paper cites.
Let’s think dot by dot: Hidden computation in transformer language models
Jacob Pfau, William Merrill, and Samuel R Bowman · 2024
Earlier work this paper cites.
The benefits of a concise chain of thought on problem-solving in large language models
Matthew Renze and Erhan Guven · 2024
Earlier work this paper cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al · 2024
Earlier work this paper cites.
Keep the cost down: A review on methods to optimize llm’s kv-cache consumption
Luohe Shi, Hongyi Zhang, Yao Yao, Zuchao Li, and Hai Zhao · 2024
Earlier work this paper cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Earlier work this paper cites.
Fast best-of-n decoding via speculative rejection
Hanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang, Jiahao Qiu, Ming Yin, Mengdi Wang, Peter Bartlett, and Andrea Zanette · 2024
Earlier work this paper cites.
Reasoning aware self-consistency: Leveraging reasoning paths for efficient llm sampling
Guangya Wan, Yuqi Wu, Jie Chen, and Sheng Li · 2024
Earlier work this paper cites.
Make every penny count: Difficulty-adaptive self-consistency for cost-efficient reasoning
Xinglin Wang, Shaoxiong Feng, Yiwei Li, Peiwen Yuan, Yueqi Zhang, Chuyi Tan, Boyuan Pan, Yao Hu, and Kan Li · 2024
Earlier work this paper cites.
Autotrust: Benchmarking trustworthiness in large vision language models for autonomous driving
Shuo Xing, Hongyuan Hua, Xiangbo Gao, Shenzhe Zhu, Renjie Li, Kexin Tian, Xiaopeng Li, Heng Huang, Tianbao Yang, Zhangyang Wang, et al · 2024
Earlier work this paper cites.
Model merging in llms, mllms, and beyond: Methods, theories, applications and opportunities, 2024
Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, and Dacheng Tao · 2024
Earlier work this paper cites.
Distilling system 2 into system 1
Ping Yu, Jing Xu, Jason Weston, and Ilia Kulikov · 2024
Earlier work this paper cites.
KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches
Jiayi Yuan, Hongyi Liu, Shaochen Zhong, Yu-Neng Chuang, Songchen Li, Guanchu Wang, Duy Le, Hongye Jin, Vipin Chaudhary, Zhaozhuo Xu, Zirui Liu, and Xia Hu · 2024
Earlier work this paper cites.
Small language models need strong verifiers to self-correct reasoning
Yunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee, and Lu Wang · 2024
Earlier work this paper cites.
Probe then retrieve and reason: Distilling probing and reasoning capabilities into smaller language models
Yichun Zhao, Shuheng Zhou, and Huijia Zhu · 2024
Earlier work this paper cites.
Path-consistency: Prefix enhancement for efficient inference in llm
Jiace Zhu, Yingtao Shen, Jie Zhao, and An Zou · 2024
Earlier work this paper cites.
Xunyu Zhu, Jian Li, Can Ma, and Weiping Wang · 2024
Earlier work this paper cites.
First finish search: Efficient test-time scaling in large language models, 2025
Aradhye Agarwal, Ayan Sengupta, and Tanmoy Chakraborty · 2025
Earlier work this paper cites.
L1: Controlling how long a reasoning model thinks with reinforcement learning
Pranjal Aggarwal and Sean Welleck · 2025
Earlier work this paper cites.
Don’t think longer, think wisely: Optimizing thinking dynamics for large reasoning models, 2025
Sohyun An, Ruochen Wang, Tianyi Zhou, and Cho-Jui Hsieh · 2025
Earlier work this paper cites.
Claude 3.7 sonnet, 2023
Anthropic · 2025
Earlier work this paper cites.
Training language models to reason efficiently
Daman Arora and Andrea Zanette · 2025
Earlier work this paper cites.
Sketch-of-thought: Efficient llm reasoning with adaptive cognitive-inspired sketching
Simon A Aytes, Jinheon Baek, and Sung Ju Hwang · 2025
Earlier work this paper cites.
Activation steering for chain-of-thought compression, 2025
Seyedarmin Azizi, Erfan Baghaei Potraghloo, and Massoud Pedram · 2025
Earlier work this paper cites.
Aware first, think less: Dynamic boundary self-awareness drives extreme reasoning efficiency in large language models, 2025
Qiguang Chen, Dengyun Peng, Jinhao Liu, HuiKang Su, Jiannan Guan, Libo Qin, and Wanxiang Che · 2025
Earlier work this paper cites.
Seal: Steerable reasoning calibration of large language models for free
Runjin Chen, Zhenyu Zhang, Junyuan Hong, Souvik Kundu, and Zhangyang Wang · 2025
Earlier work this paper cites.
Inner thinking transformer: Leveraging dynamic depth scaling to foster adaptive internal thinking
Yilong Chen, Junyuan Shang, Zhenyu Zhang, Yanxi Xie, Jiawei Sheng, Tingwen Liu, Shuohuan Wang, Yu Sun, Hua Wu, and Haifeng Wang · 2025
Earlier work this paper cites.
R-stitch: Dynamic trajectory stitching for efficient reasoning, 2025
Zhuokun Chen, Zeren Chen, Jiahao He, Mingkui Tan, Jianfei Cai, and Bohan Zhuang · 2025
Earlier work this paper cites.
Verithinker: Learning to verify makes reasoning model efficient
Zigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu, and Xinchao Wang · 2025
Earlier work this paper cites.
Incentivizing dual process thinking for efficient large language model reasoning
Xiaoxue Cheng, Junyi Li, Zhenduo Zhang, Xinyu Tang, Wayne Xin Zhao, Xinyu Kong, and Zhiqiang Zhang · 2025
Earlier work this paper cites.
Optimizing length compression in large reasoning models, 2025
Zhengxiang Cheng, Dongping Chen, Mingyang Fu, and Tianyi Zhou · 2025
Earlier work this paper cites.
Confident or seek stronger: Exploring uncertainty-based on-device llm routing from benchmarking to generalization, 2025
Yu-Neng Chuang, Leisheng Yu, Guanchu Wang, Lizhe Zhang, Zirui Liu, Xuanting Cai, Yang Sui, Vladimir Braverman, and Xia Hu · 2025
Earlier work this paper cites.
Learning to route llms with confidence tokens, 2025
Yu-Neng Chuang, Helen Zhou, Prathusha Kameswara Sarma, Parikshit Gopalan, John Boccio, Sara Bolouki, and Xia Hu · 2025
Earlier work this paper cites.
Codeforces - competitive programming platform, 2025
Codeforces · 2025
Earlier work this paper cites.
The danger of overthinking: Examining the reasoning-action dilemma in agentic tasks, 2025
Alejandro Cuadron, Dacheng Li, Wenjie Ma, Xingyao Wang, Yichuan Wang, Siyuan Zhuang, Shu Liu, Luis Gaspar Schroeder, Tian Xia, Huanzhi Mao, Nicholas Thumiger, Aditya Desai, Ion Stoica, Ana Klimovic, Graham Neubig, and Joseph E. Gonzalez · 2025
Earlier work this paper cites.
Yingqian Cui, Pengfei He, Jingying Zeng, Hui Liu, Xianfeng Tang, Zhenwei Dai, Yan Han, Chen Luo, Jing Huang, Zhen Li, et al · 2025
Earlier work this paper cites.
Stable reinforcement learning for efficient reasoning, 2025
Muzhi Dai, Shixuan Liu, and Qingyi Si · 2025
Earlier work this paper cites.
S-grpo: Early exit via reinforcement learning in reasoning models
Muzhi Dai, Chenxu Yang, and Qingyi Si · 2025
Earlier work this paper cites.
Do thinking tokens help or trap? towards more efficient large reasoning model, 2025
Bowen Ding, Yuhan Chen, Futing Wang, Lingfeng Ming, and Tao Lin · 2025
Cited alongside, same era.
Best-route: Adaptive llm routing with test-time optimal compute, 2025
Dujian Ding, Ankur Mallick, Shaokun Zhang, Chi Wang, Daniel Madrigal, Mirian Del Carmen Hipolito Garcia, Menglin Xia, Laks V. S. Lakshmanan, Qingyun Wu, and Victor Rühle · 2025
Cited alongside, same era.
Dynamic parallel tree search for efficient llm reasoning
Yifu Ding, Wentao Jiang, Shunyu Liu, Yongcheng Jing, Jinyang Guo, Yingjie Wang, Jing Zhang, Zengmao Wang, Ziwei Liu, Bo Du, et al · 2025
Cited alongside, same era.
Conciserl: Conciseness-guided reinforcement learning for efficient reasoning models
Razvan-Gabriel Dumitru, Darius Peteleaza, Vikas Yadav, and Liangming Pan · 2025
Cited alongside, same era.
Overclocking llm reasoning: Monitoring and controlling thinking path lengths in llms, 2025
Optimizing anytime reasoning via budget relative policy optimization
Penghui Qi, Zichen Liu, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin · 2025
Closest in time.
Concise: Confidence-guided compression in step-by-step efficient reasoning
Ziqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Fandong Meng, Jie Zhou, Ju Ren, and Yaoxue Zhang · 2025
Closest in time.
Optimizing test-time compute via meta reinforcement fine-tuning
Yuxiao Qu, Matthew YR Yang, Amrith Setlur, Lewis Tunstall, Edward Emanuel Beeching, Ruslan Salakhutdinov, and Aviral Kumar · 2025
Closest in time.
Putting the value back in rl: Better test-time scaling by unifying llm reasoners with verifiers, 2025
Kusha Sareen, Morgane M Moss, Alessandro Sordoni, Rishabh Agarwal, and Arian Hosseini · 2025
Closest in time.
Reasoning with latent thoughts: On the power of looped transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Roy Eisenstadt, Itamar Zimerman, and Lior Wolf · 2025
Cited alongside, same era.
Debate only when necessary: Adaptive multiagent collaboration for efficient llm reasoning, 2025
Sugyeong Eo, Hyeonseok Moon, Evelyn Hayoon Zi, Chanjun Park, and Heuiseok Lim · 2025
Cited alongside, same era.
Missing premise exacerbates overthinking: Are reasoning models losing critical thinking skill?
Chenrui Fan, Ming Li, Lichao Sun, and Tianyi Zhou · 2025
Cited alongside, same era.
Cothink: Token-efficient reasoning via instruct models guiding reasoning models, 2025
Siqi Fan, Peng Han, Shuo Shang, Yequan Wang, and Aixin Sun · 2025
Cited alongside, same era.
Thinkless: Llm learns when to think, 2025
Gongfan Fang, Xinyin Ma, and Xinchao Wang · 2025
Cited alongside, same era.
Safemlrm: Demystifying safety in multi-modal large reasoning models
Junfeng Fang, Yukai Wang, Ruipeng Wang, Zijun Yao, Kun Wang, An Zhang, Xiang Wang, and Tat-Seng Chua · 2025
Cited alongside, same era.
Concise reasoning via reinforcement learning
Mehdi Fatemi, Banafsheh Rafiee, Mingjie Tang, and Kartik Talamadupula · 2025
Cited alongside, same era.
Reasoning without self-doubt: More efficient chain-of-thought through certainty probing
Yichao Fu, Junda Chen, Yonghao Zhuang, Zheyu Fu, Ion Stoica, and Hao Zhang · 2025
Cited alongside, same era.
Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, and Sashank J Reddi · 2025
Closest in time.
Hawkeye:efficient reasoning with model collaboration, 2025
Jianshu She, Zhuohao Li, Zhemin Huang, Qi Li, Peiran Xu, Haonan Li, and Qirong Ho · 2025
Closest in time.
Efficient reasoning with hidden thinking
Xuan Shen, Yizhou Wang, Xiangxi Shi, Yanzhi Wang, Pu Zhao, and Jiuxiang Gu · 2025
Closest in time.
Dast: Difficulty-adaptive slow-thinking for large reasoning models
Yi Shen, Jian Zhang, Jieyun Huang, Shuming Shi, Wenjing Zhang, Jiangze Yan, Ning Wang, Kai Wang, and Shiguo Lian · 2025
Closest in time.
Codi: Compressing chain-of-thought into continuous space via self-distillation
Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, and Yulan He · 2025
Closest in time.
Sample more to think less: Group filtered policy optimization for concise reasoning
Vaishnavi Shrivastava, Ahmed Awadallah, Vidhisha Balachandran, Shivam Garg, Harkirat Behl, and Dimitris Papailiopoulos · 2025
Closest in time.
Reasoning path compression: Compressing generation trajectories for efficient llm reasoning
Jiwon Song, Dongwon Jo, Yulhwa Kim, and Jae-Joon Kim · 2025
Closest in time.
Walk before you run! concise llm reasoning via reinforcement learning
Mingyang Song and Mao Zheng · 2025
Closest in time.
Accelerated test-time scaling with model-free speculative sampling
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin, Aram Galstyan, and Sravan Babu Bodapati · 2025
Closest in time.
Towards reasoning ability of small language models
Gaurav Srivastava, Shuxiang Cao, and Xuan Wang · 2025
Closest in time.
Token assorted: Mixing latent and text tokens for improved language model reasoning
DiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao, Yuandong Tian, and Qinqing Zheng · 2025
Closest in time.
Meta-reasoner: Dynamic guidance for optimized inference-time reasoning in large language models
Yuan Sui, Yufei He, Tri Cao, Simeng Han, and Bryan Hooi · 2025
Closest in time.
Tinyr1-32b-preview: Boosting accuracy with branch-merge distillation
Lin Sun, Guangxiang Zhao, Xiaoqi Jian, Yuhan Wu, Weihong Lin, Yongfu Zhu, Linglin Zhang, Jinzhu Wu, Junfeng Ran, Sai-er Hu, et al · 2025
Closest in time.
Think silently, think fast: Dynamic latent compression of llm reasoning chains
Wenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, and Ruihua Song · 2025
Closest in time.
Think before recommend: Unleashing the latent reasoning power for sequential recommendation
Jiakai Tang, Sunhao Dai, Teng Shi, Jun Xu, Xu Chen, Wen Chen, Wu Jian, and Yuning Jiang · 2025
Closest in time.
Concisehint: Boosting efficient reasoning via continuous concise hints during generation, 2025
Siao Tang, Xinyin Ma, Gongfan Fang, and Xinchao Wang · 2025
Closest in time.
Confidence improves self-consistency in llms
Amir Taubenfeld, Tom Sheffer, Eran Ofek, Amir Feder, Ariel Goldstein, Zorik Gekhman, and Gal Yona · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
Qwq-32b-preview
Qwen Team · 2025
Closest in time.
Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl, 2025
Songjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian, Linjing Li, Xiangyuan Lan, and Dongbin Zhao · 2025
Closest in time.
Adapthink: Adaptive thinking preferences for reasoning language model, 2025
Xu Wan, Wei Wang, Wenyue Xu, Wotao Yin, Jie Song, and Mingyang Sun · 2025
Closest in time.
Wait, we don’t need to "wait"! removing thinking tokens improves reasoning efficiency, 2025
Chenlong Wang, Yuanning Feng, Dongping Chen, Zhaoyang Chu, Ranjay Krishna, and Tianyi Zhou · 2025
Closest in time.
Efficient reasoning for llms through speculative chain-of-thought, 2025
Jikai Wang, Juntao Li, Jianye Hou, Bowen Yan, Lijun Wu, and Min Zhang · 2025
Closest in time.
Think deep, think fast: Investigating efficiency of verifier-free inference-time-scaling methods, 2025
Junlin Wang, Shang Zhu, Jon Saad-Falcon, Ben Athiwaratkun, Qingyang Wu, Jue Wang, Shuaiwen Leon Song, Ce Zhang, Bhuwan Dhingra, and James Zou · 2025
Closest in time.
Value-guided search for efficient chain-of-thought reasoning, 2025
Kaiwen Wang, Jin Peng Zhou, Jonathan Chang, Zhaolin Gao, Nathan Kallus, Kianté Brantley, and Wen Sun · 2025
Closest in time.
Every rollout counts: Optimal resource allocation for efficient test-time scaling, 2025
Xinglin Wang, Yiwei Li, Shaoxiong Feng, Peiwen Yuan, Yueqi Zhang, Jiayi Shi, Chuyi Tan, Boyuan Pan, Yao Hu, and Kan Li · 2025
Closest in time.
R1-compress: Long chain-of-thought compression via chunk compression and search
Yibo Wang, Li Shen, Huanjin Yao, Tiansheng Huang, Rui Liu, Naiqiang Tan, Jiaxing Huang, Kai Zhang, and Dacheng Tao · 2025
Closest in time.
Sampling-efficient test-time scaling: Self-estimating the best-of-n sampling in early decoding
Yiming Wang, Pei Zhang, Siyuan Huang, Baosong Yang, Zhuosheng Zhang, Fei Huang, and Rui Wang · 2025
Closest in time.
Accelerating large language model reasoning via speculative search, 2025
Zhihai Wang, Jie Wang, Jilai Pan, Xilin Xia, Huiling Zhen, Mingxuan Yuan, Jianye Hao, and Feng Wu · 2025
Closest in time.
Light-r1: Curriculum sft, dpo and rl for long cot from scratch and beyond
Liang Wen, Yunke Cai, Fenrui Xiao, Xin He, Qi An, Zhenyu Duan, Yimin Du, Junchen Liu, Lifu Tang, Xiaowei Lv, et al · 2025
Closest in time.
Unlocking efficient long-to-short llm reasoning with model merging, 2025
Han Wu, Yuxuan Yao, Shuqi Liu, Zehua Liu, Xiaojin Fu, Xiongwei Han, Xing Li, Hui-Ling Zhen, Tao Zhong, and Mingxuan Yuan · 2025
Closest in time.
When more is less: Understanding chain-of-thought length in llms
Yuyang Wu, Yifei Wang, Tianqi Du, Stefanie Jegelka, and Yisen Wang · 2025
Closest in time.
Tokenskip: Controllable chain-of-thought compression in llms
Heming Xia, Yongqi Li, Chak Tou Leong, Wenjie Wang, and Wenjie Li · 2025
Closest in time.
Can atomic step decomposition enhance the self-structured reasoning of multimodal large models?
Kun Xiang, Zhili Liu, Zihao Jiang, Yunshuang Nie, Kaixin Cai, Yiyang Yin, Runhui Huang, Haoxiang Fan, Hanhui Li, Weiran Huang, et al · 2025
Closest in time.
Just enough thinking: Efficient reasoning with adaptive length penalties reinforcement learning, 2025
Violet Xiang, Chase Blagden, Rafael Rafailov, Nathan Lile, Sang Truong, Chelsea Finn, and Nick Haber · 2025
Closest in time.
Limopro: Reasoning refinement for efficient and effective test-time scaling
Yang Xiao, Jiashuo Wang, Ruifeng Yuan, Chunpu Xu, Kaishuai Xu, Wenjie Li, and Pengfei Liu · 2025
Closest in time.
Openemma: Open-source multimodal model for end-to-end autonomous driving
Shuo Xing, Chengyuan Qian, Yuping Wang, Hongyuan Hua, Kexin Tian, Yang Zhou, and Zhengzhong Tu · 2025
Closest in time.
Can large vision language models read maps like a human?
Shuo Xing, Zezhou Sun, Shuangyu Xie, Kaiyuan Chen, Yanjia Huang, Yuping Wang, Jiachen Li, Dezhen Song, and Zhengzhong Tu · 2025
Closest in time.
Towards large reasoning models: A survey of reinforced reasoning with large language models
Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, et al · 2025
Closest in time.
Twt: Thinking without tokens by habitual reasoning distillation with multi-teachers’ guidance, 2025
Jingxian Xu, Mengyu Zhou, Weichang Liu, Hanbing Liu, Shi Han, and Dongmei Zhang · 2025
Closest in time.
Chain of draft: Thinking faster by writing less
Silei Xu, Wenhao Xie, Lingxiao Zhao, and Pengcheng He · 2025
Closest in time.
A*-thought: Efficient reasoning via bidirectional compression for low-resource settings, 2025
Xiaoang Xu, Shuo Wang, Xu Han, Zhenghao Liu, Huijia Wu, Peipei Li, Zhiyuan Liu, Maosong Sun, and Zhaofeng He · 2025
Closest in time.
Softcot: Soft chain-of-thought for efficient reasoning with llms
Yige Xu, Xu Guo, Zhiwei Zeng, and Chunyan Miao · 2025
Closest in time.
Scalable chain of thoughts via elastic reasoning
Yuhui Xu, Hanze Dong, Lei Wang, Doyen Sahoo, Junnan Li, and Caiming Xiong · 2025
Closest in time.
Mur: Momentum uncertainty guided reasoning for large language models, 2025
Hang Yan, Fangzhi Xu, Rongman Xu, Yifei Li, Jian Zhang, Haoran Luo, Xiaobao Wu, Luu Anh Tuan, Haiteng Zhao, Qika Lin, and Jun Liu · 2025
Closest in time.
Inftythink: Breaking the length limits of long-context reasoning in large language models
Yuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang, Mengdi Zhang, Jian Shao, and Yueting Zhuang · 2025
Closest in time.
Test-time prompt intervention, 2025
Chenxu Yang, Qingyi Si, Mz Dai, Dingyu Yao, Mingyu Zheng, Minghui Chen, Zheng Lin, and Weiping Wang · 2025
Closest in time.
Think when you need: Self-adaptive chain-of-thought learning, 2025
Junjie Yang, Ke Lin, and Xing Yu · 2025
Closest in time.
Speculative thinking: Enhancing small-model reasoning with large model guidance at inference time
Wang Yang, Xiang Yue, Vipin Chaudhary, and Xiaotian Han · 2025
Closest in time.
Towards thinking-optimal scaling of test-time compute for llm reasoning, 2025
Wenkai Yang, Shuming Ma, Yankai Lin, and Furu Wei · 2025
Closest in time.
Limo: Less is more for reasoning, 2025
Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, and Pengfei Liu · 2025
Closest in time.
Demystifying long chain-of-thought reasoning in llms
Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, and Xiang Yue · 2025
Closest in time.
Bin Yu, Hang Yuan, Haotian Li, Xueyin Xu, Yuliang Wei, Bailing Wang, Weizhen Qi, and Kai Chen · 2025
Closest in time.
Causal sufficiency and necessity improves chain-of-thought reasoning, 2025
Xiangning Yu, Zhuohan Wang, Linyi Yang, Haoxuan Li, Anjie Liu, Xiao Xue, Jun Wang, and Mengyue Yang · 2025
Closest in time.
Premise: Scalable and strategic prompt optimization for efficient mathematical reasoning in large models, 2025
Ye Yu, Yaoning Yu, and Haohan Wang · 2025
Closest in time.
Back attention: Understanding and enhancing multi-hop reasoning in large language models
Zeping Yu, Yonatan Belinkov, and Sophia Ananiadou · 2025
Closest in time.
Z1: Efficient test-time scaling with code, 2025
Zhaojian Yu, Yinghao Wu, Yilun Zhao, Arman Cohan, and Xiao-Ping Zhang · 2025
Closest in time.
Think smarter not harder: Adaptive reasoning with inference aware optimization
Zishun Yu, Tengyu Xu, Di Jin, Karthik Abinav Sankararaman, Yun He, Wenxuan Zhou, Zhouhao Zeng, Eryk Helenowski, Chen Zhu, Sinong Wang, et al · 2025
Closest in time.
Efficient rl training for reasoning models via length-aware optimization
Danlong Yuan, Tian Xie, Shaohan Huang, Zhuocheng Gong, Huishuai Zhang, Chong Luo, Furu Wei, and Dongyan Zhao · 2025
Closest in time.
Not all tokens are what you need in thinking
Hang Yuan, Bin Yu, Haotian Li, Shijun Yang, Christina Dan Wang, Zhou Yu, Xueyin Xu, Weizhen Qi, and Kai Chen · 2025
Closest in time.
Promoting efficient reasoning with verifiable stepwise reward, 2025
Chuhuai Yue, Chengqi Dong, Yinan Gao, Hang He, Jiajun Chai, Guojun Yin, and Wei Lin · 2025
Closest in time.
Pruning the unsurprising: Efficient code reasoning via first-token surprisal, 2025
Wenhao Zeng, Yaoning Wang, Chao Hu, Yuling Shi, Chengcheng Wan, Hongyu Zhang, and Xiaodong Gu · 2025
Closest in time.
Adaptthink: Reasoning models can learn when to think, 2025
Jiajie Zhang, Nianyi Lin, Lei Hou, Ling Feng, and Juanzi Li · 2025
Closest in time.
Lightthinker: Thinking step-by-step compression
Jintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo, Shuofei Qiao, Lun Du, Da Zheng, Huajun Chen, and Ningyu Zhang · 2025
Closest in time.
Alphaone: Reasoning models thinking slow and fast at test time
Junyu Zhang, Runpei Dong, Han Wang, Xuying Ning, Haoran Geng, Peihao Li, Xialin He, Yutong Bai, Jitendra Malik, Saurabh Gupta, et al · 2025
Closest in time.
Nan Zhang, Yusen Zhang, Prasenjit Mitra, and Rui Zhang · 2025
Closest in time.
Long or short cot? investigating instance-level switch of large reasoning models, 2025
Ruiqi Zhang, Changyi Xiao, and Yixin Cao · 2025
Closest in time.
Othink-r1: Intrinsic fast/slow thinking mode switching for over-reasoning mitigation, 2025
Shengjia Zhang, Junjie Wu, Jiawei Chen, Changwang Zhang, Xingyu Lou, Wangchunshu Zhou, Sheng Zhou, Can Wang, and Jun Wang · 2025
Closest in time.
Synapseroute: An auto-route switching framework on dual-state large language model, 2025
Wencheng Zhang, Shiqin Qiao, Lingjie Luo, Yinfeng Li, Chuanyang Zheng, Qian Xu, Meng Li, Yong Gui, Yijun He, Jianing Qiu, Jindong Hong, and Jiankai Sun · 2025
Closest in time.
S1-bench: A simple benchmark for evaluating system 1 thinking capability of large reasoning models
Wenyuan Zhang, Shuaiyi Nie, Xinghua Zhang, Zefeng Zhang, and Tingwen Liu · 2025
Closest in time.
When to continue thinking: Adaptive thinking mode switching for efficient reasoning, 2025
Xiaoyun Zhang, Jingqing Ruan, Xing Ma, Yawen Zhu, Haodong Zhao, Hao Li, Jiansong Chen, Ke Zeng, and Xunliang Cai · 2025
Closest in time.
Making small language models efficient reasoners: Intervention, supervision, reinforcement
Xuechen Zhang, Zijian Huang, Chenshun Ni, Ziyang Xiong, Jiasi Chen, and Samet Oymak · 2025
Closest in time.
Saber: Switchable and balanced training for efficient llm reasoning, 2025
Kai Zhao, Yanjun Zhao, Jiaming Song, Shien He, Lusheng Zhang, Qiang Zhang, and Tianjiao Li · 2025
Closest in time.
Shangziqi Zhao, Jiahao Yuan, Guisong Yang, and Usman Naseem · 2025
Closest in time.
Exploring and exploiting the inherent efficiency within large reasoning models for self-guided efficiency enhancement, 2025
Weixiang Zhao, Jiahe Guo, Yang Deng, Xingyu Sui, Yulin Hu, Yanyan Zhao, Wanxiang Che, Bing Qin, Tat-Seng Chua, and Ting Liu · 2025
Closest in time.
Bridging internal probability and self-consistency for effective and efficient llm reasoning
Zhi Zhou, Tan Yuhao, Zenan Li, Yuan Yao, Lan-Zhe Guo, Xiaoxing Ma, and Yu-Feng Li · 2025
Closest in time.
Accelerating chain-of-thought reasoning: When goal-gradient importance meets dynamic skipping
Ren Zhuang, Ben Wang, and Shuifa Sun · 2025
Closest in time.
A technical study into 0.5b reasoning language models, 2025
Xialie Zhuang, Peixian Ma, Zhikai Jia, Shiwei Liu, and Zheng Cao · 2025
Closest in time.