Fetching the paper…
Reading the bibliography…
Large reasoning models (LRMs) like OpenAI o1 and DeepSeek R1 have demonstrated impressive performance on complex reasoning tasks like mathematics and programming with long Chain-of-Thought (CoT) reasoning sequences (slow-thinking), compared with traditional large language models (fast-thinking).
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 1903
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019 · 1905
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020 · 2007
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a · 2009
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
Explanations for CommonsenseQA: New Dataset and Models
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021 · 2021
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021 · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, et al. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021 · 2021
Earlier work this paper cites.
A diverse corpus for evaluating and developing english math word problem solvers
Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2021 · 2021
Earlier work this paper cites.
Are nlp models really able to solve simple math word problems?
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2021
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022 · 2022
Earlier work this paper cites.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022 · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei. 2022 · 2022
Earlier work this paper cites.
Minif2f: a cross-system benchmark for formal olympiad-level mathematics
Kunhao Zheng, Jesse Michael Han, and Stanislas Polu. 2022 · 2022
Earlier work this paper cites.
Proofnet: Autoformalizing and formally proving undergraduate-level mathematics
Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W. Ayers, Dragomir Radev, and Jeremy Avigad. 2023 · 2023
Earlier work this paper cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023 · 2023
Earlier work this paper cites.
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023 · 2023
Earlier work this paper cites.
Teaching small language models to reason
Lucie Charlotte Magister, Jonathan Mallinson, Jakub Adamek, Eric Malmi, and Aliaksei Severyn. 2023 · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2023 · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Compressed chain of thought: Efficient reasoning through dense representations
Jeffrey Cheng and Benjamin Van Durme. 2024 · 2024
Earlier work this paper cites.
From explicit cot to implicit cot: Learning to internalize cot step by step
Yuntian Deng, Yejin Choi, and Stuart Shieber. 2024 · 2024
Earlier work this paper cites.
Hybrid llm: Cost-efficient and quality-aware query routing
Dujian Ding, Ankur Mallick, Chi Wang, Robert Sim, Subhabrata Mukherjee, Victor Ruhle, Laks V. S. Lakshmanan, and Ahmed Hassan Awadallah. 2024 · 2024
Earlier work this paper cites.
Omni-math: A universal olympiad level mathematic benchmark for large language models
Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, and Baobao Chang. 2024 · 2024
Earlier work this paper cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al. 2024 · 2024
Earlier work this paper cites.
Training large language models to reason in a continuous latent space
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. 2024 · 2024
Earlier work this paper cites.
Routerbench: A benchmark for multi-llm routing system
Qitian Jason Hu, Jacob Bieker, Xiuyu Li, Nan Jiang, Benjamin Keigwin, Gaurav Ranganath, Kurt Keutzer, and Shriyash Kaustubh Upadhyay. 2024 · 2024
Earlier work this paper cites.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024 · 2024
Earlier work this paper cites.
Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models
Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri. 2024 · 2024
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024 · 2024
Earlier work this paper cites.
The impact of reasoning step length on large language models
Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, and Mengnan Du. 2024 · 2024
Earlier work this paper cites.
C3ot: Generating shorter chain-of-thought without compromising effectiveness
Yu Kang, Xianghui Sun, Liangyu Chen, and Wei Zou. 2024 · 2024
Earlier work this paper cites.
From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Earlier work this paper cites.
Mario: Math reasoning with code interpreter output – a reproducible pipeline
Minpeng Liao, Wei Luo, Chengxi Li, Jing Wu, and Kai Fan. 2024 · 2024
Earlier work this paper cites.
OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, et al. 2024 · 2024
Earlier work this paper cites.
Dynathink: Fast or slow? a dynamic decision-making framework for large language models
Jiabao Pan, Yan Zhang, Chen Zhang, Zuozhu Liu, Hongwei Wang, and Haizhou Li. 2024 · 2024
Earlier work this paper cites.
Steering llama 2 via contrastive activation addition
Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. 2024 · 2024
Earlier work this paper cites.
Let’s think dot by dot: Hidden computation in transformer language models
Jacob Pfau, William Merrill, and Samuel R. Bowman. 2024 · 2024
Earlier work this paper cites.
Sysbench: Can large language models follow system messages?
Yanzhao Qin, Tao Zhang, Tao Zhang, Yanjun Shen, Wenjing Luo, Haoze Sun, Yan Zhang, Yujing Qiao, Weipeng Chen, Zenan Zhou, Wentao Zhang, and Bin Cui. 2024 · 2024
Earlier work this paper cites.
The benefits of a concise chain of thought on problem-solving in large language models
Matthew Renze and Erhan Guven. 2024 · 2024
Earlier work this paper cites.
Mathscale: Scaling instruction tuning for mathematical reasoning
Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei. 2024 · 2024
Earlier work this paper cites.
Measuring short-form factuality in large language models
Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. 2024 · 2024
Cited alongside, same era.
Travelplanner: A benchmark for real-world planning with language agents
Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su. 2024 · 2024
Cited alongside, same era.
Hdflow: Enhancing llm complex problem-solving with hybrid thinking and dynamic workflows
Wenlin Yao, Haitao Mi, and Dong Yu. 2024 · 2024
Cited alongside, same era.
First finish search: Efficient test-time scaling in large language models
Aradhye Agarwal, Ayan Sengupta, and Tanmoy Chakraborty. 2025 · 2025
Cited alongside, same era.
Answer convergence as a signal for early stopping in reasoning
Xin Liu and Lu Wang. 2025 · 2025
Closest in time.
Adacot: Pareto-optimal adaptive chain-of-thought triggering via reinforcement learning
Chenwei Lou, Zewei Sun, Xinnian Liang, Meng Qu, Wei Shen, Wenqi Wang, Yuntao Li, Qingping Yang, and Shuangzhi Wu. 2025 · 2025
Closest in time.
Jinghui Lu, Haiyang Yu, Siliang Xu, Shiwei Ran, Guozhi Tang, Siqi Wang, Bin Shan, Teng Fu, Hao Feng, Jingqun Tang, Han Wang, and Can Huang. 2025 · 2025
Closest in time.
Sadegh Mahdavi, Muchen Li, Kaiwen Liu, Christos Thrampoulidis, Leonid Sigal, and Renjie Liao. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pranjal Aggarwal and Sean Welleck. 2025 · 2025
Cited alongside, same era.
Nemotron-crossthink: Scaling self-learning beyond math reasoning
Syeda Nahida Akter, Shrimai Prabhumoye, Matvei Novikov, Seungju Han, Ying Lin, Evelina Bakhturina, Eric Nyberg, Yejin Choi, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2025 · 2025
Cited alongside, same era.
Reasoning on a budget: A survey of adaptive and controllable test-time compute in llms
Mohammad Ali Alomrani, Yingxue Zhang, Derek Li, Qianyi Sun, Soumyasundar Pal, Zhanguang Zhang, Yaochen Hu, Rohan Deepak Ajwani, Antonios Valkanas, Raika Karimi, Peng Cheng, Yunzhou Wang, Pengyi Liao, Hanrui Huang, Bin Wang, Jianye Hao, and Mark Coates. 2025 · 2025
Cited alongside, same era.
Don’t think longer, think wisely: Optimizing thinking dynamics for large reasoning models
Sohyun An, Ruochen Wang, Tianyi Zhou, and Cho-Jui Hsieh. 2025 · 2025
Cited alongside, same era.
https://www.anthropic.com/news/claude-3-7-sonnet
anthropic. 2025 · 2025
Cited alongside, same era.
Training language models to reason efficiently
Daman Arora and Andrea Zanette. 2025 · 2025
Cited alongside, same era.
Language models can predict their own behavior
Dhananjay Ashok and Jonathan May. 2025 · 2025
Cited alongside, same era.
Sketch-of-thought: Efficient llm reasoning with adaptive cognitive-inspired sketching
Simon A. Aytes, Jinheon Baek, and Sung Ju Hwang. 2025 · 2025
Cited alongside, same era.
Zhiting Mei, Christina Zhang, Tenny Yin, Justin Lidard, Ola Shorinwa, and Anirudha Majumdar. 2025 · 2025
Closest in time.
Minimax-m1: Scaling test-time compute efficiently with lightning attention
MiniMax, :, Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, et al. 2025 · 2025
Closest in time.
Ivan Moshkov, Darragh Hanley, Ivan Sorokin, Shubham Toshniwal, Christof Henkel, Benedikt Schifferer, Wei Du, and Igor Gitman. 2025 · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025 · 2025
Closest in time.
Self-training elicits concise reasoning in large language models
Tergel Munkhbat, Namgyu Ho, Seo Hyun Kim, Yongjin Yang, Yujin Kim, and Se-Young Yun. 2025 · 2025
Closest in time.
Concise thoughts: Impact of output length on llm reasoning and cost
Sania Nayab, Giulio Rossolini, Marco Simoni, Andrea Saracino, Giorgio Buttazzo, Nicolamaria Manes, and Fabrizio Giacomelli. 2025 · 2025
Closest in time.
Not all thoughts are generated equal: Efficient llm reasoning via multi-turn reinforcement learning
Yansong Ning, Wei Li, Jun Fang, Naiqiang Tan, and Hao Liu. 2025 · 2025
Closest in time.
Activation-informed merging of large language models
Amin Heyrani Nobari, Kaveh Alimohammadi, Ali ArjomandBigdeli, Akash Srivastava, Faez Ahmed, and Navid Azizan. 2025 · 2025
Closest in time.
Routellm: Learning to route llms with preference data
Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. 2025 · 2025
Closest in time.
Thinking slow, fast: Scaling inference compute with distilled reasoners
Daniele Paliotta, Junxiong Wang, Matteo Pagliardini, Kevin Y. Li, Aviv Bick, J. Zico Kolter, Albert Gu, François Fleuret, and Tri Dao. 2025 · 2025
Closest in time.
Inference-time computations for llm reasoning and planning: A benchmark and insights
Shubham Parashar, Blake Olson, Sambhav Khurana, Eric Li, Hongyi Ling, James Caverlee, and Shuiwang Ji. 2025 · 2025
Closest in time.
Optimizing anytime reasoning via budget relative policy optimization
Penghui Qi, Zichen Liu, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. 2025 · 2025
Closest in time.
Concise: Confidence-guided compression in step-by-step efficient reasoning
Ziqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Fandong Meng, Jie Zhou, Ju Ren, and Yaoxue Zhang. 2025 · 2025
Closest in time.
A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond
Xiaoye Qu, Yafu Li, Zhaochen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, Peng Li, Wei Wei, Jing Shao, Chaochao Lu, Yue Zhang, Xian-Sheng Hua, Bowen Zhou, and Yu Cheng. 2025 · 2025
Closest in time.
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, et al. 2025 · 2025
Closest in time.
Reasoning with latent thoughts: On the power of looped transformers
Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, and Sashank J. Reddi. 2025 · 2025
Closest in time.
Seed1.5-thinking: Advancing superb reasoning models with reinforcement learning
ByteDance Seed, :, Jiaze Chen, Tiantian Fan, Xin Liu, Lingjun Liu, et al. 2025 · 2025
Closest in time.
On reasoning strength planning in large reasoning models
Leheng Sheng, An Zhang, Zijian Wu, Weixiang Zhao, Changshuo Shen, Yi Zhang, Xiang Wang, and Tat-Seng Chua. 2025 · 2025
Closest in time.
Walk before you run! concise llm reasoning via reinforcement learning
Mingyang Song and Mao Zheng. 2025 · 2025
Closest in time.
Thinking fast and right: Balancing accuracy and reasoning length with adaptive rewards
Jinyan Su and Claire Cardie. 2025 · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, and Xia Hu. 2025 · 2025
Closest in time.
Think silently, think fast: Dynamic latent compression of llm reasoning chains
Wenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo, Jian Luan, and Ruihua Song. 2025 · 2025
Closest in time.
Concisehint: Boosting efficient reasoning via continuous concise hints during generation
Siao Tang, Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2025 · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning
Qwen Team. 2025 · 2025
Closest in time.
https://github.com/tencent-hunyuan/hunyuan-a13b
Tencent Hunyuan. 2025 · 2025
Closest in time.
Deepdistill: Enhancing llm reasoning capabilities via large-scale difficulty-graded data training
Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yiping Peng, Yunjie Ji, Han Zhao, and Xiangang Li. 2025 · 2025
Closest in time.
https://transluce.org/investigating-o3-truthfulness
transluce. 2025 · 2025
Closest in time.
Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl
Songjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian, Linjing Li, Xiangyuan Lan, and Dongbin Zhao. 2025 · 2025
Closest in time.
https://www.vectara.com/blog/deepseek-r1-hallucinates-more-than-deepseek-v3
vectara. 2025 · 2025
Closest in time.
Livebench: A challenging, contamination-limited llm benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblum. 2025 · 2025
Closest in time.
Tokenskip: Controllable chain-of-thought compression in llms
Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, and Wenjie Li. 2025 · 2025
Closest in time.
Just enough thinking: Efficient reasoning with adaptive length penalties reinforcement learning
Violet Xiang, Chase Blagden, Rafael Rafailov, Nathan Lile, Sang Truong, Chelsea Finn, and Nick Haber. 2025 · 2025
Closest in time.
Fast-slow thinking for large vision-language model reasoning
Wenyi Xiao, Leilei Gan, Weilong Dai, Wanggui He, Ziwei Huang, Haoyuan Li, Fangxun Shu, Zhelun Yu, Peng Zhang, Hao Jiang, and Fei Wu. 2025 · 2025
Closest in time.
Interleaved reasoning for large language models via reinforcement learning
Roy Xie, David Qiu, Deepak Gopinath, Dong Lin, Yanchao Sun, Chong Wang, Saloni Potdar, and Bhuwan Dhingra. 2025 · 2025
Closest in time.
Mixture of reasonings: Teach large language models to reason with adaptive strategies
Tao Xiong, Xavier Hu, Wenyan Fan, and Shengyu Zhang. 2025 · 2025
Closest in time.
Demystifying long chain-of-thought reasoning in llms
Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, and Xiang Yue. 2025 · 2025
Closest in time.
Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning
Jingyang Yi, Jiazheng Wang, and Sida Li. 2025 · 2025
Closest in time.
Xixian Yong, Xiao Zhou, Yingying Zhang, Jinlin Li, Yefeng Zheng, and Xian Wu. 2025 · 2025
Closest in time.
Efficient rl training for reasoning models via length-aware optimization
Danlong Yuan, Tian Xie, Shaohan Huang, Zhuocheng Gong, Huishuai Zhang, Chong Luo, Furu Wei, and Dongyan Zhao. 2025 · 2025
Closest in time.
Accelerating chain-of-thought reasoning: When goal-gradient importance meets dynamic skipping
Ren Zhuang, Ben Wang, and Shuifa Sun. 2025 · 2025
Closest in time.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, et al. 2025 · 2025
Closest in time.