Fetching the paper…
Reading the bibliography…
Recent advancements in reasoning with large language models (RLLMs), such as OpenAI-O1 and DeepSeek-R1, have demonstrated their impressive capabilities in complex domains like mathematics and coding.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search
Thomas Anthony, Zheng Tian, and David Barber · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
On the measure of intelligence
François Chollet · 2019
Earlier work this paper cites.
Improving policies via search in cooperative partially observable games
Adam Lerer, Hengyuan Hu, Jakob Foerster, and Noam Brown · 2020
Earlier work this paper cites.
Greedy policy search: A simple baseline for learnable test-time augmentation
Alexander Lyzhov, Yuliya Molchanova, Arsenii Ashukha, Dmitry Molchanov, and Dmitry Vetrov · 2020
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Scaling scaling laws with board games
Andy L Jones · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Qiming Bao, Alex Yuxuan Peng, Tim Hartill, Neset Tan, Zhenyun Deng, Michael Witbrock, and Jiamou Liu · 2022
Earlier work this paper cites.
Goal misgeneralization in deep reinforcement learning
Lauro Langosco Di Langosco, Jack Koch, Lee D Sharkey, Jacob Pfau, and David Krueger · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al · 2022
Earlier work this paper cites.
Can language models learn from explanations in context?
Andrew Lampinen, Ishita Dasgupta, Stephanie Chan, Kory Mathewson, Mh Tessler, Antonia Creswell, James McClelland, Jane Wang, and Felix Hill · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2022
Earlier work this paper cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike · 2022
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins · 2022
Earlier work this paper cites.
ScienceWorld: Is your agent smarter than a 5th grader?
Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté, and Prithviraj Ammanabrolu · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Chain of thought imitation with procedure cloning
Sherry Yang, Dale Schuurmans, Pieter Abbeel, and Ofir Nachum · 2022
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik R Narasimhan · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Earlier work this paper cites.
Fan Zhou, Haoyu Dong, Qian Liu, Zhoujun Cheng, Shi Han, and Dongmei Zhang · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Amc 2023
AI-MO · 2023
Earlier work this paper cites.
Learning from mistakes makes llm better reasoner
Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng, Jian-Guang Lou, and Weizhu Chen · 2023
Earlier work this paper cites.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Earlier work this paper cites.
Qiming Bao, Gael Gendron, Alex Yuxuan Peng, Wanjun Zhong, Neset Tan, Yang Chen, Michael Witbrock, and Jiamou Liu · 2023
Earlier work this paper cites.
Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen · 2023
Earlier work this paper cites.
RAFT: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, KaShun SHUM, and Tong Zhang · 2023
Earlier work this paper cites.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Earlier work this paper cites.
Complexity-based prompting for multi-step reasoning
Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot · 2023
Earlier work this paper cites.
PAL: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Earlier work this paper cites.
Self-verification improves few-shot clinical information extraction
Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley, Jianfeng Gao, and Hoifung Poon · 2023
Earlier work this paper cites.
ROSCOE: A suite of metrics for scoring step-by-step reasoning
Olga Golovneva, Moya Peng Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz · 2023
Earlier work this paper cites.
Critic: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen · 2023
Earlier work this paper cites.
Reinforced self-training (rest) for language modeling
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al · 2023
Earlier work this paper cites.
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
Michael Hanna, Ollie Liu, and Alexandre Variengien · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu · 2023
Earlier work this paper cites.
Large language models are reasoning teachers
Namgyu Ho, Laura Schmid, and Se-Young Yun · 2023
Earlier work this paper cites.
Training chain-of-thought via latent-variable inference
Matthew Douglas Hoffman, Du Phan, david dohan, Sholto Douglas, Tuan Anh Le, Aaron T Parisi, Pavel Sountsov, Charles Sutton, Sharad Vikram, and Rif A. Saurous · 2023
Earlier work this paper cites.
Not all languages are created equal in LLMs: Improving multilingual capability by cross-lingual-thought prompting
Haoyang Huang, Tianyi Tang, Dongdong Zhang, Xin Zhao, Ting Song, Yan Xia, and Furu Wei · 2023
Earlier work this paper cites.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang · 2023
Earlier work this paper cites.
Mathprompter: Mathematical reasoning using large language models
Shima Imani, Liang Du, and Harsh Shrivastava · 2023
Earlier work this paper cites.
Camels in a changing climate: Enhancing lm adaptation with tulu 2, 2023
Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi · 2023
Earlier work this paper cites.
Towards mitigating LLM hallucination via self reflection
Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung · 2023
Earlier work this paper cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Earlier work this paper cites.
LAMBADA: Backward chaining for automated reasoning in natural language
Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, and Deepak Ramachandran · 2023
Earlier work this paper cites.
Grace: Discriminator-guided chain-of-thought reasoning
Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, and Lu Wang · 2023
Earlier work this paper cites.
The CoT collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning
Seungone Kim, Se Joo, Doyoung Kim, Joel Jang, Seonghyeon Ye, Jamin Shin, and Minjoon Seo · 2023
Earlier work this paper cites.
Reflection-tuning: Data recycling improves llm instruction-tuning
Ming Li, Lichang Chen, Jiuhai Chen, Shwai He, Heng Huang, Jiuxiang Gu, and Tianyi Zhou · 2023
Earlier work this paper cites.
MoT: Memory-of-thought enables ChatGPT to self-improve
Xiaonan Li and Xipeng Qiu · 2023
Earlier work this paper cites.
Making language models better reasoners with step-aware verifier
Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen · 2023
Earlier work this paper cites.
Deductive verification of chain-of-thought reasoning
Zhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang, Mingu Lee, Roland Memisevic, and Hao Su · 2023
Earlier work this paper cites.
Tinygsm: achieving> 80% on gsm8k with small language models
Bingbin Liu, Sebastien Bubeck, Ronen Eldan, Janardhan Kulkarni, Yuanzhi Li, Anh Nguyen, Rachel Ward, and Yi Zhang · 2023
Earlier work this paper cites.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang · 2023
Earlier work this paper cites.
Faithful chain-of-thought reasoning
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch · 2023
Earlier work this paper cites.
Let’s reward step by step: Step-level reward model as the navigators for reasoning
Qianli Ma, Haotian Zhou, Tingkai Liu, Jianbo Yuan, Pengfei Liu, Yang You, and Hongxia Yang · 2023
Earlier work this paper cites.
What makes chain-of-thought prompting effective? a counterfactual study
Aman Madaan, Katherine Hermann, and Amir Yazdanbakhsh · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2023
Earlier work this paper cites.
The expressive power of transformers with chain of thought
William Merrill and Ashish Sabharwal · 2023
Earlier work this paper cites.
LEVER: Learning to verify language-to-code generation with execution
Ansong Ni, Srini Iyer, Dragomir Radev, Veselin Stoyanov, Wen-Tau Yih, Sida Wang, and Xi Victoria Lin · 2023
Earlier work this paper cites.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang · 2023
Earlier work this paper cites.
ReCEval: Evaluating reasoning chains via correctness and informativeness
Archiki Prasad, Swarnadeep Saha, Xiang Zhou, and Mohit Bansal · 2023
Earlier work this paper cites.
Why think step by step? reasoning emerges from the locality of experience
Ben Prystawski, Michael Li, and Noah Goodman · 2023
Earlier work this paper cites.
Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages
Libo Qin, Qiguang Chen, Fuxuan Wei, Shijue Huang, and Wanxiang Che · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Earlier work this paper cites.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al · 2023
Earlier work this paper cites.
Synthetic prompting: Generating chain-of-thought demonstrations for large language models
Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen · 2023
Earlier work this paper cites.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Earlier work this paper cites.
Automatic prompt augmentation and selection with chain-of-thought from labeled data
Kashun Shum, Shizhe Diao, and Tong Zhang · 2023
Earlier work this paper cites.
A survey of reasoning with foundation models
Jiankai Sun, Chuanyang Zheng, Enze Xie, Zhengying Liu, Ruihang Chu, Jianing Qiu, Jiaqi Xu, Mingyu Ding, Hongyang Li, Mengzhe Geng, et al · 2023
Earlier work this paper cites.
Challenging BIG-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, and Jason Wei · 2023
Earlier work this paper cites.
Causal abstraction for chain-of-thought reasoning in arithmetic word problems
Juanhe (TJ) Tan · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Earlier work this paper cites.
Towards understanding chain-of-thought prompting: An empirical study of what matters
Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun · 2023
Earlier work this paper cites.
Large language models are better reasoners with self-verification
Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao · 2023
Earlier work this paper cites.
System 2 attention (is something you might need too)
Jason Weston and Sainbayar Sukhbaatar · 2023
Earlier work this paper cites.
Self-evaluation guided beam search for reasoning
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie · 2023
Earlier work this paper cites.
Haotian Xu · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Earlier work this paper cites.
Mammoth: Building math generalist models through hybrid instruction tuning
Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen · 2023
Earlier work this paper cites.
A multi-modal neural geometric solver with textual clauses parsed from diagram
Ming-Liang Zhang, Fei yin, and Cheng-Lin Liu · 2023
Earlier work this paper cites.
Large language models as commonsense knowledge for large-scale task planning
Zirui Zhao, Wee Sun Lee, and David Hsu · 2023
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi · 2023
Earlier work this paper cites.
Solving math word problems via cooperative reasoning induced language models
Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Ruyi Gan, Jiaxing Zhang, and Yujiu Yang · 2023
Earlier work this paper cites.
Through the lens of core competency: Survey on evaluation of large language models
Ziyu Zhuang, Qiguang Chen, Longxuan Ma, Mingda Li, Yi Han, Yushan Qian, Haopeng Bai, Weinan Zhang, and Liu Ting · 2023
Earlier work this paper cites.
Medec: A benchmark for medical error detection and correction in clinical notes
Asma Ben Abacha, Wen-wai Yim, Yujuan Fu, Zhaoyi Sun, Meliha Yetisgen, Fei Xia, and Thomas Lin · 2024
Earlier work this paper cites.
Inference scaling vs reasoning: An empirical analysis of compute-optimal llm problem-solving
Marwan AbdElhameed and Pavly Halim · 2024
Earlier work this paper cites.
Nemotron-4 340b technical report
Bo Adler, Niket Agarwal, Ashwath Aithal, Dong H Anh, Pallab Bhattacharya, Annika Brundyn, Jared Casper, Bryan Catanzaro, Sharon Clay, Jonathan Cohen, et al · 2024
Earlier work this paper cites.
Aime 2024
AI-MO · 2024
Earlier work this paper cites.
Critique-out-loud reward models
Zachary Ankner, Mansheej Paul, Brandon Cui, Jonathan Daniel Chang, and Prithviraj Ammanabrolu · 2024
Earlier work this paper cites.
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic · 2024
Earlier work this paper cites.
Llemma: An open language model for mathematics
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen Marcus McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck · 2024
Earlier work this paper cites.
Exploring iterative enhancement for improving learnersourced multiple-choice question explanations with large language models
Qiming Bao, Juho Leinonen, Alex Yuxuan Peng, Wanjun Zhong, Gael Gendron, Timothy Pistotti, Alice Huang, Paul Denny, Michael Witbrock, and Jiamou Liu · 2024
Earlier work this paper cites.
Abstract Meaning Representation-based logic-driven data augmentation for logical reasoning
Qiming Bao, Alex Peng, Zhenyun Deng, Wanjun Zhong, Gael Gendron, Timothy Pistotti, Neset Tan, Nathan Young, Yang Chen, Yonghua Zhu, Paul Denny, Michael Witbrock, and Jiamou Liu · 2024
Earlier work this paper cites.
Titans: Learning to memorize at test time
Ali Behrouz, Peilin Zhong, and Vahab Mirrokni · 2024
Earlier work this paper cites.
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler · 2024
Earlier work this paper cites.
Deepseek llm: Scaling open-source language models with longtermism
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al · 2024
Earlier work this paper cites.
Vermcts: Synthesizing multi-step programs using a verifier, a large language model, and tree search
David Brandfonbrener, Simon Henniger, Sibi Raja, Tarun Prasad, Chloe Loughridge, Federico Cassano, Sabrina Ruixin Hu, Jianang Yang, William E Byrd, Robert Zinkov, et al · 2024
Earlier work this paper cites.
Large language monkeys: Scaling inference compute with repeated sampling
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini · 2024
Earlier work this paper cites.
ARES: Alternating reinforcement learning and supervised fine-tuning for enhanced multi-modal chain-of-thought reasoning through diverse AI feedback
Ju-Seung Byun, Jiyun Chun, Jihyung Kil, and Andrew Perrault · 2024
Earlier work this paper cites.
System-2 mathematical reasoning via enriched instruction tuning
Huanqia Cai, Yijun Yang, and Zhifeng Li · 2024
Earlier work this paper cites.
Xai meets llms: A survey of the relation between explainable ai and large language models
Erik Cambria, Lorenzo Malandri, Fabio Mercorio, Navid Nobani, and Andrea Seveso · 2024
Earlier work this paper cites.
GraphReason: Enhancing reasoning capabilities of large language models through a graph-based verification approach
Lang Cao · 2024
Earlier work this paper cites.
xcot: Cross-lingual instruction tuning for cross-lingual chain-of-thought reasoning
Linzheng Chai, Jian Yang, Tao Sun, Hongcheng Guo, Jiaheng Liu, Bing Wang, Xiannian Liang, Jiaqi Bai, Tongliang Li, Qiyao Peng, et al · 2024
Earlier work this paper cites.
Mle-bench: Evaluating machine learning agents on machine learning engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, et al · 2024
Earlier work this paper cites.
Step-level value preference optimization for mathematical reasoning
Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan · 2024
Earlier work this paper cites.
M 3 CoT: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought
Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen, Xiao Xu, and Wanxiang Che · 2024
Earlier work this paper cites.
Toward adaptive reasoning in large language models with thought rollback
Sijia Chen and Baochun Li · 2024
Earlier work this paper cites.
When is tree search useful for LLM planning? it depends on the discriminator
Ziru Chen, Michael White, Ray Mooney, Ali Payani, Yu Su, and Huan Sun · 2024
Earlier work this paper cites.
Jiale Cheng, Xiao Liu, Cunxiang Wang, Xiaotao Gu, Yida Lu, Dan Zhang, Yuxiao Dong, Jie Tang, Hongning Wang, and Minlie Huang · 2024
Earlier work this paper cites.
PuzzleVQA: Diagnosing multimodal reasoning challenges of language models with abstract visual patterns
Yew Ken Chia, Vernon Toh, Deepanway Ghosal, Lidong Bing, and Soujanya Poria · 2024
Earlier work this paper cites.
Agents thinking fast and slow: A talker-reasoner architecture
Konstantina Christakopoulou, Shibl Mourad, and Maja Mataric · 2024
Earlier work this paper cites.
Navigate through enigmatic labyrinth a survey of chain of thought reasoning: Advances, frontiers and future
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, and Ting Liu · 2024
Earlier work this paper cites.
Beyond llms: Advancing the landscape of complex reasoning
Jennifer Chu-Carroll, Andrew Beck, Greg Burnham, David OS Melville, David Nachman, A Erdem Özcan, and David Ferrucci · 2024
Earlier work this paper cites.
Mhpp: Exploring the capabilities and limitations of language models beyond basic code generation
Jianbo Dai, Jianqiao Lu, Yunlong Feng, Dong Huang, Guangtao Zeng, Rongju Ruan, Ming Cheng, Haochen Tan, and Zhijiang Guo · 2024
Earlier work this paper cites.
From explicit cot to implicit cot: Learning to internalize cot step by step
Yuntian Deng, Yejin Choi, and Stuart Shieber · 2024
Earlier work this paper cites.
Rlhf workflow: From reward modeling to online rlhf
Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang · 2024
Earlier work this paper cites.
Stepcoder: Improve code generation with reinforcement learning from compiler feedback
Shihan Dou, Yan Liu, Haoxiang Jia, Limao Xiong, Enyu Zhou, Wei Shen, Junjie Shan, Caishuang Huang, Xiao Wang, Xiaoran Fan, et al · 2024
Earlier work this paper cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Earlier work this paper cites.
How to think step-by-step: A mechanistic understanding of chain-of-thought reasoning
Subhabrata Dutta, Joykirat Singh, Soumen Chakrabarti, and Tanmoy Chakraborty · 2024
Earlier work this paper cites.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Earlier work this paper cites.
Promptbreeder: Self-referential self-improvement via prompt evolution
Chrisantha Fernando, Dylan Sunil Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel · 2024
Earlier work this paper cites.
LLM self-correction with deCRIM: Decompose, critique, and refine for enhanced following of instructions with multiple constraints
Thomas Palmeira Ferraz, Kartik Mehta, Yu-Hsiang Lin, Haw-Shiuan Chang, Shereen Oraby, Sijia Liu, Vivek Subramanian, Tagyoung Chung, Mohit Bansal, and Nanyun Peng · 2024
Earlier work this paper cites.
Stream of search (sos): Learning to search in language
Kanishk Gandhi, Denise HJ Lee, Gabriel Grand, Muxin Liu, Winson Cheng, Archit Sharma, and Noah Goodman · 2024
Earlier work this paper cites.
Bofei Gao, Zefan Cai, Runxin Xu, Peiyi Wang, Ce Zheng, Runji Lin, Keming Lu, Junyang Lin, Chang Zhou, Tianyu Liu, and Baobao Chang · 2024
Earlier work this paper cites.
Rlef: Grounding code llms in execution feedback with reinforcement learning
Jonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella, Taco Cohen, and Gabriel Synnaeve · 2024
Earlier work this paper cites.
Large language models are not strong abstract reasoners
Gaël Gendron, Qiming Bao, Michael Witbrock, and Gillian Dobbie · 2024
Earlier work this paper cites.
Puzzle solving using reasoning of large language models: A survey
Panagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, and Giorgos Stamou · 2024
Earlier work this paper cites.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai
Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, et al · 2024
Earlier work this paper cites.
Chatglm: A family of large language models from glm-130b to glm-4 all tools
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, et al · 2024
Earlier work this paper cites.
Uncertainty-guided optimization on large language model search trees
Julia Grosse, Ruotian Wu, Ahmad Rashid, Philipp Hennig, Pascal Poupart, and Agustinus Kristiadi · 2024
Earlier work this paper cites.
Xinyan Guan, Yanjiang Liu, Xinyu Lu, Boxi Cao, Ben He, Xianpei Han, Le Sun, Jie Lou, Bowen Yu, Yaojie Lu, et al · 2024
Earlier work this paper cites.
Putnam-AXIOM: A functional and static benchmark for measuring higher level mathematical reasoning
Aryan Gulati, Brando Miranda, Eric Chen, Emily Xia, Kai Fronsdal, Bruno de Moraes Dumont, and Sanmi Koyejo · 2024
Earlier work this paper cites.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al · 2024
Earlier work this paper cites.
Token-budget-aware llm reasoning
Tingxu Han, Chunrong Fang, Shiyu Zhao, Shiqing Ma, Zhenyu Chen, and Zhenting Wang · 2024
Earlier work this paper cites.
LLM reasoners: New evaluation, library, and analysis of step-by-step reasoning with large language models
Shibo Hao, Yi Gu, Haotian Luo, Tianyang Liu, Xiyan Shao, Xinyuan Wang, Shuhua Xie, Haodi Ma, Adithya Samavedhi, Qiyue Gao, Zhen Wang, and Zhiting Hu · 2024
Earlier work this paper cites.
GLore: When, where, and how to improve LLM reasoning via global and local refinements
Alexander Havrilla, Sharath Chandra Raparthy, Christoforos Nalmpantis, Jane Dwivedi-Yu, Maksym Zhuravinskyi, Eric Hambro, and Roberta Raileanu · 2024
Earlier work this paper cites.
OlympiadBench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun · 2024
Earlier work this paper cites.
Advancing process verification for large language models via tree-based preference learning
Mingqian He, Yongliang Shen, Wenqi Zhang, Zeqi Tan, and Weiming Lu · 2024
Earlier work this paper cites.
Advances in reasoning by prompting large language models: A survey
Ruixin Hong, Xinyu Pang, and Changshui Zhang · 2024
Earlier work this paper cites.
Not all LLM reasoners are created equal
Arian Hosseini, Alessandro Sordoni, Daniel Kenji Toyama, Aaron Courville, and Rishabh Agarwal · 2024
Earlier work this paper cites.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Jian Hu, Xibin Wu, Zilin Zhu, Xianyu, Weixun Wang, Dehao Zhang, and Yu Cao · 2024
Earlier work this paper cites.
Uncertainty of thoughts: Uncertainty-aware planning enhances information seeking in large language models
Zhiyuan Hu, Chumin Liu, Xidong Feng, Yilun Zhao, See-Kiong Ng, Anh Tuan Luu, Junxian He, Pang Wei Koh, and Bryan Hooi · 2024
Earlier work this paper cites.
Jen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang, Wenxuan Wang, Youliang Yuan, Wenxiang Jiao, Xing Wang, Zhaopeng Tu, and Michael R Lyu · 2024
Earlier work this paper cites.
A survey on evaluation of multimodal large language models
Jiaxing Huang and Jingyi Zhang · 2024
Earlier work this paper cites.
Advancing large language model attribution through self-improving
Lei Huang, Xiaocheng Feng, Weitao Ma, Liang Zhao, Yuchun Fan, Weihong Zhong, Dongliang Xu, Qing Yang, Hongtao Liu, and Bing Qin · 2024
Earlier work this paper cites.
Qwen2.5-coder technical report
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Keming Lu, et al · 2024
Earlier work this paper cites.
Self-explore: Enhancing mathematical reasoning in language models with fine-grained rewards
Hyeonbin Hwang, Doyoung Kim, Seungone Kim, Seonghyeon Ye, and Minjoon Seo · 2024
Earlier work this paper cites.
Mapcoder: Multi-agent code generation for competitive problem solving
Md Ashraful Islam, Mohammed Eunus Ali, and Md Rizwan Parvez · 2024
Earlier work this paper cites.
Unpacking DPO and PPO: Disentangling best practices for learning from preference feedback
Hamish Ivison, Yizhong Wang, Jiacheng Liu, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert, Noah A. Smith, Yejin Choi, and Hannaneh Hajishirzi · 2024
Earlier work this paper cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Earlier work this paper cites.
A small step towards reproducing openai o1: Progress report on the steiner open source models, October 2024
Yichao Ji · 2024
Earlier work this paper cites.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Earlier work this paper cites.
SWE-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan · 2024
Earlier work this paper cites.
Disentangling memory and reasoning ability in large language models
Mingyu Jin, Weidi Luo, Sitao Cheng, Xinyi Wang, Wenyue Hua, Ruixiang Tang, William Yang Wang, and Yongfeng Zhang · 2024
Earlier work this paper cites.
The impact of reasoning step length on large language models
Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, and Mengnan Du · 2024
Earlier work this paper cites.
Gpt-guided monte carlo tree search for symbolic regression in financial fraud detection
Prashank Kadam · 2024
Earlier work this paper cites.
Evaluating LLMs at detecting errors in LLM responses
Ryo Kamoi, Sarkar Snigdha Sarathi Das, Renze Lou, Jihyun Janice Ahn, Yilun Zhao, Xiaoxin Lu, Nan Zhang, Yusen Zhang, Haoran Ranran Zhang, Sujeeth Reddy Vummanthala, Salika Dave, Shaobo Qin, Arman Cohan, Wenpeng Yin, and Rui Zhang · 2024
Earlier work this paper cites.
Mindstar: Enhancing math reasoning in pre-trained llms at inference time
Jikun Kang, Xin Zhe Li, Xi Chen, Amirreza Kazemi, Qianyi Sun, Boxing Chen, Dong Li, Xu He, Quan He, Feng Wen, et al · 2024
Earlier work this paper cites.
On the empirical complexity of reasoning and planning in LLMs
Liwei Kang, Zirui Zhao, David Hsu, and Wee Sun Lee · 2024
Earlier work this paper cites.
Vineppo: Unlocking rl potential for llm reasoning through refined credit assignment
Amirhossein Kazemnejad, Milad Aghajohari, Eva Portelance, Alessandro Sordoni, Siva Reddy, Aaron Courville, and Nicolas Le Roux · 2024
Cited alongside, same era.
Lachesis: Predicting llm inference accuracy using structural properties of reasoning paths
Naryeong Kim, Sungmin Kang, Gabin An, and Shin Yoo · 2024
Cited alongside, same era.
Prometheus 2: An open source language model specialized in evaluating other language models
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo · 2024
Cited alongside, same era.
Tree search for language model agents
Jing Yu Koh, Stephen McAleer, Daniel Fried, and Ruslan Salakhutdinov · 2024
Cited alongside, same era.
Training language models to self-correct via reinforcement learning
Dynamic parallel tree search for efficient llm reasoning
Yifu Ding, Wentao Jiang, Shunyu Liu, Yongcheng Jing, Jinyang Guo, Yingjie Wang, Jing Zhang, Zengmao Wang, Ziwei Liu, Bo Du, et al · 2025
Closest in time.
Stp: Self-play llm theorem provers with iterative conjecturing and proving
Kefan Dong and Tengyu Ma · 2025
Closest in time.
Emergent response planning in llm
Zhichen Dong, Zhanhui Zhou, Zhixuan Liu, Chao Yang, and Chaochao Lu · 2025
Closest in time.
Diverse inference and verification for advanced reasoning
Iddo Drori, Gaston Longhitano, Mao Mao, Seunghwan Hyun, Yuke Zhang, Sungjun Park, Zachary Meeks, Xin-Yu Zhang, Ben Segev, Howard Yong, et al · 2025
Closest in time.
Boost, disentangle, and customize: A robust system2-to-system1 pipeline for code generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, et al · 2024
Cited alongside, same era.
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms
Xin Lai, Zhuotao Tian, Yukang Chen, Senqiao Yang, Xiangru Peng, and Jiaya Jia · 2024
Cited alongside, same era.
Tulu 3: Pushing frontiers in open language model post-training, 2024
Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luca Soldaini, Noah A. Smith, Yizhong Wang, Pradeep Dasigi, and Hannaneh Hajishirzi · 2024
Cited alongside, same era.
Comat: Chain of mathematically annotated thought improves mathematical reasoning
Joshua Ong Jun Leang, Aryo Pradipta Gema, and Shay B Cohen · 2024
Cited alongside, same era.
Jung Hyun Lee, June Yong Yang, Byeongho Heo, Dongyoon Han, and Kang Min Yoo · 2024
Cited alongside, same era.
Beyond a*: Better planning with transformers via search dynamics bootstrapping
Lucas Lehnert, Sainbayar Sukhbaatar, DiJia Su, Qinqing Zheng, Paul McVay, Michael Rabbat, and Yuandong Tian · 2024
Cited alongside, same era.
MACM: Utilizing a multi-agent system for condition mining in solving complex mathematical problems
Bin Lei, Yi Zhang, Shan Zuo, Ali Payani, and Caiwen Ding · 2024
Cited alongside, same era.
Bohan Li, Jiannan Guan, Longxu Dou, Yunlong Feng, Dingzirui Wang, Yang Xu, Enbo Wang, Qiguang Chen, Bichen Wang, Xiao Xu, et al · 2024
Cited alongside, same era.
Kounianhua Du, Hanjing Wang, Jianxing Liu, Jizheng Chen, Xinyi Dai, Yasheng Wang, Ruiming Tang, Yong Yu, Jun Wang, and Weinan Zhang · 2025
Closest in time.
Competitive programming with large reasoning models
Ahmed El-Kishky, Alexander Wei, Andre Saraiva, Borys Minaev, Daniel Selsam, David Dohan, Francis Song, Hunter Lightman, Ignasi Clavera, Jakub Pachocki, et al · 2025
Closest in time.
Large language models for recommendation with deliberative user preference alignment
Yi Fang, Wenjie Wang, Yang Zhang, Fengbin Zhu, Qifan Wang, Fuli Feng, and Xiangnan He · 2025
Closest in time.
Reasoning does not necessarily improve role-playing ability
Xiachong Feng, Longxu Dou, and Lingpeng Kong · 2025
Closest in time.
Reasoning beyond limits: Advances and open problems for llms
Mohamed Amine Ferrag, Norbert Tihanyi, and Merouane Debbah · 2025
Closest in time.
Unveiling and causalizing cot: A causal pespective
Jiarun Fu, Lizhong Ding, Hao Li, Pengqi Li, Qiuning Wei, and Xu Chen · 2025
Closest in time.
Metasc: Test-time safety specification optimization for language models
Víctor Gallego · 2025
Closest in time.
Rethinking external slow-thinking: From snowball errors to probability of correct reasoning
Zeyu Gan, Yun Liao, and Yong Liu · 2025
Closest in time.
Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective stars
Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh, Nathan Lile, and Noah D Goodman · 2025
Closest in time.
A comparison of deepseek and other llms
Tianchen Gao, Jiashun Jin, Zheng Tracy Ke, and Gabriel Moryoussef · 2025
Closest in time.
Yuyao Ge, Shenghua Liu, Yiwei Wang, Lingrui Mei, Lizhe Chen, Baolong Bi, and Xueqi Cheng · 2025
Closest in time.
Scaling up test-time compute with latent reasoning: A recurrent depth approach
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian R Bartoldson, Bhavya Kailkhura, Abhinav Bhatele, and Tom Goldstein · 2025
Closest in time.
The multilingual mind: A survey of multilingual reasoning in language models
Akash Ghosh, Debayan Datta, Sriparna Saha, and Chirag Agarwal · 2025
Closest in time.
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al · 2025
Closest in time.
Capturing nuanced preferences: Preference-aligned distillation for small language models
Yanggan Gu, Junzhuo Li, Sirui Huang, Xin Zou, Zhenghua Li, and Xuming Hu · 2025
Closest in time.
Deeprag: Thinking to retrieval step by step for large language models
Xinyan Guan, Jiali Zeng, Fandong Meng, Chunlei Xin, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun, and Jie Zhou · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Can mllms reason in multimodality? emma: An enhanced multimodal reasoning benchmark
Yunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang, and Yu Cheng · 2025
Closest in time.
From code to courtroom: Llms as the new software judges
Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, and David Lo · 2025
Closest in time.
Inference retrieval-augmented multi-modal chain-of-thoughts reasoning for language models
Qiangqiang He, Shuwei Qian, Jie Zhang, and Chongjun Wang · 2025
Closest in time.
Evaluating the systematic reasoning abilities of large language models through graph coloring
Alex Heyman and Joel Zylberberg · 2025
Closest in time.
Thinkprune: Pruning long chain-of-thought of llms via reinforcement learning
Bairu Hou, Yang Zhang, Jiabao Ji, Yujian Liu, Kaizhi Qian, Jacob Andreas, and Shiyu Chang · 2025
Closest in time.
Reinforce++: A simple and efficient approach for aligning large language models
Jian Hu · 2025
Closest in time.
Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model
Jingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang, and Heung-Yeung Shum Xiangyu Zhang · 2025
Closest in time.
Lean and mean: Decoupled value policy optimization with global value guidance
Chenghua Huang, Lu Wang, Fangkai Yang, Pu Zhao, Zhixu Li, Qingwei Lin, Dongmei Zhang, Saravan Rajmohan, and Qi Zhang · 2025
Closest in time.
Test-time view selection for multi-modal decision making
Eeshaan Jain, Johann Wenckstern, Benedikt von Querfurth, and Charlotte Bunne · 2025
Closest in time.
Ke Ji, Jiahao Xu, Tian Liang, Qiuzhi Liu, Zhiwei He, Xingyu Chen, Xiaoyuan Liu, Zhijie Wang, Junying Chen, Benyou Wang, et al · 2025
Closest in time.
Boyu Jia, Junzhe Zhang, Huixuan Zhang, and Xiaojun Wan · 2025
Closest in time.
Safechain: Safety of language models with long chain-of-thought reasoning capabilities
Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran · 2025
Closest in time.
Exploring concept depth: How large language models acquire knowledge and concept at different layers?
Mingyu Jin, Qinkai Yu, Jingyuan Huang, Qingcheng Zeng, Zhenting Wang, Wenyue Hua, Haiyan Zhao, Kai Mei, Yanda Meng, Kaize Ding, Fan Yang, Mengnan Du, and Yongfeng Zhang · 2025
Closest in time.
Large language models pass the turing test
Cameron R Jones and Benjamin K Bergen · 2025
Closest in time.
Scalable best-of-n selection for large language models via self-certainty
Zhewei Kang, Xuandong Zhao, and Dawn Song · 2025
Closest in time.
Towards robust legal reasoning: Harnessing logical llms in law
Manuj Kant, Sareh Nabi, Manav Kant, Roland Scharrer, Megan Ma, and Marzieh Nabi · 2025
Closest in time.
Artyom Kharinaev, Viktor Moskvoretskii, Egor Shvetsov, Kseniia Studenikina, Bykov Mikhail, and Evgeny Burnaev · 2025
Closest in time.
Hypothesis-driven theory-of-mind reasoning for large language models
Hyunwoo Kim, Melanie Sclar, Tan Zhi-Xuan, Lance Ying, Sydney Levine, Yang Liu, Joshua B Tenenbaum, and Yejin Choi · 2025
Closest in time.
Scalable language models with posterior inference of latent thought vectors
Deqian Kong, Minglu Zhao, Dehong Xu, Bo Pang, Shu Wang, Edouardo Honig, Zhangzhang Si, Chuan Li, Jianwen Xie, Sirui Xie, et al · 2025
Closest in time.
Overthink: Slowdown attacks on reasoning llms
Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, and Eugene Bagdasarian · 2025
Closest in time.
Martin Kuo, Jianyi Zhang, Aolin Ding, Qinsi Wang, Louis DiValentin, Yujia Bao, Wei Wei, Da-Cheng Juan, Hai Li, and Yiran Chen · 2025
Closest in time.
Open-r1-multimodal
EvolvingLMMs Lab · 2025
Closest in time.
Bespoke-stratos: The unreasonable effectiveness of reasoning distillation
Bespoke Labs · 2025
Closest in time.
Multidimensional consistency improves reasoning in language models
Huiyuan Lai, Xiao Zhang, and Malvina Nissim · 2025
Closest in time.
Diverse preference optimization
Jack Lanchantin, Angelica Chen, Shehzaad Dhuliawala, Ping Yu, Jason Weston, Sainbayar Sukhbaatar, and Ilia Kulikov · 2025
Closest in time.
Anh Duc Le, Tu Vu, Nam Le Hai, Nguyen Thi Ngoc Diep, Linh Ngo Van, Trung Le, and Thien Huu Nguyen · 2025
Closest in time.
Theorem prover as a judge for synthetic data generation
Joshua Ong Jun Leang, Giwon Hong, Wenda Li, and Shay B Cohen · 2025
Closest in time.
Revise: Learning to refine at test-time via intrinsic self-verification
Hyunseok Lee, Seunghyuk Oh, Jaehyung Kim, Jinwoo Shin, and Jihoon Tack · 2025
Closest in time.
Evaluating step-by-step reasoning traces: A survey
Jinu Lee and Julia Hockenmaier · 2025
Closest in time.
Questbench: Can llms ask the right question to acquire information in reasoning tasks?
Belinda Z Li, Been Kim, and Zi Wang · 2025
Closest in time.
Enhancing reasoning through process supervision with monte carlo tree search
Shuangtao Li, Shuaihao Dong, Kexin Luan, Xinhan Di, and Chaofan Ding · 2025
Closest in time.
Xinzhe Li · 2025
Closest in time.
CMMaTH: A Chinese multi-modal math skill evaluation benchmark for foundation models
Zhongzhi Li, Ming-Liang Zhang, Pei-Jie Wang, Jian Xu, Rui-Song Zhang, Yin Fei, Zhi-Long Ji, Jin-Feng Bai, Zhen-Ru Pan, Jiaxin Zhang, and Cheng-Lin Liu · 2025
Closest in time.
Reward-guided speculative decoding for efficient llm reasoning
Baohao Liao, Yuhui Xu, Hanze Dong, Junnan Li, Christof Monz, Silvio Savarese, Doyen Sahoo, and Caiming Xiong · 2025
Closest in time.
Zebralogic: On the scaling limits of llms for logical reasoning
Bill Yuchen Lin, Ronan Le Bras, Kyle Richardson, Ashish Sabharwal, Radha Poovendran, Peter Clark, and Yejin Choi · 2025
Closest in time.
Style over substance: Distilled language models reason via stylistic replication
Philip Lippmann and Jie Yang · 2025
Closest in time.
Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, et al · 2025
Closest in time.
Dakuan Lu, Xiaoyu Tan, Rui Xu, Tianchu Yao, Chao Qu, Wei Chu, Yinghui Xu, and Yuan Qi · 2025
Closest in time.
O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning
Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, and Dacheng Tao · 2025
Closest in time.
Exploring the limit of outcome reward for learning mathematical reasoning
Chengqi Lyu, Songyang Gao, Yuzhe Gu, Wenwei Zhang, Jianfei Gao, Kuikun Liu, Ziyi Wang, Shuaibin Li, Qian Zhao, Haian Huang, et al · 2025
Closest in time.
Inference-time scaling for diffusion models beyond scaling denoising steps
Nanye Ma, Shangyuan Tong, Haolin Jia, Hexiang Hu, Yu-Chuan Su, Mingda Zhang, Xuan Yang, Yandong Li, Tommi Jaakkola, Xuhui Jia, et al · 2025
Closest in time.
Millions scale dataset distilled from r1-32b
Sathwik Tejaswi Madhusudhan, Shruthan Radhakrishna, Jash Mehta, and Toby Liang · 2025
Closest in time.
Sadegh Mahdavi, Muchen Li, Kaiwen Liu, Christos Thrampoulidis, Leonid Sigal, and Renjie Liao · 2025
Closest in time.
Cos (m+ o) s: Curiosity and rl-enhanced mcts for exploring story space via language models
Tobias Materzok · 2025
Closest in time.
Synthetic-1: Two million collaboratively generated reasoning traces from deepseek-r1, 2025
Justus Mattern, Sami Jaghouar, Manveer Basra, Jannik Straube, Matthew Di Ferrante, Felix Gabriel, Jack Min Ong, Vincent Weisser, and Johannes Hagemann · 2025
Closest in time.
Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning
Fanqing Meng, Lingxiao Du, Zongkai Liu, Zhixiang Zhou, Quanfeng Lu, Daocheng Fu, Botian Shi, Wenhai Wang, Junjun He, Kaipeng Zhang, Ping Luo, Yu Qiao, Qiaosheng Zhang, and Wenqi Shao · 2025
Closest in time.
GSM-symbolic: Understanding the limitations of mathematical reasoning in large language models
Seyed Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.
Self-training elicits concise reasoning in large language models
Tergel Munkhbat, Namgyu Ho, Seohyun Kim, Yongjin Yang, Yujin Kim, and Se-Young Yun · 2025
Closest in time.
Toolcomp: A multi-tool reasoning & process supervision benchmark
Vaskar Nath, Pranav Raja, Claire Yoon, and Sean Hendryx · 2025
Closest in time.
Aime 2025
OpenCompass · 2025
Closest in time.
How do llms acquire new knowledge? a knowledge circuits perspective on continual pre-training
Yixin Ou, Yunzhi Yao, Ningyu Zhang, Hui Jin, Jiacheng Sun, Shumin Deng, Zhenguo Li, and Huajun Chen · 2025
Closest in time.
Coat: Chain-of-associated-thoughts framework for enhancing large language models reasoning
Jianfeng Pan, Senyou Deng, and Shaomang Huang · 2025
Closest in time.
Tinyzero
Jiayi Pan, Junjie Zhang, Xingyao Wang, Lifan Yuan, Hao Peng, and Alane Suhr · 2025
Closest in time.
Bolt: Bootstrap long chain-of-thought in language models without distillation
Bo Pang, Hanze Dong, Jiacheng Xu, Silvio Savarese, Yingbo Zhou, and Caiming Xiong · 2025
Closest in time.
Inference-time computations for llm reasoning and planning: A benchmark and insights
Shubham Parashar, Blake Olson, Sambhav Khurana, Eric Li, Hongyi Ling, James Caverlee, and Shuiwang Ji · 2025
Closest in time.
Chanwoo Park, Seungju Han, Xingzhi Guo, Asuman Ozdaglar, Kaiqing Zhang, and Joo-Kyung Kim · 2025
Closest in time.
Manojkumar Parmar and Yuvaraj Govindarajulu · 2025
Closest in time.
Advancing reasoning in large language models: Promising methods and approaches
Avinash Patil · 2025
Closest in time.
Avinash Patil and Amardeep Kour Gedhu · 2025
Closest in time.
Patomporn Payoungkhamdee, Pume Tuchinda, Jinheon Baek, Samuel Cahyawijaya, Can Udomcharoenchaikit, Potsawee Manakul, Peerat Limkonchotiwat, Ekapol Chuangsuwanich, and Sarana Nutanong · 2025
Closest in time.
Dengyun Peng, Yuhang Zhou, Qiguang Chen, Jinhao Liu, Jingjing Chen, and Libo Qin · 2025
Closest in time.
Proof or bluff? evaluating llms on 2025 usa math olympiad
Ivo Petrov, Jasper Dekoninck, Lyuben Baltadzhiev, Maria Drencheva, Kristian Minchev, Mislav Balunović, Nikola Jovanović, and Martin Vechev · 2025
Closest in time.
Understanding and benchmarking artificial intelligence: Openai’s o3 is not agi
Rolf Pfister and Hansueli Jud · 2025
Closest in time.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, et al · 2025
Closest in time.
Hardml: A benchmark for evaluating data science and machine learning knowledge and reasoning in ai
Tidor-Vlad Pricope · 2025
Closest in time.
A roadmap to guide the integration of llms in hierarchical planning
Israel Puerta-Merino, Carlos Núñez-Molina, Pablo Mesejo, and Juan Fernández-Olivares · 2025
Closest in time.
Isha Puri, Shivchander Sudalairaj, Guangxuan Xu, Kai Xu, and Akash Srivastava · 2025
Closest in time.
A survey of multilingual large language models
Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, and S Yu Philip · 2025
Closest in time.
Optimizing test-time compute via meta reinforcement finetuning
Yuxiao Qu, Matthew Y. R. Yang, Amrith Setlur, Lewis Tunstall, Edward Emanuel Beeching, Ruslan Salakhutdinov, and Aviral Kumar · 2025
Closest in time.
On the reasoning capacity of ai models and how to quantify it
Santosh Kumar Radha and Oktay Goktas · 2025
Closest in time.
Improving chain-of-thought reasoning via quasi-symbolic abstractions
Leonardo Ranaldi, Marco Valentino, Alexander Polonsky, and Andrè Freitas · 2025
Closest in time.
Mohammad Raza and Natasa Milic-Frayling · 2025
Closest in time.
Cer: Confidence enhanced reasoning in llms
Ali Razghandi, Seyed Mohammad Hadi Hosseini, and Mahdieh Soleymani Baghshah · 2025
Closest in time.
Reasoning to learn from latent thoughts
Yangjun Ruan, Neil Band, Chris J Maddison, and Tatsunori Hashimoto · 2025
Closest in time.
Think or step-by-step? unzipping the black box in zero-shot prompts
Nikta Gohari Sadr, Sangmitra Madhusudan, and Ali Emami · 2025
Closest in time.
Learning to plan & reason for evaluation with thinking-llm-as-a-judge
Swarnadeep Saha, Xian Li, Marjan Ghazvininejad, Jason Weston, and Tianlu Wang · 2025
Closest in time.
Reasoning with latent thoughts: On the power of looped transformers
Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, and Sashank J Reddi · 2025
Closest in time.
Rewarding progress: Scaling automated process verifiers for LLM reasoning
Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, and Aviral Kumar · 2025
Closest in time.
R1-multimodal-journey
Wenqi Shao, Qiaosheng Zhang, Lingxiao Du, Xiangyan Liu, and Fanqing Meng · 2025
Closest in time.
R-prm: Reasoning-driven process reward modeling
Shuaijie She, Junxiao Liu, Yifeng Liu, Jiajun Chen, Xin Huang, and Shujian Huang · 2025
Closest in time.
Vlm-r1: A stable and generalizable r1-style large vision-language model
Haozhan Shen, Zilun Zhang, Qianqian Zhang, Ruochen Xu, and Tiancheng Zhao · 2025
Closest in time.
Safal Shrestha, Minwu Kim, and Keith Ross · 2025
Closest in time.
Layer by layer: Uncovering hidden representations in language models
Oscar Skean, Md Rifat Arefin, Dan Zhao, Niket Patel, Jalal Naghiyev, Yann LeCun, and Ravid Shwartz-Ziv · 2025
Closest in time.
R1-searcher: Incentivizing the search capability in llms via reinforcement learning
Huatong Song, Jinhao Jiang, Yingqian Min, Jie Chen, Zhipeng Chen, Wayne Xin Zhao, Lei Fang, and Ji-Rong Wen · 2025
Closest in time.
Towards reasoning ability of small language models
Gaurav Srivastava, Shuxiang Cao, and Xuan Wang · 2025
Closest in time.
Expanding rl with verifiable rewards across diverse domains
Yi Su, Dian Yu, Linfeng Song, Juntao Li, Haitao Mi, Zhaopeng Tu, Min Zhang, and Dong Yu · 2025
Closest in time.
Visual agents as fast and slow thinkers
Guangyan Sun, Mingyu Jin, Zhenting Wang, Cheng-Long Wang, Siqi Ma, Qifan Wang, Tong Geng, Ying Nian Wu, Yongfeng Zhang, and Dongfang Liu · 2025
Closest in time.
Llm pretraining with continuous concepts
Jihoon Tack, Jack Lanchantin, Jane Yu, Andrew Cohen, Ilia Kulikov, Janice Lan, Shibo Hao, Yuandong Tian, Jason Weston, and Xian Li · 2025
Closest in time.
Reason-rft: Reinforcement fine-tuning for visual reasoning
Huajie Tan, Yuheng Ji, Xiaoshuai Hao, Minglan Lin, Pengwei Wang, Zhongyuan Wang, and Shanghang Zhang · 2025
Closest in time.
Lego-puzzles: How good are mllms at multi-step spatial reasoning?
Kexian Tang, Junyao Gao, Yanhong Zeng, Haodong Duan, Yanan Sun, Zhening Xing, Wenran Liu, Kaifeng Lyu, and Kai Chen · 2025
Closest in time.
Confidence improves self-consistency in llms
Amir Taubenfeld, Tom Sheffer, Eran Ofek, Amir Feder, Ariel Goldstein, Zorik Gekhman, and Gal Yona · 2025
Closest in time.
Dolphin R1
DolphinR1 Team · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
Think less, achieve more: Cut reasoning costs by 50% without sacrificing accuracy
NovaSky Team · 2025
Closest in time.
Sky-t1: Train your own o1 preview model within $ 450
NovaSky Team · 2025
Closest in time.
Atom of thoughts for markov llm test-time scaling
Fengwei Teng, Zhaoyang Yu, Quan Shi, Jiayi Zhang, Chenglin Wu, and Yuyu Luo · 2025
Closest in time.
Llamav-o1: Rethinking step-by-step visual reasoning in llms
Omkar Thawakar, Dinura Dissanayake, Ketan More, Ritesh Thawkar, Ahmed Heakl, Noor Ahsan, Yuhao Li, Mohammed Zumri, Jean Lahoud, Rao Muhammad Anwer, et al · 2025
Closest in time.
Webgames: Challenging general-purpose web-browsing ai agents
George Thomas, Alex J Chan, Jikun Kang, Wenqi Wu, Filippos Christianos, Fraser Greenlee, Andy Toulis, and Marvin Purtorab · 2025
Closest in time.
Think twice: Enhancing llm reasoning by scaling multi-round test-time thinking
Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yunjie Ji, Yiping Peng, Han Zhao, and Xiangang Li · 2025
Closest in time.
Interacting with ai reasoning models: Harnessing" thoughts" for ai-driven software engineering
Christoph Treude and Raula Gaikovina Kula · 2025
Closest in time.
Llms can teach themselves to better predict the future
Benjamin Turtel, Danny Franklin, and Philipp Schoenegger · 2025
Closest in time.
Measuring faithfulness of chains of thought by unlearning reasoning steps
Martin Tutek, Fateme Hashemi Chaleshtori, Ana Marasović, and Yonatan Belinkov · 2025
Closest in time.
Ignore the kl penalty! boosting exploration on critical tokens to enhance rl fine-tuning
Jean Vassoyan, Nathanaël Beau, and Roman Plaud · 2025
Closest in time.
Rema: Learning to meta-think for llms with multi-agent reinforcement learning
Ziyu Wan, Yunxiang Li, Yan Song, Hanjing Wang, Linyi Yang, Mark Schmidt, Jun Wang, Weinan Zhang, Shuyue Hu, and Ying Wen · 2025
Closest in time.
Ante Wang, Linfeng Song, Ye Tian, Dian Yu, Haitao Mi, Xiangyu Duan, Zhaopeng Tu, Jinsong Su, and Dong Yu · 2025
Closest in time.
DivIL: Unveiling and addressing over-invariance for out-of-distribution generalization
Jiaqi WANG, Yuhang Zhou, Zhixiong Zhang, Qiguang Chen, Yongqiang Chen, and James Cheng · 2025
Closest in time.
Mobile-agent-v: Learning mobile device operation through video-guided multi-agent collaboration
Junyang Wang, Haiyang Xu, Xi Zhang, Ming Yan, Ji Zhang, Fei Huang, and Jitao Sang · 2025
Closest in time.
Dynamic chain-of-thought: Towards adaptive deep reasoning
Libo Wang · 2025
Closest in time.
Anjiang Wei, Jiannan Cao, Ran Li, Hongyu Chen, Yuhui Zhang, Ziheng Wang, Yaofeng Sun, Yuan Liu, Thiago SFX Teixeira, Diyi Yang, et al · 2025
Closest in time.
Codeplan: Unlocking reasoning potential in large language models by scaling code-form planning
Jiaxin Wen, Jian Guan, Hongning Wang, Wei Wu, and Minlie Huang · 2025
Closest in time.
Livebench: A challenging, contamination-limited LLM benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Venkat Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblum · 2025
Closest in time.
Boosting multimodal reasoning with mcts-automated structured thinking
Jinyang Wu, Mingkuan Feng, Shuai Zhang, Ruihan Jin, Feihu Che, Zengqi Wen, and Jianhua Tao · 2025
Closest in time.
Tokenskip: Controllable chain-of-thought compression in llms
Heming Xia, Yongqi Li, Chak Tou Leong, Wenjie Wang, and Wenjie Li · 2025
Closest in time.
Towards system 2 reasoning in llms: Learning how to think with meta chain-of-though
Violet Xiang, Charlie Snell, Kanishk Gandhi, Alon Albalak, Anikait Singh, Chase Blagden, Duy Phung, Rafael Rafailov, Nathan Lile, Dakota Mahan, et al · 2025
Closest in time.
Robotic programmer: Video instructed policy code generation for robotic manipulation
Senwei Xie, Hongyu Wang, Zhanqi Xiao, Ruiping Wang, and Xilin Chen · 2025
Closest in time.
Towards large reasoning models: A survey of reinforced reasoning with large language models
Fengli Xu, Qianyue Hao, Zefang Zong, Jingwei Wang, Yunke Zhang, Jingyi Wang, Xiaochong Lan, Jiahui Gong, Tianjian Ouyang, Fanjin Meng, et al · 2025
Closest in time.
Kai Yan, Yufei Xu, Zhengyin Du, Xuesong Yao, Zheyu Wang, Xiaowen Guo, and Jiecao Chen · 2025
Closest in time.
Chartmimic: Evaluating LMM’s cross-modal reasoning capability via chart-to-code generation
Cheng Yang, Chufan Shi, Yaxin Liu, Bo Shui, Junjie Wang, Mohan Jing, Linran XU, Xinyu Zhu, Siheng Li, Yuxiang Zhang, Gongye Liu, Xiaomei Nie, Deng Cai, and Yujiu Yang · 2025
Closest in time.
Xinhao Yao, Ruifeng Ren, Yun Liao, and Yong Liu · 2025
Closest in time.
Multimodal chain-of-thought reasoning: A comprehensive survey
Wang Yaoting, Wu Shengqiong, Zhang Yuechen, Yan Shuicheng, Liu Ziwei, Luo Jiebo, and Fei Hao · 2025
Closest in time.
Multimodal rewardbench: Holistic evaluation of reward models for vision language models
Michihiro Yasunaga, Luke Zettlemoyer, and Marjan Ghazvininejad · 2025
Closest in time.
On the emergence of thinking in llms i: Searching for the right intuition
Guanghao Ye, Khiem Duc Pham, Xinzhi Zhang, Sivakanth Gopi, Baolin Peng, Beibin Li, Janardhan Kulkarni, and Huseyin A Inan · 2025
Closest in time.
Demystifying long chain-of-thought reasoning in llms
Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, and Xiang Yue · 2025
Closest in time.
Sppd: Self-training with process preference learning using dynamic value margin
Hao Yi, Qingyang Li, Yulan Hu, Fuzheng Zhang, Di Zhang, and Yong Liu · 2025
Closest in time.
Monte carlo tree diffusion for system 2 planning
Jaesik Yoon, Hyeonseo Cho, Doojin Baek, Yoshua Bengio, and Sungjin Ahn · 2025
Closest in time.
Uncertainty-aware search and value models: Mitigating search scaling flaws in llms
Fei Yu, Yingru Li, and Benyou Wang · 2025
Closest in time.
Advancing LLM reasoning generalists with preference trees
Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding, Xingyao Wang, Boji Shan, Zeyuan Liu, Jia Deng, Huimin Chen, Ruobing Xie, Yankai Lin, Zhenghao Liu, Bowen Zhou, Hao Peng, Zhiyuan Liu, and Maosong Sun · 2025
Closest in time.
Mammoth2: Scaling instructions from the web
Xiang Yue, Tianyu Zheng, Ge Zhang, and Wenhu Chen · 2025
Closest in time.
Optimizing generative ai by backpropagating language model feedback
Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Pan Lu, Zhi Huang, Carlos Guestrin, and James Zou · 2025
Closest in time.
Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks
YuYue, Yufeng Yuan, Qiying Yu, Xiaochen Zuo, Ruofei Zhu, Wenyuan Xu, Jiaze Chen, Chengyi Wang, TianTian Fan, Zhengyin Du, Xiangpeng Wei, Gaohong Liu, Juncai Liu, Lingjun Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Chi Zhang, Mofan Zhang, Wang Zhang, Hang Zhu, Ru Zhang, Xin Liu, Mingxuan Wang, Yonghui Wu, and Lin Yan · 2025
Closest in time.
Internlm-xcomposer2.5-reward: A simple yet effective multi-modal reward model
Yuhang Zang, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Ziyu Liu, Shengyuan Ding, Shenxi Wu, Yubo Ma, Haodong Duan, Wenwei Zhang, et al · 2025
Closest in time.
Acecoder: Acing coder rl via automated test-case synthesis
Huaye Zeng, Dongfu Jiang, Haozhe Wang, Ping Nie, Xiaotong Chen, and Wenhu Chen · 2025
Closest in time.
An evaluation of deepseek models in biomedical natural language processing
Zaifu Zhan, Shuang Zhou, Huixue Zhou, Jiawen Deng, Yu Hou, Jeremy Yeung, and Rui Zhang · 2025
Closest in time.
Codecriticbench: A holistic code critique benchmark for large language models
Alexander Zhang, Marcus Dong, Jiaheng Liu, Wei Zhang, Yejie Wang, Jian Yang, Ge Zhang, Tianyu Liu, Zhongyuan Peng, Yingshui Tan, et al · 2025
Closest in time.
BackMATH: Towards backward reasoning for solving math problems step by step
Shaowei Zhang and Deyi Xiong · 2025
Closest in time.
Sample, scrutinize and scale: Effective inference-time search by scaling verification
Eric Zhao, Pranjal Awasthi, and Sreenivas Gollapudi · 2025
Closest in time.
Yichong Zhao and Susumu Goto · 2025
Closest in time.
Enhancing llm reliability via explicit knowledge boundary modeling
Hang Zheng, Hongshen Xu, Yuncong Liu, Lu Chen, Pascale Fung, and Kai Yu · 2025
Closest in time.
Dyve: Thinking fast and slow for dynamic process verification
Jianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen, and Qiang Xu · 2025
Closest in time.
Changzhi Zhou, Xinyu Zhang, Dandan Song, Xiancai Chen, Wanli Gu, Huipeng Ma, Yuhang Tian, Mengdi Zhang, and Linmei Hu · 2025
Closest in time.
Chain-of-thought matters: Improving long-context language models with reasoning path supervision
Dawei Zhu, Xiyu Wei, Guangxiang Zhao, Wenhao Wu, Haosheng Zou, Junfeng Ran, Xun Wang, Lin Sun, Xiangzheng Zhang, and Sujian Li · 2025
Closest in time.
Reasoning on a spectrum: Aligning llms to system 1 and system 2 thinking
Alireza S Ziabari, Nona Ghazizadeh, Zhivar Sourati, Farzan Karimi-Malekabadi, Payam Piray, and Morteza Dehghani · 2025
Closest in time.
Testnuc: Enhancing test-time computing approaches through neighboring unlabeled data consistency
Henry Peng Zou, Zhengyao Gu, Yue Zhou, Yankai Chen, Weizhi Zhang, Liancheng Fang, Yibo Wang, Yangning Li, Kay Liu, and Philip S Yu · 2025
Closest in time.
Medxpertqa: Benchmarking expert-level medical reasoning and understanding
Yuxin Zuo, Shang Qu, Yifei Li, Zhangren Chen, Xuekai Zhu, Ermo Hua, Kaiyan Zhang, Ning Ding, and Bowen Zhou · 2025
Closest in time.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits · 2076
Closest in time.