Fetching the paper…
Reading the bibliography…
Reasoning language models (RLMs), also known as Large Reasoning Models (LRMs), such as OpenAI's o1 and o3, DeepSeek-R1, and Alibaba's QwQ, have redefined AI's problem-solving capabilities by extending LLMs with advanced reasoning mechanisms.
On the Measure of Intelligence, Nov. 2019
F. Chollet · 1911
Earlier work this paper cites.
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Coarse-to-Fine n-Best Parsing and MaxEnt Discriminative Reranking
E. Charniak and M. Johnson · 2005
Earlier work this paper cites.
Bandit Based Monte-Carlo Planning
L. Kocsis and C. Szepesvári · 2006
Earlier work this paper cites.
Cognitive Analysis Techniques in Business Planning and Decision Support Systems
R. Tadeusiewicz, L. Ogiela, and M. R. Ogiela · 2006
Earlier work this paper cites.
Modern Methods for the Cognitive Analysis of Economic Data and Text Documents and Their Application in Enterprise Management
R. Tadeusiewicz and L. Ogiela · 2008
Earlier work this paper cites.
Multi-Armed Bandits with Episode Context
C. D. Rosin · 2011
Earlier work this paper cites.
Learning to Solve Arithmetic Word Problems with Verb Categorization
M. J. Hosseini, H. Hajishirzi, O. Etzioni, and N. Kushman · 2014
Earlier work this paper cites.
Solving General Arithmetic Word Problems
S. Roy and D. Roth · 2015
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2015
Earlier work this paper cites.
Distinguishing Cause from Effect Using Observational Data: Methods and Benchmarks
J. M. Mooij, J. Peters, D. Janzing, J. Zscheischler, and B. Schölkopf · 2016
Earlier work this paper cites.
Mastering the Game of Go with Deep Neural Networks and Tree Search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms, Aug. 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Mastering the Game of Go without Human Knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis · 2017
Earlier work this paper cites.
The Culture of AI: Everyday Life and the Digital Revolution
A. Elliott · 2018
Earlier work this paper cites.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal · 2018
Earlier work this paper cites.
Ray: A Distributed Framework for Emerging AI Applications
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, and I. Stoica · 2018
Earlier work this paper cites.
A General Reinforcement Learning Algorithm that Masters Chess, Shogi, and Go Through Self-Play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis · 2018
Earlier work this paper cites.
Diverse Beam Search for Improved Description of Complex Scenes
A. Vijayakumar, M. Cogswell, R. Selvaraju, Q. Sun, S. Lee, D. Crandall, and D. Batra · 2018
Earlier work this paper cites.
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference
R. Zellers, Y. Bisk, R. Schwartz, and Y. Choi · 2018
Earlier work this paper cites.
MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi · 2019
Earlier work this paper cites.
PHYRE: A New Benchmark for Physical Reasoning
A. Bakhtin, L. van der Maaten, J. Johnson, L. Gustafson, and R. Girshick · 2019
Earlier work this paper cites.
Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
T. Ben-Nun and T. Hoefler · 2019
Earlier work this paper cites.
Artificial Intelligence: Reshaping Life and Business
P. Kumar · 2019
Earlier work this paper cites.
Probing Neural Network Comprehension of Natural Language Arguments
T. Niven and H.-Y. Kao · 2019
Earlier work this paper cites.
Social IQa: Commonsense Reasoning about Social Interactions
M. Sap, H. Rashkin, D. Chen, R. Le Bras, and Y. Choi · 2019
Earlier work this paper cites.
CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text
K. Sinha, S. Sodhani, J. Dong, J. Pineau, and W. L. Hamilton · 2019
Earlier work this paper cites.
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
A. Talmor, J. Herzig, N. Lourie, and J. Berant · 2019
Earlier work this paper cites.
Neuropathic Pain Diagnosis Simulator for Causal Discovery Algorithm Evaluation
R. Tu, K. Zhang, B. Bertilson, H. Kjellstrom, and C. Zhang · 2019
Earlier work this paper cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Y. Bisk, R. Zellers, R. Le Bras, J. Gao, and Y. Choi · 2020
Earlier work this paper cites.
Artificial Intelligence: Reshaping the Practice of Radiological Sciences in the 21st Century
I. El Naqa, M. A. Haider, M. L. Giger, and R. K. Ten Haken · 2020
Earlier work this paper cites.
Retrieval-Augmented Language Model Pre-Training
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M. Chang · 2020
Earlier work this paper cites.
The Curious Case of Neural Text Degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2020
Earlier work this paper cites.
Evaluating the Factual Consistency of Abstractive Text Summarization
W. Kryscinski, B. McCann, C. Xiong, and R. Socher · 2020
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela · 2020
Earlier work this paper cites.
Adversarial NLI: A New Benchmark for Natural Language Understanding
Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela · 2020
Earlier work this paper cites.
WinoGrande: An Adversarial Winograd Schema Challenge at Scale
K. Sakaguchi, R. Le Bras, C. Bhagavatula, and Y. Choi · 2020
Earlier work this paper cites.
Mastering Atari, Go, Chess and Shogi by Planning With a Learned Model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. Lillicrap, and D. Silver · 2020
Earlier work this paper cites.
Learning to Summarize with Human Feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Program Synthesis with Large Language Models, Aug. 2021
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton · 2021
Earlier work this paper cites.
GraphMineSuite: Enabling High-Performance and Programmable Graph Mining Algorithms with Set Algebra
M. Besta, Z. Vonarburg-Shmaria, Y. Schaffner, L. Schwarz, G. Kwaśniewski, L. Gianinazzi, J. Beranek, K. Janda, T. Holenstein, S. Leisinger, P. Tatkowski, E. Ozdemir, A. Balla, M. Copik, P. Lindenberger, M. Konieczny, O. Mutlu, and T. Hoefler · 2021
Earlier work this paper cites.
AI-Enabled Business-Model Innovation and Transformation in Industrial Ecosystems: A Framework, Model and Outline for Further Research
T. Burström, V. Parida, T. Lahti, and J. Wincent · 2021
Earlier work this paper cites.
GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning
J. Chen, J. Tang, J. Qin, X. Liang, L. Liu, E. Xing, and L. Lin · 2021
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code, July 2021
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems, Nov. 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
SeBS: A Serverless Benchmark Suite for Function-as-a-Service Computing
M. Copik, G. Kwaśniewski, M. Besta, M. Podstawski, and T. Hoefler · 2021
Earlier work this paper cites.
Measuring Coding Challenge Competence with APPS
D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Measuring Mathematical Problem Solving with the MATH Dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Benchmarking of Data-Driven Causality Discovery Approaches in the Interactions of Arctic Sea Ice and Atmosphere
Y. Huang, M. Kleindessner, A. Munishkin, D. Varshney, P. Guo, and J. Wang · 2021
Earlier work this paper cites.
LISA: Language Models of ISAbelle Proofs
A. Q. Jiang, W. Li, J. M. Han, and Y. Wu · 2021
Earlier work this paper cites.
Towards Demystifying Serverless Machine Learning Training
J. Jiang, S. Gan, Y. Liu, F. Wang, G. Alonso, A. Klimovic, A. Singla, W. Wu, and C. Zhang · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
X. L. Li and P. Liang · 2021
Earlier work this paper cites.
Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
P. Lu, R. Gong, S. Jiang, L. Qiu, S. Huang, X. Liang, and S.-C. Zhu · 2021
Earlier work this paper cites.
Uncertainty Estimation in Autoregressive Structured Prediction
A. Malinin and M. Gales · 2021
Earlier work this paper cites.
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
M. Shridhar, X. Yuan, M.-A. Côté, Y. Bisk, A. Trischler, and M. Hausknecht · 2021
Earlier work this paper cites.
ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language
O. Tafjord, B. Dalvi, and P. Clark · 2021
Earlier work this paper cites.
Let’s Wait Awhile: How Temporal Workload Shifting Can Reduce Carbon Emissions in the Cloud
P. Wiesner, I. Behnke, D. Scheinert, K. Gontarska, and L. Thamsen · 2021
Earlier work this paper cites.
Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation
Q. Bao, A. Peng, T. Hartill, N. Tan, Z. Deng, M. Witbrock, and J. Liu · 2022
Earlier work this paper cites.
Motif Prediction with Graph Neural Networks
M. Besta, R. Grob, C. Miglioli, N. Bernold, G. Kwaśniewski, G. Gjini, R. Kanakagiri, S. Ashkboos, L. Gianinazzi, N. Dryden, and T. Hoefler · 2022
Earlier work this paper cites.
UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression
J. Chen, T. Li, J. Qin, P. Lu, L. Lin, C. Chen, and X. Liang · 2022
Earlier work this paper cites.
Noise in the Clouds: Influence of Network Performance Variability on Application Scalability
D. De Sensi, T. De Matteis, K. Taranov, S. Di Girolamo, T. Rahn, and T. Hoefler · 2022
Earlier work this paper cites.
CRASS: A Novel Data Set and Benchmark to Test Counterfactual Reasoning of Large Language Models
J. Frohberg and F. Binder · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
Relationships 5.0: How AI, VR, and Robots Will Reshape Our Emotional Lives
E. Kislev · 2022
Cited alongside, same era.
WANLI: Worker and AI Collaboration for Natural Language Inference Dataset Creation
A. Liu, S. Swayamdipta, N. A. Smith, and Y. Choi · 2022
Cited alongside, same era.
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
A. Masry, X. L. Do, J. Q. Tan, S. Joty, and E. Hoque · 2022
Cited alongside, same era.
Self-Critiquing Models for Assisting Human Evaluators, June 2022
X. Guan, Y. Liu, X. Lu, B. Cao, B. He, X. Han, L. Sun, J. Lou, B. Yu, Y. Lu, and H. Lin · 2024
Later among the works it cites.
FOLIO: Natural Language Reasoning with First-Order Logic
S. Han, H. Schoelkopf, Y. Zhao, Z. Qi, M. Riddell, W. Zhou, J. Coady, D. Peng, Y. Qiao, L. Benson, L. Sun, A. Wardle-Solano, H. Szabó, E. Zubova, M. Burtell, J. Fan, Y. Liu, B. Wong, M. Sailor, A. Ni, L. Nan, J. Kasai, T. Yu, R. Zhang, A. Fabbri, W. M. Kryscinski, S. Yavuz, Y. Liu, X. V. Lin, S. Joty, Y. Zhou, C. Xiong, R. Ying, A. Cohan, and D. Radev · 2024
Later among the works it cites.
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
A. Havrilla, S. C. Raparthy, C. Nalmpantis, J. Dwivedi-Yu, M. Zhuravinskyi, E. Hambro, and R. Raileanu · 2024
Later among the works it cites.
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
C. He, R. Luo, Y. Bai, S. Hu, Z. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, J. Liu, L. Qi, Z. Liu, and M. Sun · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Saunders, C. Yeh, J. Wu, S. Bills, L. Ouyang, J. Ward, and J. Leike · 2022
Cited alongside, same era.
Solving Math Word Problems with Process-and Outcome-Based Feedback, Nov. 2022
J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins · 2022
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou · 2022
Cited alongside, same era.
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
S. Yao, H. Chen, J. Yang, and K. Narasimhan · 2022
Cited alongside, same era.
AbductionRules: Training Transformers to Explain Unexpected Inputs
N. Young, Q. Bao, J. Bensemann, and M. Witbrock · 2022
Cited alongside, same era.
MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual Data
Y. Zhao, Y. Li, C. Li, and R. Zhang · 2022
Cited alongside, same era.
miniF2F: A Cross-System Benchmark for Formal Olympiad-Level Mathematics
K. Zheng, J. M. Han, and S. Polu · 2022
Cited alongside, same era.
Large Language Models Cannot Self-Correct Reasoning Yet
J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou · 2024
Later among the works it cites.
SWE-bench: Can Language Models Resolve Real-World Github Issues?
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. R. Narasimhan · 2024
Later among the works it cites.
OpenAI Unveils New A.I. That Can ‘Reason’ Through Math and Science Problems
W. Knight · 2024
Later among the works it cites.
MARIO: MAth Reasoning with code Interpreter Output - A Reproducible Pipeline
M. Liao, C. Li, W. Luo, W. Jing, and K. Fan · 2024
Later among the works it cites.
Let’s Verify Step by Step
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2024
Later among the works it cites.
AgentBench: Evaluating LLMs as Agents
X. Liu, H. Yu, H. Zhang, Y. Xu, X. Lei, H. Lai, Y. Gu, H. Ding, K. Men, K. Yang, S. Zhang, X. Deng, A. Zeng, Z. Du, C. Zhang, S. Shen, T. Zhang, Y. Su, H. Sun, M. Huang, Y. Dong, and J. Tang · 2024
Later among the works it cites.
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao · 2024
Later among the works it cites.
Improve Mathematical Reasoning in Language Models by Automated Process Supervision, Dec. 2024
L. Luo, Y. Liu, R. Liu, S. Phatale, M. Guo, H. Lara, Y. Li, L. Shu, Y. Zhu, L. Meng, J. Sun, and A. Rastogi · 2024
Later among the works it cites.
M. Luo, S. Kumbhar, M. Shen, M. Parmar, N. Varshney, P. Banerjee, S. Aditya, and C. Baral · 2024
Later among the works it cites.
Improving Language Modeling by Increasing Test-Time Planning Compute
F. Mai, N. Cornille, and M.-F. Moens · 2024
Later among the works it cites.
R. Manvi, A. Singh, and S. Ermon · 2024
Later among the works it cites.
CHAMP: A Competition-Level Dataset for Fine-Grained Analyses of LLMs’ Mathematical Reasoning Capabilities
Y. Mao, Y. Kim, and Y. Zhou · 2024
Later among the works it cites.
GAIA: A Benchmark for General AI Assistants
G. Mialon, C. Fourrier, T. Wolf, Y. LeCun, and T. Scialom · 2024
Later among the works it cites.
SpotServe: Serving Generative Large Language Models on Preemptible Instances
X. Miao, C. Shi, J. Duan, X. Xi, D. Lin, B. Cui, and Z. Jia · 2024
Later among the works it cites.
Introducing ChatGPT
OpenAI · 2024
Later among the works it cites.
Introducing OpenAI o1
OpenAI · 2024
Later among the works it cites.
Iterative Reasoning Preference Optimization
R. Y. Pang, W. Yuan, H. He, K. Cho, S. Sukhbaatar, and J. E. Weston · 2024
Later among the works it cites.
O1 Replication Journey: A Strategic Progress Report – Part 1, Oct. 2024
Y. Qin, X. Li, H. Zou, Y. Liu, S. Xia, Z. Huang, Y. Ye, W. Yuan, H. Liu, Y. Li, and P. Liu · 2024
Later among the works it cites.
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
Y. Qu, T. Zhang, N. Garg, and A. Kumar · 2024
Later among the works it cites.
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2024
Later among the works it cites.
S. Srivastava, A. M. B, A. P. V, S. Menon, A. Sukumar, A. S. T, A. Philipose, S. Prince, and S. Thomas · 2024
Later among the works it cites.
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
Z. Tang, X. Zhang, B. Wang, and F. Wei · 2024
Later among the works it cites.
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
Y. Tian, B. Peng, L. Song, L. Jin, D. Yu, L. Han, H. Mi, and D. Yu · 2024
Later among the works it cites.
AlphaZero-Like Tree-Search Can Guide Large Language Model Decoding and Training
Z. Wan, X. Feng, M. Wen, S. M. McAleer, Y. Wen, W. Zhang, and J. Wang · 2024
Later among the works it cites.
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models, Oct. 2024
J. Wang, M. Fang, Z. Wan, M. Wen, J. Zhu, A. Liu, Z. Gong, Y. Song, L. Chen, L. M. Ni, L. Yang, Y. Wen, and W. Zhang · 2024
Later among the works it cites.
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
K. Wang, H. Ren, A. Zhou, Z. Lu, S. Luo, W. Shi, R. Zhang, L. Song, M. Zhan, and H. Li · 2024
Later among the works it cites.
Math-Shepherd: Verify and Reinforce LLMs Step-by-Step without Human Annotations
P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui · 2024
Later among the works it cites.
X. Wang, L. Song, Y. Tian, D. Yu, B. Peng, H. Mi, F. Huang, and D. Yu · 2024
Later among the works it cites.
AgentGym: Evolving Large Language Model-Based Agents Across Diverse Environments, June 2024
Z. Xi, Y. Ding, W. Chen, B. Hong, H. Guo, J. Wang, D. Yang, C. Liao, X. Guo, W. He, S. Gao, L. Chen, R. Zheng, Y. Zou, T. Gui, Q. Zhang, X. Qiu, X. Huang, Z. Wu, and Y.-G. Jiang · 2024
Later among the works it cites.
Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Y. Xie, A. Goyal, W. Zheng, M.-Y. Kan, T. P. Lillicrap, K. Kawaguchi, and M. Shieh · 2024
Later among the works it cites.
Benchmarking Retrieval-Augmented Generation for Medicine
G. Xiong, Q. Jin, Z. Lu, and A. Zhang · 2024
Later among the works it cites.
Y. Yan, J. Su, J. He, F. Fu, X. Zheng, Y. Lyu, K. Wang, S. Wang, Q. Wen, and X. Hu · 2024
Later among the works it cites.
Free Process Rewards without Process Labels, Dec. 2024
L. Yuan, W. Li, H. Chen, G. Cui, N. Ding, K. Zhang, B. Zhou, Z. Liu, and H. Peng · 2024
Later among the works it cites.
Z. Zeng, Q. Cheng, Z. Yin, B. Wang, S. Li, Y. Zhou, Q. Guo, X. Huang, and X. Qiu · 2024
Later among the works it cites.
D. Zhang, X. Huang, D. Zhou, Y. Li, and W. Ouyang · 2024
Later among the works it cites.
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
D. Zhang, S. Zhoubian, Z. Hu, Y. Yue, Y. Dong, and J. Tang · 2024
Later among the works it cites.
Generative Verifiers: Reward Modeling as Next-Token Prediction
L. Zhang, A. Hosseini, H. Bansal, M. Kazemi, A. Kumar, and R. Agarwal · 2024
Later among the works it cites.
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions, Nov. 2024
Y. Zhao, H. Yin, B. Zeng, H. Wang, T. Shi, C. Lyu, L. Wang, W. Luo, and K. Zhang · 2024
Later among the works it cites.
WebArena: A Realistic Web Environment for Building Autonomous Agents
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, T. Ou, Y. Bisk, D. Fried, U. Alon, and G. Neubig · 2024
Later among the works it cites.
AIME 2024
AI-MO · 2025
Closest in time.
AMC 2024
AI-MO · 2025
Closest in time.
Multi-Head RAG: Solving Multi-Aspect Problems with LLMs, June 2025
M. Besta, A. Kubicek, R. Gerstenberger, M. Chrapek, R. Niggli, P. Okanovic, Y. Zhu, P. Iff, M. Podstawski, L. Weitzendorf, M. Chi, J. Gajda, P. Nyczyk, J. Müller, H. Niewiadomski, and T. Hoefler · 2025
Closest in time.
Demystifying Chains, Trees, and Graphs of Thoughts, Feb. 2025
M. Besta, F. Memedi, Z. Zhang, R. Gerstenberger, N. Blach, P. Nyczyk, M. Copik, G. Kwaśniewski, J. Müller, L. Gianinazzi, et al · 2025
Closest in time.
CheckEmbed: Effective Verification of LLM Solutions to Open-Ended Tasks, June 2025
M. Besta, L. Paleari, M. Copik, R. Gerstenberger, A. Kubicek, P. Nyczyk, P. Iff, E. Schreiber, T. Srindran, T. Lehmann, H. Niewiadomski, and T. Hoefler · 2025
Closest in time.
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs, Jan. 2025
K. Chernyshev, V. Polshkov, E. Artemova, A. Myasnikov, V. Stepanov, A. Miasnikov, and S. Tilga · 2025
Closest in time.
Process Reinforcement Through Implicit Rewards, Feb. 2025
G. Cui, L. Yuan, Z. Wang, H. Wang, W. Li, B. He, Y. Fan, T. Yu, Q. Xu, W. Chen, J. Yuan, H. Chen, K. Zhang, X. Lv, S. Wang, Y. Yao, X. Han, H. Peng, Y. Cheng, Z. Liu, M. Sun, B. Zhou, and N. Ding · 2025
Closest in time.
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking, Jan. 2025
X. Guan, L. L. Zhang, Y. Liu, N. Shang, Y. Sun, Y. Zhu, F. Yang, and M. Yang · 2025
Closest in time.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, Jan. 2025
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Search-o1: Agentic Search-Enhanced Large Reasoning Models, Jan. 2025
X. Li, G. Dong, J. Jin, Y. Zhang, Y. Zhou, Y. Zhu, P. Zhang, and Z. Dou · 2025
Closest in time.
DeepSeek-V3 Technical Report, Feb. 2025
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al · 2025
Closest in time.
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models
Y. Lyu, Z. Li, S. Niu, F. Xiong, B. Tang, W. Wang, H. Wu, H. Liu, T. Xu, and E. Chen · 2025
Closest in time.
In Two Moves, AlphaGo and Lee Sedol Redefined the Future
C. Metz · 2025
Closest in time.
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
S. I. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M. Farajtabar · 2025
Closest in time.
Hello GPT-4o
OpenAI · 2025
Closest in time.
Scaling Test-Time Compute Optimally Can Be More Effective than Scaling Model Parameters
C. V. Snell, J. Lee, K. Xu, and A. Kumar · 2025
Closest in time.
A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
J. Sun, C. Zheng, E. Xie, Z. Liu, R. Chu, J. Qiu, J. Xu, M. Ding, H. Li, M. Geng, Y. Wu, W. Wang, J. Chen, Z. Yin, X. Ren, J. Fu, J. He, Y. Wu, Q. Liu, X. Liu, Y. Li, H. Dong, Y. Cheng, M. Zhang, P. A. Heng, J. Dai, P. Luo, J. Wang, J.-R. Wen, X. Qiu, Y. Guo, H. Xiong, Q. Liu, and Z. Li · 2025
Closest in time.
QwQ: Reflect Deeply on the Boundaries of the Unknown
Q. Team · 2025
Closest in time.
LLaMA-Berry: Pairwise Optimization for Olympiad-Level Mathematical Reasoning via O1-like Monte Carlo Tree Search
D. Zhang, J. Wu, J. Lei, T. Che, J. Li, T. Xie, X. Huang, S. Zhang, M. Pavone, Y. Li, W. Ouyang, and D. Zhou · 2025
Closest in time.
D.-H. Zhu, Y.-J. Xiong, J.-C. Zhang, X.-J. Xie, and C.-M. Xia · 2025
Closest in time.