Fetching the paper…
Reading the bibliography…
Mathematical reasoning is a crucial capability for Large Language Models (LLMs), yet generating detailed and accurate reasoning traces remains a significant challenge.
Training verifiers to solve math word problems
Cobbe, K · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D · 2021
Earlier work this paper cites.
Zelikman, E · 2022
Earlier work this paper cites.
Lightman, H · 2023
Earlier work this paper cites.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P · 2023
Earlier work this paper cites.
Beyond human data: Scaling self-training for problem-solving with language models
Singh, A · 2023
Earlier work this paper cites.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L · 2023
Cited alongside, same era.
Scaling relationship on learning mathematical reasoning with large language models
Yuan, Z · 2023
Cited alongside, same era.
Omni-math: A universal olympiad level mathematic benchmark for large language models
Gao, B · 2024
Cited alongside, same era.
V-star: Training verifiers for self-taught reasoners
Hosseini, A · 2024
Cited alongside, same era.
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms
Liu, H · 2024
Closest in time.
Online joint fine-tuning of multi-agent flows
Mineiro, P · 2024
Closest in time.
Iterative reasoning preference optimization
Pang, R. Y · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R · 2024
Closest in time.
Quiet-star: Language models can teach themselves to think before speaking
Zelikman, E · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lai, X · 2024
Cited alongside, same era.
Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Wang, P
Cited in the paper.
Wang, Z
Cited in the paper.
Llama-berry: Pairwise optimization for o1-like olympiad-level mathematical reasoning
Zhang, D
Cited in the paper.
Rest-mcts*: Llm self-training via process reward guided tree search
Zhang, D
Cited in the paper.
Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
Zhang, R
Cited in the paper.