Fetching the paper…
Reading the bibliography…
Recent studies have shown that large language models' (LLMs) mathematical problem-solving capabilities can be enhanced by integrating external tools, such as code interpreters, and employing multi-turn Chain-of-Thought (CoT) reasoning.
Problem complexity and method efficiency in optimization
A. S. Nemirovskij and D. B. Yudin · 1983
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
R. J. Williams and J. Peng · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search
T. Anthony, Z. Tian, and D. Barber · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
On the weaknesses of reinforcement learning for neural machine translation
L. Choshen, L. Fox, Z. Aizenbud, and O. Abend · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Q. Cai, Z. Yang, C. Jin, and Z. Wang · 2020
Earlier work this paper cites.
Implementation matters in deep policy gradients: A case study on ppo and trpo
L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Minif2f: a cross-system benchmark for formal olympiad-level mathematics
K. Zheng, J. M. Han, and S. Polu · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
W. Chen, X. Ma, X. Wang, and W. W. Cohen · 2022
Earlier work this paper cites.
Lila: A unified benchmark for mathematical reasoning
S. Mishra, M. Finlayson, P. Lu, L. Tang, S. Welleck, C. Baral, T. Rajpurohit, O. Tafjord, A. Sabharwal, P. Clark, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Chaining simultaneous thoughts for numerical reasoning
Z. Shao, F. Huang, and M. Huang · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
J. Uesato, N. Kushman, R. Kumar, F. Song, N. Siegel, L. Wang, A. Creswell, G. Irving, and I. Higgins · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
The role of coverage in online reinforcement learning
T. Xie, D. J. Foster, Y. Bai, N. Jiang, and S. M. Kakade · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. Goodman · 2022
Earlier work this paper cites.
Gec: A unified framework for interactive decision making in mdp, pomdp, and beyond
H. Zhong, W. Xiong, S. Zheng, L. Wang, Z. Wang, Z. Yang, and T. Zhang · 2022
Earlier work this paper cites.
Least-to-most prompting enables complex reasoning in large language models
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. Le, et al · 2022
Earlier work this paper cites.
Solving math word problems via cooperative reasoning induced language models
X. Zhu, J. Wang, L. Zhang, Y. Zhang, Y. Huang, R. Gan, J. Zhang, and Y. Yang · 2022
Earlier work this paper cites.
VO Q Q L: Towards optimal regret in model-free rl with nonlinear function approximation
A. Agarwal, Y. Jin, and T. Zhang · 2023
Earlier work this paper cites.
Introducing claude
Anthropic · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, M. Rowland, B. Piot, D. Guo, D. Calandriello, M. Valko, and R. Munos · 2023
Cited alongside, same era.
Reward model ensembles help mitigate overoptimization
T. Coste, U. Anwar, R. Kirk, and D. Krueger · 2023
Cited alongside, same era.
RAFT: Reward ranked finetuning for generative foundation model alignment
H. Dong, W. Xiong, D. Goyal, Y. Zhang, W. Chow, R. Pan, S. Diao, J. Zhang, K. SHUM, and T. Zhang · 2023
Cited alongside, same era.
Reinforced self-training (rest) for language modeling
C. Gulcehre, T. L. Paine, S. Srinivasan, K. Konyushkova, L. Weerts, A. Sharma, A. Siddhant, A. Ahern, M. Wang, C. Gu, et al · 2023
Cited alongside, same era.
Learning planning-based reasoning by trajectories collection and process reward synthesizing
F. Jiao, C. Qin, Z. Liu, N. F. Chen, and S. Joty · 2024
Closest in time.
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms
X. Lai, Z. Tian, Y. Chen, S. Yang, X. Peng, and J. Jia · 2024
Closest in time.
Augmenting math word problems via iterative question composing
H. Liu and A. C.-C. Yao · 2024
Closest in time.
Step-controlled dpo: Leveraging stepwise error for enhanced mathematical reasoning
Z. Lu, A. Zhou, K. Wang, H. Ren, W. Shi, J. Pan, and M. Zhan · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Cited alongside, same era.
Y. Lin, L. Tan, H. Lin, Z. Zheng, R. Pi, J. Zhang, S. Diao, H. Wang, H. Zhao, Y. Yao, et al · 2023
Cited alongside, same era.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
H. Luo, Q. Sun, C. Xu, P. Zhao, J. Lou, C. Tao, X. Geng, Q. Lin, S. Chen, and D. Zhang · 2023
Cited alongside, same era.
Nash learning from human feedback
R. Munos, M. Valko, D. Calandriello, M. G. Azar, M. Rowland, Z. D. Guo, Y. Tang, M. Geist, T. Mesnard, A. Michi, et al · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Cited alongside, same era.
Y. Meng, M. Xia, and D. Chen · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta · 2024
Closest in time.
Orca-math: Unlocking the potential of slms in grade school math
A. Mitra, H. Khanpour, C. Rosset, and A. Awadallah · 2024
Closest in time.
West-of-n: Synthetic preference generation for improved reward modeling
A. Pace, J. Mallinson, E. Malmi, S. Krause, and A. Severyn · 2024
Closest in time.
Iterative reasoning preference optimization
R. Y. Pang, W. Yuan, K. Cho, H. He, S. Sukhbaatar, and J. Weston · 2024
Closest in time.
Strengthening multimodal large language model with bootstrapped preference optimization
R. Pi, T. Han, W. Xiong, J. Zhang, R. Liu, R. Pan, and T. Zhang · 2024
Closest in time.
From r to q*: Your language model is secretly a q-function
R. Rafailov, J. Hejna, R. Park, and C. Finn · 2024
Closest in time.
Offline regularised reinforcement learning for large language models alignment
P. H. Richemond, Y. Tang, D. Guo, D. Calandriello, M. G. Azar, R. Rafailov, B. A. Pires, E. Tarassov, L. Spangher, W. Ellsworth, et al · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
C. Rosset, C.-A. Cheng, A. Mitra, M. Santacroce, A. Awadallah, and T. Xie · 2024
Closest in time.
Multi-turn reinforcement learning from preference human feedback
L. Shani, A. Rosenberg, A. Cassel, O. Lang, D. Calandriello, A. Zipori, H. Noga, O. Keller, B. Piot, I. Szpektor, et al · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. Li, Y. Wu, and D. Guo · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
G. Swamy, C. Dann, R. Kidambi, Z. S. Wu, and A. Agarwal · 2024
Closest in time.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
F. Tajwar, A. Singh, A. Sharma, R. Rafailov, J. Schneider, T. Xie, S. Ermon, C. Finn, and A. Kumar · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Y. Tang, Z. D. Guo, Z. Zheng, D. Calandriello, R. Munos, M. Rowland, P. H. Richemond, M. Valko, B. Á. Pires, and B. Piot · 2024
Closest in time.
Codegemma: Open code models based on gemma
C. Team · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Dart-math: Difficulty-aware rejection tuning for mathematical problem-solving
Y. Tong, X. Zhang, R. Wang, R. Wu, and J. He · 2024
Closest in time.
Openmathinstruct-1: A 1.8 million math instruction tuning dataset
S. Toshniwal, I. Moshkov, S. Narenthiran, D. Gitman, F. Jia, and I. Gitman · 2024
Closest in time.
Mint: Multi-turn interactive evaluation for tool-augmented llms with language feedback
X. Wang, Z. Wang, J. Liu, Y. Chen, L. Yuan, H. Peng, and H. Ji · 2024
Closest in time.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
C. Ye, W. Xiong, Y. Zhang, N. Jiang, and T. Zhang · 2024
Closest in time.
Advancing llm reasoning generalists with preference trees
L. Yuan, G. Cui, H. Wang, N. Ding, X. Wang, J. Deng, B. Shan, H. Chen, R. Xie, Y. Lin, et al · 2024
Closest in time.
Mammoth2: Scaling instructions from the web
X. Yue, T. Zheng, G. Zhang, and W. Chen · 2024
Closest in time.
Weak-to-strong extrapolation expedites alignment
C. Zheng, Z. Wang, H. Ji, M. Huang, and N. Peng · 2024
Closest in time.
Dpo meets ppo: Reinforced token optimization for rlhf
H. Zhong, G. Feng, W. Xiong, L. Zhao, D. He, J. Bian, and L. Wang · 2024
Closest in time.