Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks, yet generating reliable reasoning processes remains a significant challenge.
Fine-tuning language models from human preferences
Ziegler, D. M · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A · 1952
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
Dempster, A. P · 1977
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovskij, A. S · 1983
Earlier work this paper cites.
A view of the em algorithm that justifies incremental, sparse, and other variants
Neal, R. M · 1998
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Wirth, C · 2017
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Earlier work this paper cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K · 2021
Earlier work this paper cites.
Efficient (soft) q-learning for text generation with limited good data
Guo, H · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J · 2021
Earlier work this paper cites.
Dueling rl: reinforcement learning with trajectory preferences
Pacchiano, A · 2021
Earlier work this paper cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Chen, X · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A · 2022
Earlier work this paper cites.
Quark: Controllable text generation with reinforced unlearning
Lu, X · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Zelikman, E · 2022
Cited alongside, same era.
Least-to-most prompting enables complex reasoning in large language models
Zhou, D · 2022
Cited alongside, same era.
Introducing claude. https://www.anthropic.com/index/introducing-claude
Anthropic · 2023
Cited alongside, same era.
Reinforced self-training (rest) for language modeling
Gulcehre, C · 2023
Cited alongside, same era.
Introducing openai o1
OpenAI · 2024
Later among the works it cites.
Iterative reasoning preference optimization
Pang, R. Y · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R · 2024
Later among the works it cites.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D · 2024
Later among the works it cites.
Speculations on test-time scaling
Rush, S · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, E. J · 2023
Cited alongside, same era.
Jiang, A. Q · 2023
Cited alongside, same era.
Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Liu, J · 2023
Cited alongside, same era.
American mathematics competitions amc 12, 2023. https://maa.org/math-competitions/amc-12
Mathematical Association of America · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Beyond human data: Scaling self-training for problem-solving with language models
Singh, A · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H · 2023
Cited alongside, same era.
Hybridflow: A flexible and efficient rlhf framework
Sheng, G · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C · 2024
Later among the works it cites.
Generalized preference optimization: A unified approach to offline alignment
Tang, Y · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models. https://qwenlm.github.io/blog/qwen2.5/
Team, Q · 2024
Later among the works it cites.
Offline reinforcement learning for llm multi-step reasoning
Wang, H · 2024
Later among the works it cites.
Thinking llms: General instruction following with thought generation
Wu, T · 2024
Later among the works it cites.
Exploratory preference optimization: Harnessing implicit q*-approximation for sample-efficient rlhf
Xie, T · 2024
Later among the works it cites.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Xiong, W · 2024
Later among the works it cites.
Yang, R · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S · 2024
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback
Yuan, H · 2024
Later among the works it cites.
Quiet-star: Language models can teach themselves to think before speaking
Zelikman, E · 2024
Later among the works it cites.
Self-exploring language models: Active preference elicitation for online alignment
Zhang, S · 2024
Later among the works it cites.
Dpo meets ppo: Reinforced token optimization for rlhf
Zhong, H · 2024
Later among the works it cites.
A theoretical analysis of optimistic proximal policy optimization in linear markov decision processes
Zhong, H · 2024
Later among the works it cites.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Zhuo, T. Y · 2024
Later among the works it cites.