Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) plays a crucial role in aligning large language models (LLMs) with human preferences and improving their ability to perform complex tasks.
Rank analysis of incomplete block designs: I. the method of paired comparisons
1952
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
2017
Earlier work this paper cites.
Proximal policy optimization algorithms
2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
2018
Earlier work this paper cites.
Learning by playing solving sparse reward tasks from scratch
2018
Earlier work this paper cites.
On the weaknesses of reinforcement learning for neural machine translation
2020
Earlier work this paper cites.
Training verifiers to solve math word problems
2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
2022
Earlier work this paper cites.
Monte carlo augmented actor-critic for sparse reward deep reinforcement learning from suboptimal demonstrations
2022
Earlier work this paper cites.
Gpt-4 technical report
2023
Cited alongside, same era.
Raft: Reward ranked finetuning for generative foundation model alignment
2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
2023
Cited alongside, same era.
Let’s verify step by step
2023
Cited alongside, same era.
Offline rl for natural language generation with implicit language q learning
2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
2023
Cited alongside, same era.
Breadcrumbs to the goal: goal-conditioned exploration from human-in-the-loop feedback
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms
2024
Closest in time.
Numinamath
2024
Closest in time.
Improving multi-step reasoning abilities of large language models with direct advantage policy optimization
2024
Closest in time.
Selfcheck: Using llms to zero-shot check their own step-by-step reasoning
2024
Closest in time.
Nash learning from human feedback
2024
Closest in time.
Hiql: Offline goal-conditioned rl with latent states as actions
2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
2023
Cited alongside, same era.
Zephyr: Direct distillation of lm alignment
2023
Cited alongside, same era.
Slic-hf: Sequence likelihood calibration with human feedback
2023
Cited alongside, same era.
Enhancing textbook question answering task with large language models and retrieval augmented generation
2024
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
2024
Cited alongside, same era.
2024
Closest in time.
Offline regularised reinforcement learning for large language models alignment
2024
Closest in time.
Multi-turn reinforcement learning from preference human feedback
2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
2024
Closest in time.
Hybridflow: A flexible and efficient rlhf framework
2024
Closest in time.
Gemma: Open models based on gemini research and technology
2024
Closest in time.
Self-play preference optimization for language model alignment
2024
Closest in time.
Monte carlo tree search boosts reasoning via iterative preference learning
2024
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
2024
Closest in time.