Fetching the paper…
Reading the bibliography…
Preference optimization, particularly through Reinforcement Learning from Human Feedback (RLHF), has achieved significant success in aligning Large Language Models (LLMs) to adhere to human intentions.
Hellaswag: Can a machine really finish your sentence?
Zellers, R · 1905
Earlier work this paper cites.
The rating of chessplayers: Past and present
Elo, A. E · 1978
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Strens, M · 2000
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
From ε \varepsilon -entropy to kl-entropy: Analysis of minimum information complexity density estimation
Zhang, T · 2006
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Russo, D · 2013
Earlier work this paper cites.
Ensemble sampling
Lu, X · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T · 2018
Earlier work this paper cites.
Flambe: Structural complexity and representation learning of low rank mdps
Agarwal, A · 2020
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y · 2022
Earlier work this paper cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Chen, X · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Earlier work this paper cites.
Conservative dual policy optimization for efficient model-based reinforcement learning
Zhang, S · 2022
Earlier work this paper cites.
Gec: A unified framework for interactive decision making in mdp, pomdp, and beyond
Zhong, H · 2022
Earlier work this paper cites.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G · 2023
Earlier work this paper cites.
Ultrafeedback: Boosting language models with high-quality feedback
Cui, G · 2023
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L · 2023
Cited alongside, same era.
Reinforced self-training (rest) for language modeling
Gulcehre, C · 2023
Cited alongside, same era.
Gunasekar, S · 2023
Cited alongside, same era.
Camels in a changing climate: Enhancing lm adaptation with tulu 2
Ivison, H · 2023
Cited alongside, same era.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Dubois, Y · 2024
Closest in time.
Efficient exploration for llms
Dwaracherla, V · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K · 2024
Closest in time.
Direct language model alignment from online ai feedback
Guo, S · 2024
Closest in time.
Snorkel-mistral-pairrm-dpo. https://huggingface.co/snorkelai/Snorkel-Mistral-PairRM-DPO
Hoang Tran, B. H., Chris Glaze · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, D · 2023
Cited alongside, same era.
Josifoski, M · 2023
Cited alongside, same era.
Sample efficient reinforcement learning from human feedback via active exploration
Mehta, V · 2023
Cited alongside, same era.
Approximate thompson sampling via epistemic neural networks
Osband, I · 2023
Cited alongside, same era.
Eq-bench: An emotional intelligence benchmark for large language models
Paech, S. J · 2023
Cited alongside, same era.
Peng, B · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Taori, R · 2023
Cited alongside, same era.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Hu, J · 2024
Closest in time.
sdpo: Don’t use your data all at once
Kim, D · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Köpf, A · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date. https://ai.meta.com/blog/meta-llama-3/
Meta · 2024
Closest in time.
Language model alignment with elastic reset
Noukhovitch, M · 2024
Closest in time.
Epistemic neural networks
Osband, I · 2024
Closest in time.
Smaug: Fixing failure modes of preference optimisation with dpo-positive
Pal, A · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C · 2024
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Sun, Z · 2024
Closest in time.
Understanding the performance gap between online and offline alignment algorithms
Tang, Y · 2024
Closest in time.
Self-play preference optimization for language model alignment
Wu, Y · 2024
Closest in time.
Exploratory preference optimization: Harnessing implicit q*-approximation for sample-efficient rlhf
Xie, T · 2024
Closest in time.
Is dpo superior to ppo for llm alignment? a comprehensive study
Xu, S · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Young, A · 2024
Closest in time.
Self-rewarding language models
Yuan, W · 2024
Closest in time.
Iterative reasoning preference optimization
Yuanzhe Pang, R · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L · 2024
Closest in time.
Dpo meets ppo: Reinforced token optimization for rlhf
Zhong, H · 2024
Closest in time.