Fetching the paper…
Reading the bibliography…
Standard reinforcement learning from human feedback (RLHF) approaches relying on parametric models like the Bradley-Terry model fall short in capturing the intransitivity and irrationality in human preferences.
Hellaswag: Can a machine really finish your sentence?
Zellers, R · 1905
Earlier work this paper cites.
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
Bradley, R. A · 1952
Earlier work this paper cites.
Intransitivity of preferences
Tversky, A · 1969
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y · 1997
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Freund, Y · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S · 1999
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D · 2009
Earlier work this paper cites.
Contextual dueling bandits
Dudík, M · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T · 2018
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K · 2021
Earlier work this paper cites.
Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing
He, P · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K · 2021
Cited alongside, same era.
Active ranking without strong stochastic transitivity
Lou, H · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G · 2023
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Zhu, B · 2023
Later among the works it cites.
Human alignment of large language models through online preference optimisation
Calandriello, D · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models
Chen, Z · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K · 2024
Closest in time.
Rebel: Reinforcement learning via regressing relative rewards
Gao, Z · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cui, G · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L · 2023
Cited alongside, same era.
Statistical rejection sampling improves preference optimization
Liu, T · 2023
Cited alongside, same era.
Nash learning from human feedback
Munos, R · 2023
Cited alongside, same era.
OpenAI, J., Achiam · 2023
Cited alongside, same era.
Beyond human data: Scaling self-training for problem-solving with language models
Singh, A · 2023
Cited alongside, same era.
Borda regret minimization for generalized linear dueling bandits
Wu, Y · 2023
Cited alongside, same era.
Reference-free monolithic preference optimization with odds ratio
Hong, J · 2024
Closest in time.
Reinforcement learning from human feedback with active queries
Ji, K · 2024
Closest in time.
From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline
Li, T · 2024
Closest in time.
Smaug: Fixing failure modes of preference optimisation with dpo-positive
Pal, A · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Swamy, G · 2024
Closest in time.
Is rlhf more difficult than standard rl? a theoretical perspective
Wang, Y · 2024
Closest in time.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
Ye, C · 2024
Closest in time.
Self-rewarding language models
Yuan, W · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L · 2024
Closest in time.