Fetching the paper…
Reading the bibliography…
State-of-the-art language model fine-tuning techniques, such as Direct Preference Optimization (DPO), restrict user control by hard-coding predefined behaviors into the model.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Earlier work this paper cites.
Distilled self-critique of llms with synthetic data: a bayesian perspective
V. Gallego · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models, 2023
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson · 2023
Cited alongside, same era.
Suppressing pink elephants with direct principle feedback
L. Castricato, N. Lile, S. Anand, H. Schoelkopf, S. Verma, and S. Biderman · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…