Fetching the paper…
Reading the bibliography…
Learning from preference feedback has emerged as an essential step for improving the generation quality and performance of modern language models (LMs).
Policy Gradient Methods for Reinforcement Learning with Function Approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Program Synthesis with Large Language Models, 2021
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Understanding Dataset Difficulty with 𝒱 \mathcal{V} -Usable Information
K. Ethayarajh, Y. Choi, and S. Swayamdipta · 2022
Earlier work this paper cites.
TOXIGEN: Controlling Language Models to Generate Implied and Adversarial Toxicity
T. Hartvigsen, S. Gabriel, H. Palangi, M. Sap, D. Ray, and E. Kamar · 2022
Earlier work this paper cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Earlier work this paper cites.
Rainier: Reinforced knowledge introspector for commonsense question answering
J. Liu, S. Hallinan, X. Lu, P. He, S. Welleck, H. Hajishirzi, and Y. Choi · 2022
Earlier work this paper cites.
Quark: Controllable Text Generation with Reinforced Unlearning
X. Lu, S. Welleck, L. Jiang, J. Hessel, L. Qin, P. West, P. Ammanabrolu, and Y. Choi · 2022
Earlier work this paper cites.
Challenging BIG-Bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Earlier work this paper cites.
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, M. Rowland, B. Piot, D. Guo, D. Calandriello, M. Valko, and R. Munos · 2023
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
B. bench authors · 2023
Earlier work this paper cites.
Ultrafeedback: Boosting language models with high-quality feedback
G. Cui, L. Yuan, N. Ding, G. Yao, W. Zhu, Y. Ni, G. Xie, Z. Liu, and M. Sun · 2023
Cited alongside, same era.
Amplify-Instruct: Synthetically Generated Diverse Multi-turn Conversations for Effecient LLM Training
L. Daniele and Suphavadeeprasit · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2023
Cited alongside, same era.
EasyLM: A Simple And Scalable Training Framework for Large Language Models, 2023
X. Geng · 2023
Cited alongside, same era.
HuggingFace H4 Stack Exchange Preference Dataset, 2023
N. Lambert, L. Tunstall, N. Rajani, and T. Thrush · 2023
Cited alongside, same era.
AlpacaEval: An Automatic Evaluator of Instruction-following Models
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, 2024
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, and I. Stoica · 2024
Closest in time.
Reward model ensembles help mitigate overoptimization
T. Coste, U. Anwar, R. Kirk, and D. Krueger · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela · 2024
Closest in time.
ORPO: Monolithic Preference Optimization without Reference Model, 2024
J. Hong, N. Lee, and J. Thorne · 2024
Closest in time.
Mixtral of Experts, 2024
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Li, T. Zhang, Y. Dubois, R. Taori, I. Gulrajani, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Cited alongside, same era.
Let’s Verify Step by Step, 2023
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Cited alongside, same era.
Supervised Fine-Tuning and Direct Preference Optimization on Intel Gaudi2, 2023
K. Lv, W. Zhang, and H. Shen · 2023
Cited alongside, same era.
OctoPack: Instruction Tuning Code Large Language Models
N. Muennighoff, Q. Liu, A. Zebaze, Q. Zheng, B. Hui, T. Y. Zhuo, S. Singh, X. Tang, L. von Werra, and S. Longpre · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Cited alongside, same era.
ChatGPT: Optimizing Language Models for Dialogue
J. Schulman, B. Zoph, C. Kim, and more · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Cited alongside, same era.
RewardBench: Evaluating Reward Models for Language Modeling, 2024
N. Lambert, V. Pyatkin, J. Morrison, L. Miranda, B. Y. Lin, K. Chandu, N. Dziri, S. Kumar, T. Zick, Y. Choi, N. A. Smith, and H. Hajishirzi · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date, May 2024
Meta · 2024
Closest in time.
WARM: On the Benefits of Weight Averaged Reward Models, 2024
A. Ramé, N. Vieillard, L. Hussenot, R. Dadashi, G. Cideron, O. Bachem, and J. Ferret · 2024
Closest in time.
Preference Tuning LLMs with Direct Preference Optimization Methods
K. Rasul, E. Beeching, L. Tunstall, L. von Werra, and O. Sanseviero · 2024
Closest in time.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models, 2024
P. Röttger, H. R. Kirk, B. Vidgen, G. Attanasio, F. Bianchi, and D. Hovy · 2024
Closest in time.
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
A. Saeidi, S. Verma, and C. Baral · 2024
Closest in time.
Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, 2024
F. Tajwar, A. Singh, A. Sharma, R. Rafailov, J. Schneider, T. Xie, S. Ermon, C. Finn, and A. Kumar · 2024
Closest in time.
Understanding the performance gap between online and offline alignment algorithms, 2024
Y. Tang, D. Z. Guo, Z. Zheng, D. Calandriello, Y. Cao, E. Tarassov, R. Munos, B. Ávila Pires, M. Valko, Y. Cheng, and W. Dabney · 2024
Closest in time.
Gemma: Open Models Based on Gemini Research and Technology, 2024
G. Team · 2024
Closest in time.
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, 2024
P. Wang, L. Li, Z. Shao, R. X. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui · 2024
Closest in time.
Advancing llm reasoning generalists with preference trees, 2024
L. Yuan, G. Cui, H. Wang, N. Ding, X. Wang, J. Deng, B. Shan, H. Chen, R. Xie, Y. Lin, Z. Liu, B. Zhou, H. Peng, Z. Liu, and M. Sun · 2024
Closest in time.
(InThe)WildChat: 570K ChatGPT Interaction Logs In The Wild
W. Zhao, X. Ren, J. Hessel, C. Cardie, Y. Choi, and Y. Deng · 2024
Closest in time.