Fetching the paper…
Reading the bibliography…
The alignment of large language models (LLMs) with human values is critical as these models become increasingly integrated into various societal and decision-making processes.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Bradley RA, Terry ME · 1952
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Peters J, Schaal S · 2007
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Li Y, Liang Y · 2018
Earlier work this paper cites.
Fine-Tuning Language Models from Human Preferences
Ziegler DM, Stiennon N, Wu J, Brown TB, Radford A, Amodei D, et al · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler DM, Stiennon N, Wu J, Brown TB, Radford A, Amodei D, et al · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng XB, Kumar A, Zhang G, Levine S · 2019
Earlier work this paper cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Shin T, Razeghi Y, Logan IV RL, Wallace E, Singh S · 2020
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester B, Al-Rfou R, Constant N · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li XL, Liang P · 2021
Earlier work this paper cites.
Factual Probing Is [MASK]: Learning vs. Learning to Recall
Zhong Z, Friedman D, Chen D · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C, Mishkin P, et al · 2022
Earlier work this paper cites.
Clip-Tuning: Towards Derivative-free Prompt Learning with a Mixture of Rewards
Chai Y, Wang S, Sun Y, Tian H, Wu H, Wang H · 2022
Earlier work this paper cites.
Black-box tuning for language-model-as-a-service
Sun T, Shao Y, Qian H, Huang X, Qiu X · 2022
Earlier work this paper cites.
BBTv2: Pure Black-Box Optimization Can Be Comparable to Gradient Descent for Few-Shot Learning
Sun T, He Z, Qian H, Huang X, Qiu X · 2022
Earlier work this paper cites.
Available from: https://arxiv.org/abs/2212.09251
Perez E, Ringer S, Lukošiūtė K, Nguyen K, Chen E, Heiner S, et al.. Discovering Language Model Behaviors with Model-Written Evaluations. arXiv; 2022 · 2022
Earlier work this paper cites.
A survey of reinforcement learning from human feedback
Kaufmann T, Weng P, Bengs V, Hüllermeier E · 2023
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Dai J, Pan X, Sun R, Ji J, Xu X, Liu M, et al · 2023
Earlier work this paper cites.
Principled Reinforcement Learning with Human Feedback from Pairwise or K K -wise Comparisons
Zhu B, Jiao J, Jordan MI · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar MG, Rowland M, Piot B, Guo D, Calandriello D, Valko M, et al · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper S, Davies X, Shi C, Gilbert TK, Scheurer J, Rando J, et al · 2023
Cited alongside, same era.
Black-Box Prompt Learning for Pre-trained Language Models
Diao S, Huang Z, Xu R, Li X, Yong L, Zhou X, et al · 2023
Cited alongside, same era.
PromptAgent: Strategic planning with language models enables expert-level prompt optimization
Wang X, Li C, Wang Z, Bai F, Luo H, Zhang J, et al · 2023
Cited alongside, same era.
Slic-hf: Sequence likelihood calibration with human feedback
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois Y, Li CX, Taori R, Zhang T, Gulrajani I, Ba J, et al · 2024
Later among the works it cites.
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
Chaudhari S, Aggarwal P, Murahari V, Rajpurohit T, Kalyan A, Narasimhan K, et al · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov R, Sharma A, Mitchell E, Manning CD, Ermon S, Finn C · 2024
Later among the works it cites.
Direct Preference Optimization with an Offset
Amini A, Vieira T, Cotterell R · 2024
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Azar MG, Guo ZD, Piot B, Munos R, Rowland M, Valko M, et al · 2024
Later among the works it cites.
Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhao Y, Joshi R, Liu T, Khalman M, Saleh M, Liu PJ · 2023
Cited alongside, same era.
Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints
Wang C, Jiang Y, Yang C, Liu H, Chen Y · 2023
Cited alongside, same era.
Toward Human Readable Prompt Tuning: Kubrick’s The Shining is a good movie, and a good prompt too?
Shi W, Han X, Gonen H, Holtzman A, Tsvetkov Y, Zettlemoyer L · 2023
Cited alongside, same era.
MultiPrompter: Cooperative Prompt Optimization with Multi-Agent Reinforcement Learning
Kim DK, Sohn S, Logeswaran L, Shim D, Lee H · 2023
Cited alongside, same era.
Large Language Models Are Human-Level Prompt Engineers
Zhou Y, Muresanu AI, Han Z, Paster K, Pitis S, Chan H, et al · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov R, Sharma A, Mitchell E, Ermon S, Manning CD, Finn C · 2023
Cited alongside, same era.
Helpsteer: Multi-attribute helpfulness dataset for steerlm
Wang Z, Dong Y, Zeng J, Adams V, Sreedhar MN, Egert D, et al · 2023
Cited alongside, same era.
Gou Q, Nguyen CT · 2024
Later among the works it cites.
LiPO: Listwise Preference Optimization through Learning-to-Rank
Liu T, Qin Z, Wu J, Shen J, Khalman M, Joshi R, et al · 2024
Later among the works it cites.
Filtered Direct Preference Optimization
Morimura T, Sakamoto M, Jinnai Y, Abe K, Air K · 2024
Later among the works it cites.
Generalized Preference Optimization: A Unified Approach to Offline Alignment
Tang Y, Guo ZD, Zheng Z, Calandriello D, Munos R, Rowland M, et al · 2024
Later among the works it cites.
Efficient Exploration for LLMs
Dwaracherla V, Asghari SM, Hao B, Van Roy B · 2024
Later among the works it cites.
Available from: https://arxiv.org/abs/2403.07691
Hong J, Lee N, Thorne J. ORPO: Monolithic Preference Optimization without Reference Model; 2024 · 2024
Later among the works it cites.
Available from: https://arxiv.org/abs/2405.11870
Hua E, Qi B, Zhang K, Yu Y, Ding N, Lv X, et al.. Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process; 2024 · 2024
Later among the works it cites.
Curiosity-driven red-teaming for large language models
Hong ZW, Shenfeld I, Wang TH, Chuang YS, Pareja A, Glass J, et al · 2024
Later among the works it cites.
Gradient-based language model red teaming
Wichers N, Denison C, Beirami A · 2024
Later among the works it cites.
Learning diverse attacks on large language models for robust red-teaming and safety tuning
Lee S, Kim M, Cherif L, Dobre D, Lee J, Hwang SJ, et al · 2024
Later among the works it cites.
LIAR: Leveraging Alignment (Best-of-N) to Jailbreak LLMs in Seconds
Beetham J, Chakraborty S, Wang M, Huang F, Bedi AS, Shah M · 2024
Later among the works it cites.
ULTRAFEEDBACK: Boosting Language Models with Scaled AI Feedback
Cui G, Yuan L, Ding N, Yao G, He B, Zhu W, et al · 2024
Later among the works it cites.