Fetching the paper…
Reading the bibliography…
Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data.
“Rank analysis of incomplete block designs: I. The method of paired comparisons”
Ralph Bradley and Milton Terry · 1952
Earlier work this paper cites.
“On general minimax theorems.”, 1958
Maurice Sion · 1958
Earlier work this paper cites.
“Robust stochastic approximation approach to stochastic programming”
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan and Alexander Shapiro · 2009
Earlier work this paper cites.
“Stochastic gradient methods for distributionally robust optimization with f-divergences”
Hongseok Namkoong and John Duchi · 2016
Earlier work this paper cites.
“Deep reinforcement learning from human preferences”
Paul Christiano et al · 2017
Earlier work this paper cites.
“Proximal policy optimization algorithms”
John Schulman et al · 2017
Earlier work this paper cites.
“Decoupled Weight Decay Regularization”
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
“Distributionally robust language modeling”
Yonatan Oren, Shiori Sagawa, Tatsunori Hashimoto and Percy Liang · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
Shiori Sagawa, Pang Koh, Tatsunori Hashimoto and Percy Liang · 2019
Earlier work this paper cites.
“Fine-tuning language models from human preferences”
Daniel Ziegler et al · 2019
Earlier work this paper cites.
“Learning to summarize with human feedback”
Nisan Stiennon et al · 2020
Earlier work this paper cites.
“Decision transformer: Reinforcement learning via sequence modeling”
Lili Chen et al · 2021
Earlier work this paper cites.
“Webgpt: Browser-assisted question-answering with human feedback”
Reiichiro Nakano et al · 2021
Earlier work this paper cites.
“Constitutional ai: Harmlessness from ai feedback”
Yuntao Bai et al · 2022
Earlier work this paper cites.
“Training a helpful and harmless assistant with reinforcement learning from human feedback”
Yuntao Bai et al · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Earlier work this paper cites.
“A general theoretical paradigm to understand learning from human preferences”
Mohammad Azar et al · 2023
Earlier work this paper cites.
“ULMA: Unified Language Model Alignment with Demonstration and Point-wise Human Preference”
Tianchi Cai et al · 2023
Earlier work this paper cites.
“Open problems and fundamental limitations of reinforcement learning from human feedback”
Stephen Casper et al · 2023
Cited alongside, same era.
“Raft: Reward ranked finetuning for generative foundation model alignment”
Hanze Dong et al · 2023
Cited alongside, same era.
“Towards measuring the representation of subjective global opinions in language models”
Esin Durmus et al · 2023
Cited alongside, same era.
“Scaling laws for reward model overoptimization”
Leo Gao, John Schulman and Jacob Hilton · 2023
Cited alongside, same era.
“Reinforced self-training (rest) for language modeling”
Caglar Gulcehre et al · 2023
Cited alongside, same era.
“Slic-hf: Sequence likelihood calibration with human feedback”
Yao Zhao et al · 2023
Later among the works it cites.
“MaxMin-RLHF: Towards Equitable Alignment of Large Language Models with Diverse Human Preferences”
Souradip Chakraborty et al · 2024
Closest in time.
“Self-play fine-tuning converts weak language models to strong language models”
Zixiang Chen et al · 2024
Closest in time.
“Provably Sample Efficient RLHF via Active Preference Optimization”
Nirjhar Das, Souradip Chakraborty, Aldo Pacchiano and Sayak Chowdhury · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Aligning language models with offline reinforcement learning from human feedback”
Jian Hu, Li Tao, June Yang and Chandler Zhou · 2023
Cited alongside, same era.
“A survey of reinforcement learning from human feedback”
Timo Kaufmann, Paul Weng, Viktor Bengs and Eyke Hüllermeier · 2023
Cited alongside, same era.
“Entangled preferences: The history and risks of reinforcement learning and human feedback”
Nathan Lambert, Thomas Gilbert and Tom Zick · 2023
Cited alongside, same era.
“Rlaif: Scaling reinforcement learning from human feedback with ai feedback”
Harrison Lee et al · 2023
Cited alongside, same era.
“Policy Optimization in RLHF: The Impact of Out-of-preference Data”
Ziniu Li, Tian Xu and Yang Yu · 2023
Cited alongside, same era.
“Sample Efficient Reinforcement Learning from Human Feedback via Active Exploration”, 2023
Viraj Mehta et al · 2023
Cited alongside, same era.
“Nash learning from human feedback”
Rémi Munos et al · 2023
Cited alongside, same era.
Kawin Ethayarajh et al · 2024
Closest in time.
“Reference-free monolithic preference optimization with odds ratio”
Jiwoo Hong, Noah Lee and James Thorne · 2024
Closest in time.
“Understanding the Learning Dynamics of Alignment with Human Feedback”
Shawn Im and Yixuan Li · 2024
Closest in time.
“Reinforcement Learning from Human Feedback with Active Queries”
Kaixuan Ji, Jiafan He and Quanquan Gu · 2024
Closest in time.
“Direct nash optimization: Teaching language models to self-improve with general preferences”
Corby Rosset et al · 2024
Closest in time.
“Preference ranking optimization for human alignment”
Feifan Song et al · 2024
Closest in time.
“A minimaximalist approach to reinforcement learning from human feedback”
Gokul Swamy et al · 2024
Closest in time.
“Generalized Preference Optimization: A Unified Approach to Offline Alignment”
Yunhao Tang et al · 2024
Closest in time.
“Gemma: Open models based on gemini research and technology”
Gemma Team et al · 2024
Closest in time.
Haoxiang Wang et al · 2024
Closest in time.
“Self-Play Preference Optimization for Language Model Alignment”
Yue Wu et al · 2024
Closest in time.
Rui Yang et al · 2024
Closest in time.
Chenlu Ye et al · 2024
Closest in time.
“Self-rewarding language models”
Weizhe Yuan et al · 2024
Closest in time.
“DPO Meets PPO: Reinforced Token Optimization for RLHF”
Han Zhong et al · 2024
Closest in time.