Fetching the paper…
Reading the bibliography…
The growing safety concerns surrounding large language models raise an urgent need to align them with diverse human preferences to simultaneously enhance their helpfulness and safety.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Applications of characteristic functions
E. Lukacs and R. G. Laha · 1964
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time. iv
M. D. Donsker and S. S. Varadhan · 1983
Earlier work this paper cites.
From graph to manifold laplacian: The convergence rate
A. Singer · 2006
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
S. Bubeck et al · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Earlier work this paper cites.
Nonlinear programming
D. P. Bertsekas · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Challenges of real-world reinforcement learning
G. Dulac-Arnold, D. Mankowitz, and T. Hester · 2019
Earlier work this paper cites.
Negative momentum for improved game dynamics
G. Gidel, R. A. Hemmat, M. Pezeshki, R. Le Priol, G. Huang, S. Lacoste-Julien, and I. Mitliagkas · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Constrained Markov decision processes
E. Altman · 2021
Earlier work this paper cites.
Extracting training data from large language models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Safe policies for reinforcement learning via primal-dual methods
S. Paternain, M. Calvo-Fullana, L. F. Chamon, and A. Ribeiro · 2022
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, Z. D. Guo, B. Piot, R. Munos, M. Rowland, M. Valko, and D. Calandriello · 2024
Closest in time.
MaxMin-RLHF: Alignment with diverse human preferences
S. Chakraborty, J. Qiu, H. Yuan, A. Koppel, D. Manocha, F. Huang, A. Bedi, and M. Wang · 2024
Closest in time.
Dataset reset policy optimization for RLHF
J. D. Chang, W. Shan, O. Oertell, K. Brantley, D. Misra, J. D. Lee, and W. Sun · 2024
Closest in time.
Safe RLHF: Safe reinforcement learning from human feedback
J. Dai, X. Pan, R. Sun, J. Ji, X. Xu, M. Liu, Y. Wang, and Y. Yang · 2024
Closest in time.
Uncertainty in language models: Assessment through rank-calibration
X. Huang, S. Li, M. Yu, M. Sesia, H. Hassani, I. Lee, O. Bastani, and E. Dobriban · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
PAL: Program-aided language models
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models, 2023
X. Li, T. Zhang, Y. Dubois, R. Taori, I. Gulrajani, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Cited alongside, same era.
ReLOAD: Reinforcement learning with optimistic ascent-descent for last-iterate convergence in constrained MDPs
T. Moskovitz, B. O’Donoghue, V. Veeriah, S. Flennerhag, S. Singh, and T. Zahavy · 2023
Cited alongside, same era.
LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action
D. Shah, B. Osiński, S. Levine, et al · 2023
Cited alongside, same era.
Prompting large language model for machine translation: A case study
B. Zhang, B. Haddow, and A. Birch · 2023
Cited alongside, same era.
Beyond one-preference-for-all: Multi-objective direct preference optimization
Z. Zhou, J. Liu, C. Yang, J. Shao, Y. Liu, X. Yue, W. Ouyang, and Y. Qiao · 2023
Cited alongside, same era.
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang · 2024
Closest in time.
Enhancing LLM safety via constrained direct preference optimization
Z. Liu, X. Sun, and Z. Zheng · 2024
Closest in time.
Confronting reward model overoptimization with constrained RLHF
T. Moskovitz, A. K. Singh, D. Strouse, T. Sandholm, R. Salakhutdinov, A. Dragan, and S. M. McAleer · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Rewarded soups: Towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
A. Rame, G. Couairon, C. Dancette, J.-B. Gaya, M. Shukor, L. Soulier, and M. Cord · 2024
Closest in time.
Stepwise alignment for constrained language model policy optimization
A. Wachi, T. Q. Tran, R. Sato, T. Tanabe, and Y. Akimoto · 2024
Closest in time.
Iterative preference learning from human feedback: Bridging theory and practice for RLHF under KL-constraint
W. Xiong, H. Dong, C. Ye, Z. Wang, H. Zhong, H. Ji, N. Jiang, and T. Zhang · 2024
Closest in time.
Rewards-in-Context: Multi-objective alignment of foundation models with dynamic preference adjustment
R. Yang, X. Pan, F. Luo, S. Qiu, H. Zhong, D. Yu, and J. Chen · 2024
Closest in time.