Fetching the paper…
Reading the bibliography…
Modern large language models (LLMs) are optimized for human-aligned responses using Reinforcement Learning from Human Feedback (RLHF).
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Logistic regression
Kleinbaum, D. G., Dietz, K., Gail, M., Klein, M., and Klein, M · 2002
Earlier work this paper cites.
Self-concordant analysis for logistic regression
Bach, F · 2010
Earlier work this paper cites.
The structure of musical preferences: a five-factor model
Rentfrow, P. J., Goldberg, L. R., and Levitin, D. J · 2011
Earlier work this paper cites.
Numerical methods for scientists and engineers
Hamming, R · 2012
Earlier work this paper cites.
Logistic matrix factorization for implicit feedback data
Johnson, C. C. et al · 2014
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Improved optimistic algorithms for logistic bandits
Faury, L., Abeille, M., Calauzènes, C., and Fercoq, O · 2020
Earlier work this paper cites.
Evaluating approaches to personalizing language models
King, M. and Cook, P · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
A survey of deep active learning
Ren, P., Xiao, Y., Chang, X., Huang, P.-Y., Li, Z., Gupta, B. B., Chen, X., and Wang, X · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Cited alongside, same era.
Personalized soups: Personalized large language model alignment via post-hoc parameter merging
Jang, J., Kim, S., Lin, B. Y., Wang, Y., Hessel, J., Zettlemoyer, L., Hajishirzi, H., Choi, Y., and Ammanabrolu, P · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Quantile regression for distributional reward models in rlhf
Dorka, N · 2024
Later among the works it cites.
Scaling synthetic data creation with 1,000,000,000 personas
Ge, T., Chan, X., Wang, X., Yu, D., Mi, H., and Yu, D · 2024
Later among the works it cites.
Controllable preference optimization: Toward controllable multi-objective alignment
Guo, Y., Cui, G., Yuan, L., Ding, N., Sun, Z., Sun, B., Chen, H., Xie, R., Zhou, J., Lin, Y., et al · 2024
Later among the works it cites.
Value augmented sampling for language model alignment and personalization
Han, S., Shenfeld, I., Srivastava, A., Kim, Y., and Agrawal, P · 2024
Later among the works it cites.
Reinforcement learning from human feedback with active queries
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Controlled decoding from language models
Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., et al · 2023
Cited alongside, same era.
Integrating summarization and retrieval for enhanced personalization via large language models
Richardson, C., Zhang, Y., Gillespie, K., Kar, S., Singh, A., Raeesy, Z., Khan, O. Z., and Sethy, A · 2023
Cited alongside, same era.
Dueling rl: Reinforcement learning with trajectory preferences
Saha, A., Pacchiano, A., and Lee, J · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Cited alongside, same era.
Beyond one-preference-for-all: Multi-objective direct preference optimization
Zhou, Z., Liu, J., Yang, C., Shao, J., Liu, Y., Yue, X., Ouyang, W., and Qiao, Y · 2023
Cited alongside, same era.
Active preference optimization for sample efficient rlhf
Das, N., Chakraborty, S., Pacchiano, A., and Chowdhury, S. R · 2024
Cited alongside, same era.
Can llm be a personalized judge?
Dong, Y. R., Hu, T., and Collier, N · 2024
Cited alongside, same era.
Ji, K., He, J., and Gu, Q · 2024
Later among the works it cites.
Args: Alignment as reward-guided search
Khanov, M., Burapacheep, J., and Li, Y · 2024
Later among the works it cites.
User-llm: Efficient llm contextualization with user embeddings
Ning, L., Liu, L., Wu, J., Wu, N., Berlowitz, D., Prakash, S., Green, B., O’Banion, S., and Xie, J · 2024
Later among the works it cites.
Personalizing reinforcement learning from human feedback with variational preference learning
Poddar, S., Wan, Y., Ivison, H., Gupta, A., and Jaques, N · 2024
Later among the works it cites.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Rame, A., Couairon, G., Dancette, C., Gaya, J.-B., Shukor, M., Soulier, L., and Cord, M · 2024
Later among the works it cites.
A roadmap to pluralistic alignment
Sorensen, T., Moore, J., Fisher, J., Gordon, M., Mireshghallah, N., Rytting, C. M., Ye, A., Jiang, L., Lu, X., Dziri, N., et al · 2024
Later among the works it cites.
Understanding the role of user profile in the personalization of large language models
Wu, B., Shi, Z., Rahmani, H. A., Ramineni, V., and Yilmaz, E · 2024
Later among the works it cites.
Personalization of large language models: A survey
Zhang, Z., Rossi, R. A., Kveton, B., Shao, Y., Yang, D., Zamani, H., Dernoncourt, F., Barrow, J., Yu, T., Kim, S., et al · 2024
Later among the works it cites.