Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) has shown promise in aligning large language models (LLMs).
Fine-tuning language models from human preferences, 2020
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
On general minimax theorems
Sion, M · 1958
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
Polyak, B · 1963
Earlier work this paper cites.
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks, 2021
Liu, C., Zhu, L., and Belkin, M · 2003
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Learning to summarize from human feedback, 2022
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Understanding learned reward functions, 2020
Michaud, E. J., Gleave, A., and Russell, S · 2012
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Collective Choice and Social Welfare: An Expanded Edition
Sen, A · 2017
Earlier work this paper cites.
Kreutzer, J., Uyheng, J., and Riezler, S · 2018
Earlier work this paper cites.
Biased stochastic first-order methods for conditional stochastic optimization and applications in meta learning
Hu, Y., Zhang, S., Chen, X., and He, N · 2020
Earlier work this paper cites.
Karimi, H., Nutini, J., and Schmidt, M · 2020
Earlier work this paper cites.
Deep reinforcement learning for multiobjective optimization
Li, K., Zhang, T., and Wang, R · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Cited alongside, same era.
Trl: Transformer reinforcement learning
von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., Huang, S., Rasul, K., and Gallouédec, Q · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Learning optimal controllers by policy gradient: Global optimality via convex parameterization
Sun, Y. and Fazel, M · 2021
Cited alongside, same era.
Aligning text-to-image models using human feedback, 2023
Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S · 2023
Later among the works it cites.
Ramé, A., Couairon, G., Shukor, M., Dancette, C., Gaya, J.-B., Soulier, L., and Cord, M · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Later among the works it cites.
Maxmin-RLHF: Alignment with diverse human preferences
Chakraborty, S., Qiu, J., Yuan, H., Koppel, A., Manocha, D., Huang, F., Bedi, A., and Wang, M · 2024
Later among the works it cites.
Direct preference optimization with unobserved preference heterogeneity, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Cited alongside, same era.
Fine-tuning language models to find agreement among humans with diverse preferences, 2022
Bakker, M. A., Chadwick, M. J., Sheahan, H. R., Tessler, M. H., Campbell-Gillingham, L., Balaguer, J., McAleese, N., Glaese, A., Aslanides, J., Botvinick, M. M., and Summerfield, C · 2022
Cited alongside, same era.
Timelms: Diachronic language models from twitter, 2022
Loureiro, D., Barbieri, F., Neves, L., Anke, L. E., and Camacho-Collados, J · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
The effects of reward misspecification: Mapping and mitigating misaligned models, 2022
Pan, A., Bhatia, K., and Steinhardt, J · 2022
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback, 2023
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Bıyık, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D · 2023
Cited alongside, same era.
Deep reinforcement learning from human preferences, 2023
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2023
Cited alongside, same era.
Chidambaram, K., Seetharaman, K. V., and Syrgkanis, V · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Multi-level monte-carlo gradient methods for stochastic optimization with biased oracles
Hu, Y., Wang, J., Chen, X., and He, N · 2024
Later among the works it cites.
Statistical rejection sampling improves preference optimization, 2024
Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J · 2024
Later among the works it cites.
Rlhf from heterogeneous feedback via personalization and preference aggregation, 2024
Park, C., Liu, M., Kong, D., Zhang, K., and Ozdaglar, A · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model, 2024
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2024
Later among the works it cites.
Decoding-time language model alignment with multiple objectives, 2024
Shi, R., Chen, Y., Hu, Y., Liu, A., Hajishirzi, H., Smith, N. A., and Du, S. S · 2024
Later among the works it cites.
Wang, H., Lin, Y., Xiong, W., Yang, R., Diao, S., Qiu, S., Zhao, H., and Zhang, T · 2024
Later among the works it cites.
Yang, R., Pan, X., Luo, F., Qiu, S., Zhong, H., Yu, D., and Chen, J · 2024
Later among the works it cites.
Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization, 2024
Zhou, Z., Liu, J., Shao, J., Yue, X., Yang, C., Ouyang, W., and Qiao, Y · 2024
Later among the works it cites.