Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) is the mainstream paradigm used to align large language models (LLMs) with human preferences.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Learning with noisy labels
Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A · 2013
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Weaver, L. and Tao, N · 2013
Earlier work this paper cites.
Classification with noisy labels by importance reweighting
Liu, T. and Tao, D · 2015
Earlier work this paper cites.
Concrete problems in ai safety, 2016
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog, 2019
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations, 2020
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models, 2020
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2020
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment, 2021
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., and Kaplan, J · 2021
Earlier work this paper cites.
Alignment of language agents, 2021
Kenton, Z., Everitt, T., Weidinger, L., Gabriel, I., Mikulik, V., and Irving, G · 2021
Earlier work this paper cites.
Policy learning using weak supervision
Wang, J., Guo, H., Zhu, Z., and Liu, Y · 2021
Earlier work this paper cites.
Clusterability as an alternative to anchor points when learning with noisy labels
Zhu, Z., Song, Y., and Liu, Y · 2021
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J · 2022
Cited alongside, same era.
Scaling laws for reward model overoptimization, 2022
Gao, L., Schulman, J., and Hilton, J · 2022
Cited alongside, same era.
Accelerate: Training and inference at scale made simple, efficient and adaptable
Gugger, S., Debut, L., Wolf, T., Schmid, P., Mueller, Z., Mangrulkar, S., Sun, M., and Bossan, B · 2022
Cited alongside, same era.
Rl with kl penalties is better viewed as bayesian inference, 2022
Korbak, T., Perez, E., and Buckley, C. L · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Contrastive prefence learning: Learning from human feedback without rl
Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D · 2023
Later among the works it cites.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion, 2023
Jiang, D., Ren, X., and Lin, B. Y · 2023
Later among the works it cites.
Beyond reward: Offline preference-guided policy optimization, 2023
Kang, Y., Shi, D., Liu, J., He, L., and Wang, D · 2023
Later among the works it cites.
The history and risks of reinforcement learning and human feedback, 2023
Lambert, N., Gilbert, T. K., and Zick, T · 2023
Later among the works it cites.
Policy optimization in rlhf: The impact of out-of-preference data
Li, Z., Xu, T., and Yu, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Defining and characterizing reward hacking, 2022
Skalse, J., Howe, N. H. R., Krasheninnikov, D., and Krueger, D · 2022
Cited alongside, same era.
Learning to summarize from human feedback, 2022
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P · 2022
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R · 2023
Cited alongside, same era.
Red-teaming large language models using chain of utterances for safety-alignment, 2023
Bhardwaj, R. and Poria, S · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback, 2023
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T., Marks, S., Segerie, C.-R., Carroll, M., Peng, A., Christoffersen, P., Damani, M., Slocum, S., Anwar, U., Siththaranjan, A., Nadeau, M., Michaud, E. J., Pfau, J., Krasheninnikov, D., Chen, X., Langosco, L., Hase, P., Bıyık, E., Dragan, A., Krueger, D., Sadigh, D., and Hadfield-Menell, D · 2023
Cited alongside, same era.
Adversarial preference optimization
Cheng, P., Yang, Y., Li, J., Dai, Y., and Du, N · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Failure modes of learning reward models for llms and other sequence models
Pitis, S · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
A long way to go: Investigating length correlations in rlhf
Singhal, P., Goyal, T., Xu, J., and Durrett, G · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment, 2023
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T · 2023
Later among the works it cites.
Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment
Wu, T., Zhu, B., Zhang, R., Wen, Z., Ramchandran, K., and Jiao, J · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback, 2023
Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J · 2023
Later among the works it cites.
Zhu, Z., Wang, J., Cheng, H., and Liu, Y · 2023
Later among the works it cites.
Self-play fine-tuning converts weak language models to strong language models, 2024
Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q · 2024
Closest in time.
Statistical rejection sampling improves preference optimization, 2024
Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J · 2024
Closest in time.
Secrets of rlhf in large language models part ii: Reward modeling, 2024
Wang, B., Zheng, R., Chen, L., Liu, Y., Dou, S., Huang, C., Shen, W., Jin, S., Zhou, E., Shi, C., Gao, S., Xu, N., Zhou, Y., Fan, X., Xi, Z., Zhao, J., Wang, X., Ji, T., Yan, H., Shen, L., Chen, Z., Gui, T., Zhang, Q., Qiu, X., Huang, X., Wu, Z., and Jiang, Y.-G · 2024
Closest in time.
Self-rewarding language models, 2024
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Closest in time.