Fetching the paper…
Reading the bibliography…
While direct policy optimization methods exist, pioneering LLMs are fine-tuned with reinforcement learning from human feedback (RLHF) to generate better responses under the supervision of a reward model learned from preference data.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Applied regression analysis
Draper, N · 1998
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Ape210k: A large-scale and template-rich dataset of math word problems
Zhao, W., Shang, M., Liu, Y., Wang, L., and Liu, J · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Defining and characterizing reward gaming
Skalse, J., Howe, N., Krasheninnikov, D., and Krueger, D · 2022
Earlier work this paper cites.
Introducing claude, 2023
Anthropic, A · 2023
Earlier work this paper cites.
Reward model ensembles help mitigate overoptimization
Coste, T., Anwar, U., Kirk, R., and Krueger, D · 2023
Earlier work this paper cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Earlier work this paper cites.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Cited alongside, same era.
The history and risks of reinforcement learning and human feedback
Lambert, N., Krendl Gilbert, T., and Zick, T · 2023
Cited alongside, same era.
Statistical rejection sampling improves preference optimization
Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J · 2023
Cited alongside, same era.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Cited alongside, same era.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D · 2023
Cited alongside, same era.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y., et al · 2024
Closest in time.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Hu, J., Wu, X., Wang, W., Zhang, D., Cao, Y., et al · 2024
Closest in time.
Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback
Ivison, H., Wang, Y., Liu, J., Wu, Z., Pyatkin, V., Lambert, N., Smith, N. A., Choi, Y., and Hajishirzi, H · 2024
Closest in time.
Calibrated language models must hallucinate
Kalai, A. T. and Vempala, S. S · 2024
Closest in time.
Smaug: Fixing failure modes of preference optimisation with dpo-positive
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the fragility of learned reward functions
McKinney, L., Duan, Y., Krueger, D., and Gleave, A · 2023
Cited alongside, same era.
Confronting reward model overoptimization with constrained rlhf
Moskovitz, T., Singh, A. K., Strouse, D., Sandholm, T., Salakhutdinov, R., Dragan, A. D., and McAleer, S · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Failure modes of learning reward models for llms and other sequence models
Pitis, S · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Cited alongside, same era.
Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment
Wu, T., Zhu, B., Zhang, R., Wen, Z., Ramchandran, K., and Jiao, J · 2023
Cited alongside, same era.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Cited alongside, same era.
Pal, A., Karkhanis, D., Dooley, S., Roberts, M., Naidu, S., and White, C · 2024
Closest in time.
Iterative reasoning preference optimization
Pang, R. Y., Yuan, W., Cho, K., He, H., Sukhbaatar, S., and Weston, J · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
Bond: Aligning llms with best-of-n distillation
Sessa, P. G., Dadashi, R., Hussenot, L., Ferret, J., Vieillard, N., Ramé, A., Shariari, B., Perrin, S., Friesen, A., Cideron, G., et al · 2024
Closest in time.
Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold
Setlur, A., Garg, S., Geng, X., Garg, N., Smith, V., and Kumar, A · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., and Guo, D · 2024
Closest in time.
Advanced tricks for training large language models with proximal policy optimization
Shen, W., Hu, J., Zhao, P., He, X., and Chen, L · 2024
Closest in time.
Preference ranking optimization for human alignment
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H · 2024
Closest in time.
Introducing qwen1.5, February 2024
Team, Q · 2024
Closest in time.
Secrets of rlhf in large language models part ii: Reward modeling
Wang, B., Zheng, R., Chen, L., Liu, Y., Dou, S., Huang, C., Shen, W., Jin, S., Zhou, E., Shi, C., et al · 2024
Closest in time.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., et al · 2024
Closest in time.
Improving reinforcement learning from human feedback with efficient reward model ensemble
Zhang, S., Chen, Z., Chen, S., Shen, Y., Sun, Z., and Gan, C · 2024
Closest in time.
Leveraging web-crawled data for high-quality fine-tuning
Zhou, J., Jiang, C., Shen, W., Zhou, X., and He, X · 2024
Closest in time.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., Gao, H., Ma, S., et al · 2024
Closest in time.