Fetching the paper…
Reading the bibliography…
RL-based post-training of language models is almost exclusively done using on-policy methods such as PPO.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 1909
Earlier work this paper cites.
General duality between optimal control and estimation
Todorov, E · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., Dey, A. K., et al · 2008
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Kappen, B., Gomez, V., and Opper, M · 2009
Earlier work this paper cites.
Sequence tutor: Conservative fine-tuning of sequence generation models with KL-control
Jaques, N., Gu, S., Bahdanau, D., Hernández-Lobato, J. M., Turner, R. E., and Eck, D · 2016
Earlier work this paper cites.
High-Dimensional Continuous Control Using Generalized Advantage Estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Earlier work this paper cites.
Calculus on MDPs: Potential Shaping as a Gradient
Jenner, E., van Hoof, H., and Gleave, A · 2022
Cited alongside, same era.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L. E., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. J · 2022
Cited alongside, same era.
Q-transformer: Scalable offline reinforcement learning via autoregressive Q-functions
Chebotar, Y., Vuong, Q., Irpan, A., Hausman, K., Xia, F., Lu, Y., Kumar, A., Yu, T., Herzog, A., Pertsch, K., Gopalakrishnan, K., Ibarz, J., Nachum, O., Sontakke, S., Salazar, G., Tran, H., Peralta, J., Tan, C., Manjunath, D., Singht, J., Zitkovich, B., Jackson, T., Rao, K., Finn, C., and Levine, S · 2023
Cited alongside, same era.
Rlef: Grounding code llms in execution feedback with reinforcement learning
Gehring, J., Zheng, K., Copet, J., Mella, V., Cohen, T., and Synnaeve, G · 2024
Later among the works it cites.
VinePPO: Unlocking RL potential for LLM reasoning through refined credit assignment
Kazemnejad, A., Aghajohari, M., Portelance, E., Sordoni, A., Reddy, S., Courville, A., and Roux, N. L · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Llama Team, AI Meta · 2024
Later among the works it cites.
What type of inference is planning?
Lázaro-Gredilla, M., Ku, L. Y., Murphy, K. P., and George, D · 2024
Later among the works it cites.
Asynchronous RLHF: Faster and more efficient off-policy RL for language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, R., Fu, J., Zhang, B.-W., Huang, T., Sun, Z., Lyu, C., Liu, G., Jin, Z., and Li, G · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Cited alongside, same era.
VA-learning as a more efficient alternative to Q-learning
Tang, Y., Munos, R., Rowland, M., and Valko, M · 2023
Cited alongside, same era.
Abdin, M., Aneja, J., Behl, H., Bubeck, S., Eldan, R., Gunasekar, S., Harrison, M., Hewett, R. J., Javaheripi, M., Kauffmann, P., Lee, J. R., Lee, Y. T., Li, Y., Liu, W., Mendes, C. C. T., Nguyen, A., Price, E., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Wang, X., Ward, R., Wu, Y., Yu, D., Zhang, C., and Zhang, Y · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
Process Reinforcement through IMplicit rEwards
Cui, G., Yuan, L., Wang, Z., Wang, H., Li, W., He, B., Fan, Y., Yu, T., Xu, Q., Chen, W., Yuan, J., Chen, H., Zhang, K., Lv, X., Wang, S., Yao, Y., Han, X., Peng, H., Cheng, Y., Liu, Z., Sun, M., Zhou, B., and Ding, N
Cited in the paper.
Equivalence between policy gradients and soft Q-learning
Schulman, J., Chen, X., and Abbeel, P
Cited in the paper.
Proximal Policy Optimization Algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O
Cited in the paper.
Noukhovitch, M., Huang, S., Xhonneux, S., Hosseini, A., Agarwal, R., and Courville, A · 2024
Later among the works it cites.
Offline regularised reinforcement learning for large language models alignment
Richemond, P. H., Tang, Y., Guo, D., Calandriello, D., Azar, M. G., Rafailov, R., Pires, B. A., Tarassov, E., Spangher, L., Ellsworth, W., Severyn, A., Mallinson, J., Shani, L., Shamir, G., Joshi, R., Liu, T., Munos, R., and Piot, B · 2024
Later among the works it cites.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y. K., Wu, Y., and Guo, D · 2024
Later among the works it cites.
Wu, Y., Sun, Z., Li, S., Welleck, S., and Yang, Y · 2024
Later among the works it cites.
Free process rewards without process labels
Yuan, L., Li, W., Chen, H., Cui, G., Ding, N., Zhang, K., Zhou, B., Liu, Z., and Peng, H · 2024
Later among the works it cites.