Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) emerges as a promising paradigm for aligning large language models (LLMs).
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Random forests
Breiman, L · 2001
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Naeini, M. P., Cooper, G., and Hauskrecht, M · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Janner, M., Fu, J., Zhang, M., and Levine, S · 2019
Earlier work this paper cites.
Self-supervised exploration via disagreement
Pathak, D., Gandhi, D., and Gupta, A · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X., Kumar, A., Zhang, G., and Levine, S · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T · 2020
Cited alongside, same era.
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information
Ethayarajh, K., Choi, Y., and Swayamdipta, S · 2022
Cited alongside, same era.
Uncertainty estimation for language reward models
Gleave, A. and Irving, G · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2022
Cited alongside, same era.
On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting
Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M · 2022
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Closest in time.
A survey of uncertainty in deep neural networks
Gawlikowski, J., Tassi, C. R. N., Ali, M., Lee, J., Humt, M., Feng, J., Kruspe, A., Triebel, R., Jung, P., Roscher, R., et al · 2023
Closest in time.
Aligning language models with preferences through f-divergence minimization
Go, D., Korbak, T., Kruszewski, G., Rozen, J., Ryu, N., and Dymetman, M · 2023
Closest in time.
Reinforced self-training (rest) for language modeling
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al · 2023
Closest in time.
Confronting reward model overoptimization with constrained rlhf
Moskovitz, T., Singh, A. K., Strouse, D., Sandholm, T., Salakhutdinov, R., Dragan, A. D., and McAleer, S · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kreps, S., McCain, R. M., and Brundage, M · 2022
Cited alongside, same era.
A review of uncertainty for deep reinforcement learning
Lockwood, O. and Si, M · 2022
Cited alongside, same era.
Revisiting design choices in offline model based reinforcement learning
Lu, C., Ball, P., Parker-Holder, J., Osborne, M., and Roberts, S. J · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Cited alongside, same era.
Quantifying uncertainty in foundation models via ensembles
Sun, M., Yan, W., Abbeel, P., and Mordatch, I · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Closest in time.
Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization
Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y · 2023
Closest in time.
Offline rl for natural language generation with implicit language q learning
Snell, C. V., Kostrikov, I., Su, Y., Yang, S., and Levine, S · 2023
Closest in time.
Preference ranking optimization for human alignment
Song, F., Yu, B., Li, M., Yu, H., Huang, F., Li, Y., and Wang, H · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Lora ensembles for large language model fine-tuning
Wang, X., Aitchison, L., and Rudolph, M · 2023
Closest in time.
Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment
Wu, T., Zhu, B., Zhang, R., Wen, Z., Ramchandran, K., and Jiao, J · 2023
Closest in time.
Baichuan 2: Open large-scale language models
Yang, A., Xiao, B., Wang, B., Zhang, B., Bian, C., Yin, C., Lv, C., Pan, D., Wang, D., Yan, D., et al · 2023
Closest in time.
Deepspeed-chat: Easy, fast and affordable rlhf training of chatgpt-like models at all scales
Yao, Z., Aminabadi, R. Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A. A., Rasley, J., Zhang, M., Li, C., Holmes, C., et al · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Closest in time.
Fine-tuning language models with advantage-induced policy alignment
Zhu, B., Sharma, H., Frujeri, F. V., Dong, S., Zhu, C., Jordan, M. I., and Jiao, J · 2023
Closest in time.