Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has demonstrated great promise in aligning large language models (LLMs) with human preference.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019) · 1909
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E. (1952) · 1952
Earlier work this paper cites.
A new family of optimal adaptive controllers for markov chains
Kumar, P. and Becker, A. (1982) · 1982
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L., Robbins, H., et al. (1985) · 1985
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
Boltzmann exploration done right
Cesa-Bianchi, N., Gentile, C., Lugosi, G., and Neu, G. (2017) · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. (2018) · 2018
Earlier work this paper cites.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020) · 2020
Earlier work this paper cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2020) · 2020
Earlier work this paper cites.
Exploration through reward biasing: Reward-biased maximum likelihood estimation for stochastic multi-armed bandits
Liu, X., Hsieh, P.-C., Hung, Y. H., Bhattacharya, A., and Kumar, P. (2020) · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020) · 2020
Earlier work this paper cites.
A survey of uncertainty in deep neural networks
Gawlikowski, J., Tassi, C. R. N., Ali, M., Lee, J., Humt, M., Feng, J., Kruspe, A., Triebel, R., Jung, P., Roscher, R., et al. (2021) · 2021
Earlier work this paper cites.
Reward-biased maximum likelihood estimation for linear stochastic bandits
Hung, Y.-H., Hsieh, P.-C., Liu, X., and Kumar, P. (2021) · 2021
Earlier work this paper cites.
Reward biased maximum likelihood estimation for reinforcement learning
Mete, A., Singh, R., Liu, X., and Kumar, P. (2021) · 2021
Earlier work this paper cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M. and Sun, W. (2021) · 2021
Cited alongside, same era.
On the optimality of batch policy optimization algorithms
Xiao, C., Wu, Y., Mei, J., Dai, B., Lattimore, T., Li, L., Szepesvari, C., and Schuurmans, D. (2021) · 2021
Cited alongside, same era.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A. (2021) · 2021
Cited alongside, same era.
Fast global convergence of natural policy gradient methods with entropy regularization
Cen, S., Cheng, C., Chen, Y., Wei, Y., and Chi, Y. (2022) · 2022
Cited alongside, same era.
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. (2022) · 2022
Slic-hf: Sequence likelihood calibration with human feedback
Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J. (2023) · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Zhu, B., Jordan, M., and Jiao, J. (2023) · 2023
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Guo, Z. D., Piot, B., Munos, R., Rowland, M., Valko, M., and Calandriello, D. (2024) · 2024
Closest in time.
Dataset reset policy optimization for RLHF
Chang, J. D., Shan, W., Oertell, O., Brantley, K., Misra, D., Lee, J. D., and Sun, W. (2024) · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models
Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The power of exploiter: Provable multi-agent rl in large state spaces
Jin, C., Liu, Q., and Yu, T. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Cited alongside, same era.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2022) · 2022
Cited alongside, same era.
Pessimistic Q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L., Li, G., Wei, Y., Chen, Y., and Chi, Y. (2022) · 2022
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. (2023) · 2023
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M. (2023) · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J. (2023) · 2023
Cited alongside, same era.
Closest in time.
KTO: Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D. (2024) · 2024
Closest in time.
Direct language model alignment from online ai feedback
Guo, S., Zhang, B., Liu, T., Liu, T., Khalman, M., Llinares, F., Rame, A., Mesnard, T., Zhao, Y., Piot, B., et al. (2024) · 2024
Closest in time.
Smaug: Fixing failure modes of preference optimisation with dpo-positive
Pal, A., Karkhanis, D., Dooley, S., Roberts, M., Naidu, S., and White, C. (2024) · 2024
Closest in time.
Iterative reasoning preference optimization
Pang, R. Y., Yuan, W., Cho, K., He, H., Sukhbaatar, S., and Weston, J. (2024) · 2024
Closest in time.
From r r to q ⋆ q^{\star} : Your language model is secretly a Q-function
Rafailov, R., Hejna, J., Park, R., and Finn, C. (2024) · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C., Cheng, C.-A., Mitra, A., Santacroce, M., Awadallah, A., and Xie, T. (2024) · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Swamy, G., Dann, C., Kidambi, R., Wu, Z. S., and Agarwal, A. (2024) · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P. H., Valko, M., Pires, B. Á., and Piot, B. (2024) · 2024
Closest in time.
Self-play preference optimization for language model alignment
Wu, Y., Sun, Z., Yuan, H., Ji, K., Yang, Y., and Gu, Q. (2024) · 2024
Closest in time.
Exploratory preference optimization: Harnessing implicit q*-approximation for sample-efficient rlhf
Xie, T., Foster, D. J., Krishnamurthy, A., Rosset, C., Awadallah, A., and Rakhlin, A. (2024) · 2024
Closest in time.
Faster WIND: Accelerating iterative best-of- n n distillation for LLM alignment
Yang, T., Mei, J., Dai, H., Wen, Z., Cen, S., Schuurmans, D., Chi, Y., and Dai, B. (2024) · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. (2024) · 2024
Closest in time.
Dpo meets ppo: Reinforced token optimization for rlhf
Zhong, H., Feng, G., Xiong, W., Zhao, L., He, D., Bian, J., and Wang, L. (2024) · 2024
Closest in time.
Incentivize without bonus: Provably efficient model-based online multi-agent RL for Markov games
Yang, T., Dai, B., Xiao, L., and Chi, Y. (2025) · 2025
Closest in time.