Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has demonstrated remarkable effectiveness in aligning large language models (LLMs) with human preferences.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Intransitivity, utility, and the aggregation of preference patterns
May, K. O · 1954
Earlier work this paper cites.
Intransitivity of preferences
Tversky, A · 1969
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Freund, Y. and Schapire, R. E · 1999
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 2007
Earlier work this paper cites.
Near-optimal no-regret algorithms for zero-sum games
Daskalakis, C., Deckelbaum, A., and Kim, A · 2011
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Rakhlin, S. and Sridharan, K · 2013
Earlier work this paper cites.
Fast convergence of regularized learning in games
Syrgkanis, V., Agarwal, A., Luo, H., and Schapire, R. E · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Online reinforcement learning in stochastic games
Wei, C.-Y., Hong, Y.-T., and Lu, C.-J · 2017
Earlier work this paper cites.
Cycles in zero-sum differential games and biological diversity
Mai, T., Mihail, M., Panageas, I., Ratcliff, W., Vazirani, V., and Yunker, P · 2018
Earlier work this paper cites.
On the weaknesses of reinforcement learning for neural machine translation
Choshen, L., Fox, L., Aizenbud, Z., and Abend, O · 2019
Earlier work this paper cites.
Online and bandit algorithms for nonstationary stochastic saddle-point optimization
Roy, A., Chen, Y., Balasubramanian, K., and Mohapatra, P · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Near-optimal reinforcement learning with self-play
Bai, Y., Jin, C., and Yu, T · 2020
Earlier work this paper cites.
Hedging in games: Faster convergence of external and swap regrets
Chen, X. and Peng, B · 2020
Earlier work this paper cites.
Faster algorithms for extensive-form game solving via improved smoothing functions
Kroer, C., Waugh, K., Kılınç-Karzan, F., and Sandholm, T · 2020
Earlier work this paper cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C · 2020
Earlier work this paper cites.
Linear last-iterate convergence in constrained saddle-point optimization
Wei, C.-Y., Lee, C.-W., Zhang, M., and Luo, H · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
Near-optimal no-regret learning in general games
Daskalakis, C., Fishelson, M., and Golowich, N · 2021
Cited alongside, same era.
V-learning–a simple, efficient, decentralized algorithm for multiagent rl
Jin, C., Liu, Q., Wang, Y., and Yu, T · 2021
Cited alongside, same era.
Model-free learning for two-player zero-sum partially observable markov games with perfect recall
Kozuno, T., Ménard, P., Munos, R., and Valko, M · 2021
Cited alongside, same era.
Last-iterate convergence in extensive-form games
Lee, C.-W., Kroer, C., and Luo, H · 2021
Cited alongside, same era.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Guo, Z. D., Piot, B., Munos, R., Rowland, M., Valko, M., and Calandriello, D · 2024
Later among the works it cites.
Human alignment of large language models through online preference optimisation
Calandriello, D., Guo, D., Munos, R., Rowland, M., Tang, Y., Pires, B. A., Richemond, P. H., Lan, C. L., Valko, M., Liu, T., et al · 2024
Later among the works it cites.
Rlhf workflow: From reward modeling to online rlhf
Dong, H., Xiong, W., Pang, B., Wang, H., Zhao, H., Zhou, Y., Jiang, N., Sahoo, D., Xiong, C., and Zhang, T · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Cited alongside, same era.
A sharp analysis of model-based reinforcement learning with self-play
Liu, Q., Yu, T., Bai, Y., and Jin, C · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
Rl with kl penalties is better viewed as bayesian inference
Korbak, T., Perez, E., and Buckley, C. L · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Cited alongside, same era.
Koala: A dialogue model for academic research
Geng, X., Gudibande, A., Liu, H., Wallace, E., Abbeel, P., Levine, S., and Song, D · 2023
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Dubois, Y., Galambosi, B., Liang, P., and Hashimoto, T. B · 2024
Later among the works it cites.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D · 2024
Later among the works it cites.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., et al · 2024
Later among the works it cites.
From live data to high-quality benchmarks: The arena-hard pipeline, 2024
Li, T., Chiang, W.-L., Frick, E., Dunlap, L., Zhu, B., Gonzalez, J. E., and Stoica, I · 2024
Later among the works it cites.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C., Cheng, C.-A., Mitra, A., Santacroce, M., Awadallah, A., and Xie, T · 2024
Later among the works it cites.
Multi-turn reinforcement learning from preference human feedback
Shani, L., Rosenberg, A., Cassel, A., Lang, O., Calandriello, D., Zipori, A., Noga, H., Keller, O., Piot, B., Szpektor, I., et al · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., et al · 2024
Later among the works it cites.
Self-play preference optimization for language model alignment
Wu, Y., Sun, Z., Yuan, H., Ji, K., Yang, Y., and Gu, Q · 2024
Later among the works it cites.
Exploratory preference optimization: Harnessing implicit q*-approximation for sample-efficient rlhf
Xie, T., Foster, D. J., Krishnamurthy, A., Rosset, C., Awadallah, A., and Rakhlin, A · 2024
Later among the works it cites.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T · 2024
Later among the works it cites.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
Ye, C., Xiong, W., Zhang, Y., Jiang, N., and Zhang, T · 2024
Later among the works it cites.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Later among the works it cites.
Iterative nash policy optimization: Aligning llms with general preferences via no-regret learning
Zhang, Y., Yu, D., Peng, B., Song, L., Tian, Y., Huo, M., Jiang, N., Mi, H., and Yu, D · 2024
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2024
Later among the works it cites.