Fetching the paper…
Reading the bibliography…
We present Self-Play Preference Optimization (SPO), an algorithm for reinforcement learning from human feedback.
A difficulty in the concept of social welfare
Arrow, K. J · 1950
Earlier work this paper cites.
Non-cooperative games
Nash, J · 1951
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Intransitivity, utility, and the aggregation of preference patterns
May, K. O · 1954
Earlier work this paper cites.
The magical number seven, plus or minus two: Some limits on our capacity for processing information
Miller, G. A · 1956
Earlier work this paper cites.
On general minimax theorems
Sion, M · 1958
Earlier work this paper cites.
Aggregation of preference orderings
Kreweras, G · 1965
Earlier work this paper cites.
On Defining Areas of Voter Choice: Professor Tullock on Stable Voting
Simpson, P. B · 1969
Earlier work this paper cites.
Intransitivity of preferences
Tversky, A · 1969
Earlier work this paper cites.
Mathematical games, Dec 1970
Gardner, M · 1970
Earlier work this paper cites.
On a class of equilibrium conditions for majority rule
Kramer, G. H · 1973
Earlier work this paper cites.
Strategy-proofness and arrow’s conditions: Existence and correspondence theorems for voting procedures and social welfare functions
Satterthwaite, M. A · 1975
Earlier work this paper cites.
Probabilistic social choice based on simple voting comparisons
Fishburn, P. C · 1984
Earlier work this paper cites.
Social choice theory
Sen, A · 1986
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y. and Schapire, R. E · 1997
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Policy search by dynamic programming
Bagnell, J. A., Kakade, S., Ng, A. Y., and Schneider, J · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Efficient algorithms for online decision problems
Kalai, A. and Vempala, S · 2005
Earlier work this paper cites.
Generalised weakened fictitious play
Leslie, D. S. and Collins, E. J · 2006
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 2007
Earlier work this paper cites.
Preference learning in recommender systems
De Gemmis, M., Iaquinta, L., Lops, P., Musto, C., Narducci, F., and Semeraro, G · 2009
Earlier work this paper cites.
Online markov decision processes
Even-Dar, E., Kakade, S. M., and Mansour, Y · 2009
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yue, Y. and Joachims, T · 2009
Earlier work this paper cites.
Preference-based learning to rank
Ailon, N. and Mohri, M · 2010
Earlier work this paper cites.
Optimal bayesian recommendation sets and myopically optimal choice query sets
Viappiani, P. and Boutilier, C · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
Human preferences for robot-human hand-over configurations
Cakmak, M., Srinivasa, S. S., Lee, M. K., Forlizzi, J., and Kiesler, S · 2011
Cited alongside, same era.
Follow-the-regularized-leader and mirror descent: Equivalence theorems and l1 regularization
McMahan, B · 2011
Cited alongside, same era.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Cited alongside, same era.
April: Active preference learning-based reinforcement learning
Akrour, R., Schoenauer, M., and Sebag, M · 2012
Cited alongside, same era.
Symmetric games with only asymmetric equilibria
Fey, M · 2012
Cited alongside, same era.
The k-armed dueling bandits problem
Yue, Y., Broder, J., Kleinberg, R., and Joachims, T · 2012
Cited alongside, same era.
Fine-tuning language models from human preferences, 2020
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2021
Later among the works it cites.
Preference-based online learning with dueling bandits: A survey
Bengs, V., Busa-Fekete, R., Mesaoudi-Paul, A. E., and HÃllermeier, E · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Later among the works it cites.
Adversarial dueling bandits
Saha, A., Koren, T., and Mansour, Y · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimization and learning for rough terrain legged locomotion
Zucker, M., Ratliff, N., Stolle, M., Chestnutt, J., Bagnell, J. A., Atkeson, C. G., and Kuffner, J · 2012
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Contextual dueling bandits
Dudík, M., Hofmann, K., Schapire, R. E., Slivkins, A., and Zoghi, M · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Consistent probabilistic social choice
Brandl, F., Brandt, F., and Seedig, H. G · 2016
Cited alongside, same era.
Of moments and matching: A game-theoretic framework for closing the imitation gap
Swamy, G., Choudhury, S., Bagnell, J. A., and Wu, S · 2021
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
Afsar, M. M., Crump, T., and Far, B · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Efficient and optimal algorithms for contextual dueling bandits under realizability
Saha, A. and Krishnamurthy, A · 2022
Later among the works it cites.
A ranking game for imitation learning
Sikchi, H., Saran, A., Goo, W., and Niekum, S · 2022
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R · 2023
Later among the works it cites.
Uncoupled and convergent learning in two-player zero-sum markov games with bandit feedback, 2023
Cai, Y., Luo, H., Wei, C.-Y., and Zheng, W · 2023
Later among the works it cites.
Contrastive prefence learning: Learning from human feedback without rl
Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D · 2023
Later among the works it cites.
Nash learning from human feedback, 2023
Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., Tang, Y., Geist, M., Mesnard, T., Michi, A., Selvi, M., Girgin, S., Momchev, N., Bachem, O., Mankowitz, D. J., Precup, D., and Piot, B · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI, :, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mély, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Michael, Pokorny, Pokrass, M., Pong, V., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B · 2023
Later among the works it cites.
Dueling rl: Reinforcement learning with trajectory preferences, 2023
Pacchiano, A., Saha, A., and Lee, J · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model, 2023
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
Is rlhf more difficult than standard rl?, 2023
Wang, Y., Liu, Q., and Jin, C · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J · 2023
Later among the works it cites.
Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023
Zhu, B., Frick, E., Wu, T., Zhu, H., and Jiao, J · 2023
Later among the works it cites.
Human alignment of large language models through online preference optimisation
Calandriello, D., Guo, D., Munos, R., Rowland, M., Tang, Y., Pires, B. A., Richemond, P. H., Lan, C. L., Valko, M., Liu, T., et al · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models, 2024
Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q · 2024
Closest in time.
Rebel: Reinforcement learning via regressing relative rewards
Gao, Z., Chang, J. D., Zhan, W., Oertell, O., Swamy, G., Brantley, K., Joachims, T., Bagnell, J. A., Lee, J. D., and Sun, W · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C., Cheng, C.-A., Mitra, A., Santacroce, M., Awadallah, A., and Xie, T · 2024
Closest in time.
Understanding preference fine-tuning through the lens of coverage, 2024
Song, Y., Swamy, G., Singh, A., Bagnell, J. A., and Sun, W · 2024
Closest in time.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Tajwar, F., Singh, A., Sharma, A., Rafailov, R., Schneider, J., Xie, T., Ermon, S., Finn, C., and Kumar, A · 2024
Closest in time.
Understanding the performance gap between online and offline alignment algorithms
Tang, Y., Guo, D. Z., Zheng, Z., Calandriello, D., Cao, Y., Tarassov, E., Munos, R., Pires, B. Á., Valko, M., Cheng, Y., et al · 2024
Closest in time.
Is dpo superior to ppo for llm alignment? a comprehensive study
Xu, S., Fu, W., Gao, J., Ye, W., Liu, W., Mei, Z., Wang, G., Yu, C., and Wu, Y · 2024
Closest in time.