Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has emerged as the main paradigm for aligning large language models (LLMs) with human preferences.
Zur Theorie der Gesellschaftsspiele
von Neumann, J · 1928
Earlier work this paper cites.
The use of confidence or fiducial limits illustrated in the case of the binomial
Clopper, C. J. and Pearson, E. S · 1934
Earlier work this paper cites.
Iterative solution of games by fictitious play
Brown, G. W · 1951
Earlier work this paper cites.
An iterative method of solving a game
Robinson, J · 1951
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
On general minimax theorems
Sion, M · 1958
Earlier work this paper cites.
Existence and uniqueness of equilibrium points for concave n n -person games
Rosen, J. B · 1965
Earlier work this paper cites.
Intransitivity of preferences
Tversky, A · 1969
Earlier work this paper cites.
The paradox of the nontransitive dice
Gardner, M · 1970
Earlier work this paper cites.
The extragradient method for finding saddle points and other problems
Korpelevich, G · 1976
Earlier work this paper cites.
The Rating of Chessplayers, Past and Present
Elo, A. E · 1978
Earlier work this paper cites.
Information Theory: Coding Theorems for Discrete Memoryless Systems
Csiszar, I. and Korner, J · 1982
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovski, A. and Yudin, D · 1983
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins–Monro process
Ruppert, D · 1988
Earlier work this paper cites.
New stochastic approximation type procedures
Polyak, B. T · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
The Theory of Learning in Games
Fudenberg, D. and Levine, D. K · 1998
Earlier work this paper cites.
Excessive gap technique in nonsmooth convex minimization
Nesterov, Y · 2005
Earlier work this paper cites.
Predition, Learning, and Games
Cesa-Biachi, N. and Lugosi, G · 2006
Earlier work this paper cites.
Best response dynamics for continuous zero-sum games
Hofbauer, J. and Sorin, S · 2006
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 2007
Earlier work this paper cites.
Smoothing techniques for computing Nash equilibria of sequential games
Hoda, S., Gilpin, A., Peña, J., and Sandholm, T · 2010
Earlier work this paper cites.
Preference-based policy learning
Akrour, R., Schoenauer, M., and Sebag, M · 2011
Earlier work this paper cites.
Preference-based policy iteration: Leveraging preference learning for reinforcement learning
Cheng, W., Fürnkranz, J., Hüllermeier, E., and Park, S.-H · 2011
Earlier work this paper cites.
Near-optimal no-regret algorithms for zero-sum games
Daskalakis, C., Deckelbaum, A., and Kim, A · 2011
Earlier work this paper cites.
A Bayesian approach for policy learning from trajectory preference queries
Wilson, A., Fern, A., and Tadepalli, P · 2012
Earlier work this paper cites.
Preference-based evolutionary direct policy search
Busa-Fekete, R., Szörenyi, B., Weng, P., Cheng, W., and Hüllermeier, E · 2013
Earlier work this paper cites.
Optimization, learning, and games with predictable sequences
Rakhlin, S. and Sridharan, K · 2013
Earlier work this paper cites.
Preference-based reinforcement learning: Evolutionary direct policy search using a preference-based racing algorithm
Busa-Fekete, R., Szörényi, B., Weng, P., Cheng, W., and Hüllermeier, E · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Cited alongside, same era.
Fictitious self-play in extensive-form games
Heinrich, J., Lanctot, M., and Silver, D · 2015
Cited alongside, same era.
Intransitivity in theory and in the real world
Klimenko, A. Y · 2015
Cited alongside, same era.
Fast convergence of regularized learning in games
Syrgkanis, V., Agarwal, A., Luo, H., and Schapire, R. E · 2015
Cited alongside, same era.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
Frank-Wolfe algorithms for saddle point problems
Gidel, G., Jebara, T., and Lacoste-Julien, S · 2016
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Later among the works it cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Chen, X., Zhong, H., Yang, Z., Wang, Z., and Wang, L · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., Campbell-Gillingham, L., Uesato, J., Huang, P.-S., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L. A., and Irving, G · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., Glaese, M., Young, S., Campbell-Gillingham, L., Irving, G., and McAleese, N · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Heinrich, J. and Silver, D · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
TL;DR: Mining Reddit to learn automatic summarization
Völske, M., Potthast, M., Syed, S., and Stein, B · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., and Fürnkranz, J · 2017
Cited alongside, same era.
Faster rates for convex-concave games
Abernethy, J., Lai, K. A., Levy, K. Y., and Wang, J.-K · 2018
Cited alongside, same era.
Introducing ChatGPT, 2022
OpenAI · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Mastering the game of Stratego with model-free multiagent reinforcement learning
Perolat, J., Vylder, B. D., Hennes, D., Tarassov, E., Strub, F., de Boer, V., Muller, P., Connor, J. T., Burch, N., Anthony, T., McAleer, S., Elie, R., Cen, S. H., Wang, Z., Gruslys, A., Malysheva, A., Khan, M., Ozair, S., Timbers, F., Pohlen, T., Eccles, T., Rowland, M., Lanctot, M., Lespiau, J.-B., Piot, B., Omidshafiei, S., Lockhart, E., Sifre, L., Beauguerlange, N., Munos, R., Silver, D., Singh, S., Hassabis, D., and Tuyls, K · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., and Schmidt, L · 2022
Later among the works it cites.
PaLM 2 technical report, 2023
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., Chu, E., Clark, J. H., Shafey, L. E., Huang, Y., Meier-Hellstern, K., Mishra, G., Moreira, E., Omernick, M., Robinson, K., Ruder, S., Tay, Y., Xiao, K., Xu, Y., Zhang, Y., Abrego, G. H., Ahn, J., Austin, J., Barham, P., Botha, J., Bradbury, J., Brahma, S., Brooks, K., Catasta, M., Cheng, Y., Cherry, C., Choquette-Choo, C. A., Chowdhery, A., Crepy, C., Dave, S., Dehghani, M., Dev, S., Devlin, J., Díaz, M., Du, N., Dyer, E., Feinberg, V., Feng, F., Fienber, V., Freitag, M., Garcia, X., Gehrmann, S., Gonzalez, L., Gur-Ari, G., Hand, S., Hashemi, H., Hou, L., Howland, J., Hu, A., Hui, J., Hurwitz, J., Isard, M., Ittycheriah, A., Jagielski, M., Jia, W., Kenealy, K., Krikun, M., Kudugunta, S., Lan, C., Lee, K., Lee, B., Li, E., Li, M., Li, W., Li, Y., Li, J., Lim, H., Lin, H., Liu, Z., Liu, F., Maggioni, M., Mahendru, A., Maynez, J., Misra, V., Moussalem, M., Nado, Z., Nham, J., Ni, E., Nystrom, A., Parrish, A., Pellat, M., Polacek, M., Polozov, A., Pope, R., Qiao, S., Reif, E., Richter, B., Riley, P., Ros, A. C., Roy, A., Saeta, B., Samuel, R., Shelby, R., Slone, A., Smilkov, D., So, D. R., Sohn, D., Tokumine, S., Valter, D., Vasudevan, V., Vodrahalli, K., Wang, X., Wang, P., Wang, Z., Wang, T., Wieting, J., Wu, Y., Xu, K., Xu, Y., Xue, L., Yin, P., Yu, J., Zhang, Q., Zheng, S., Zheng, C., Zhou, W., Zhou, D., Petrov, S., and Wu, Y · 2023
Closest in time.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R · 2023
Closest in time.
On the limitations of the Elo: Real-world games are transitive, not additive
Bertrand, Q., Czarnecki, W. M., and Gidel, G · 2023
Closest in time.
How to scale your EMA
Busbridge, D., Ramapuram, J., Ablin, P., Likhomanenko, T., Dhekane, E. G., Suau, X., and Webb, R · 2023
Closest in time.
RAFT: Reward rAnked FineTuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., SHUM, K., and Zhang, T · 2023
Closest in time.
Reinforced self-training (ReST) for language modeling, 2023
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., and de Freitas, N · 2023
Closest in time.
RLAIF: Scaling reinforcement learning from human feedback with AI feedback
Lee, H., Phatale, S., Mansoor, H., Lu, K., Mesnard, T., Bishop, C., Carbune, V., and Rastogi, A · 2023
Closest in time.
Confronting reward model overoptimization with constrained RLHF
Moskovitz, T., Singh, A. K., Strouse, D., Sandholm, T., Salakhutdinov, R., Dragan, A. D., and McAleer, S · 2023
Closest in time.
GPT-4 technical report
OpenAI · 2023
Closest in time.
Dueling RL: Reinforcement learning with trajectory preferences
Pacchiano, A., Saha, A., and Lee, J · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Closest in time.
Rewarded soups: Towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Rame, A., Couairon, G., Shukor, M., Dancette, C., Gaya, J.-B., Soulier, L., and Cord, M · 2023
Closest in time.
Scaling up models and data with t5x
Roberts, A., Chung, H. W., Levskaya, A., Mishra, G., Bradbury, J., Andor, D., Narang, S., Lester, B., Gaffney, C., Mohiuddin, A., Hawthorne, C., Lewkowycz, A., Salcianu, A., van Zee, M., Austin, J., Goodman, S., Soares, L. B., Hu, H., Tsvyashchenko, S., Chowdhery, A., Bastings, J., Bulian, J., Garcia, X., Ni, J., Chen, A., Kenealy, K., Clark, J. H., Lee, S., Garrette, D., Lee-Thorp, J., Raffel, C., Shazeer, N., Ritter, M., Bosma, M., Passos, A., Maitin-Shepard, J., Fiedel, N., Omernick, M., Saeta, B., Sepassi, R., Spiridonov, A., Newlan, J., and Gesmundo, A · 2023
Closest in time.
Factually consistent summarization via reinforcement learning with textual entailment feedback
Roit, P., Ferret, J., Shani, L., Aharoni, R., Cideron, G., Dadashi, R., Geist, M., Girgin, S., Hussenot, L., Keller, O., Momchev, N., Ramos, S., Stanczyk, P., Vieillard, N., Bachem, O., Elidan, G., Hassidim, A., Pietquin, O., and Szpektor, I · 2023
Closest in time.
A unified approach to reinforcement learning, quantal response equilibria, and two-player zero-sum games
Sokota, S., D’Orazio, R., Kolter, J. Z., Loizou, N., Lanctot, M., Mitliagkas, I., Brown, N., and Kroer, C · 2023
Closest in time.
Is RLHF more difficult than standard RL?
Wang, Y., Liu, Q., and Jin, C · 2023
Closest in time.
SLiC-HF: Sequence likelihood calibration with human feedback
Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J · 2023
Closest in time.
Human alignment of large language models through online preference optimisation
Calandriello, D., Guo, D., Munos, R., Rowland, M., Tang, Y., Pires, B. A., Richemond, P. H., Lan, C. L., Valko, M., Liu, T., Joshi, R., Zheng, Z., and Piot, B · 2024
Closest in time.
Multi-turn reinforcement learning from preference human feedback
Shani, L., Rosenberg, A., Cassel, A., Lang, O., Calandriello, D., Zipori, A., Noga, H., Keller, O., Piot, B., Szpektor, I., Hassidim, A., Matias, Y., and Munos, R · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment, 2024
Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P. H., Valko, M., Ávila Pires, B., and Piot, B · 2024
Closest in time.