Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has emerged as a key technique for aligning the output of large language models (LLMs) with human preferences.
1909
Earlier work this paper cites.
[author] Bradley, Ralph AllanR. A. and Terry, Milton EM. E. (1952). Rank analysis of incomplete block designs: I. The method of paired comparisons. Biometrika 39 324–345
1952
Earlier work this paper cites.
[author] May, Kenneth O.K. O. (1954). Intransitivity, Utility, and the Aggregation of Preference Patterns. Econometrica 22 1
1954
Earlier work this paper cites.
[author] Tversky, AmosA. (1969). Intransitivity of preferences. Psychological Review 76 31-48
1969
Earlier work this paper cites.
[author] Gardner, MM. (1970). Mathematical games, the paradox of the nontransitive dice and the elusive principle of indifference
1970
Earlier work this paper cites.
2005
Earlier work this paper cites.
[author] Tsiatis, Anastasios AA. A. (2006). Semiparametric theory and missing data 4. Springer
2006
Earlier work this paper cites.
Maas, A. L
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
[author] Agranov, MarinaM. and Ortoleva, PietroP. (2015). Stochastic Choice and Preferences for Randomization. Journal of Political Economy 125 40 - 68
2015
Earlier work this paper cites.
[author] Christiano, Paul FP. F., Leike, JanJ., Brown, TomT., Martic, MiljanM., Legg, ShaneS. and Amodei, DarioD. (2017). Deep reinforcement learning from human preferences. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
[author] Vaswani, AshishA., Shazeer, NoamN., Parmar, NikiN., Uszkoreit, JakobJ., Jones, LlionL., Gomez, Aidan NA. N., Kaiser, ŁukaszŁ. and Polosukhin, IlliaI. (2017). Attention is all you need. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Völske, M
2017
Earlier work this paper cites.
[author] Chakrabortty, AbhishekA. and Cai, TianxiT. (2018). Efficient and adaptive linear regression in semi-supervised settings. Annals of Statistics
2018
Earlier work this paper cites.
[author] Sutton, Richard SR. S., Barto, Andrew GA. G. et al. (2018). Reinforcement learning: An introduction. MIT press Cambridge
2018
Earlier work this paper cites.
[author] Agarwal, AlekhA., Jiang, NanN., Kakade, Sham MS. M. and Sun, WenW. (2019). Reinforcement learning: Theory and algorithms. CS Dept., UW Seattle, Seattle, WA, USA, Tech. Rep 32 96
2019
Earlier work this paper cites.
[author] Zhang, AnruA., Brown, Lawrence DL. D. and Cai, T TonyT. T. (2019). Semi-supervised Inference: General Theory and Estimation of Means. The Annals of Statistics 47 2538–2566
2019
Earlier work this paper cites.
[author] Milano, SilviaS., Taddeo, MariarosariaM. and and, Luciano FloridiL. F. (2021). Ethical aspects of multi-stakeholder recommendation systems. The Information Society 37 35–45. 10.1080/01972243.2020.1832636
2020
Earlier work this paper cites.
Stiennon, N
2020
Earlier work this paper cites.
[author] von Werra, LeandroL., Belkada, YounesY., Tunstall, LewisL., Beeching, EdwardE., Thrush, TristanT., Lambert, NathanN., Huang, ShengyiS., Rasul, KashifK. and Gallouédec, QuentinQ. (2020). TRL: Transformer Reinforcement Learning. https://github.com/huggingface/trl
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
[author] Bakker, MichielM., Chadwick, MartinM., Sheahan, HannahH., Tessler, MichaelM., Campbell-Gillingham, LucyL., Balaguer, JanJ., McAleese, NatN., Glaese, AmeliaA., Aslanides, JohnJ., Botvinick, MattM. et al. (2022). Fine-tuning language models to find agreement among humans with diverse preferences. Advances in Neural Information Processing Systems 35 38176–38189
2022
Earlier work this paper cites.
[author] Hartmann, JochenJ., Heitmann, MarkM., Siebert, ChristianC. and Schamp, ChristinaC. (2023). More than a Feeling: Accuracy and Application of Sentiment Analysis. International Journal of Research in Marketing 40 75-87. https://doi.org/10.1016/j.ijresmar.2022.05.005
2022
Earlier work this paper cites.
[author] Huang, ShengyiS., Dossa, Rousslan Fernand JulienR. F. J., Ye, ChangC., Braga, JeffJ., Chakraborty, DipamD., Mehta, KinalK. and Araújo, João G. M.J. G. M. (2022). CleanRL: High-quality Single-file Implementations of Deep Reinforcement Learning Algorithms. Journal of Machine Learning Research 23 1–18
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
[author] Ouyang, LongL., Wu, JeffreyJ., Jiang, XuX., Almeida, DiogoD., Wainwright, CarrollC., Mishkin, PamelaP., Zhang, ChongC., Agarwal, SandhiniS., Slama, KatarinaK., Ray, AlexA. et al. (2022). Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 27730–27744
2022
Cited alongside, same era.
2022
Cited alongside, same era.
[author] Zhang, YuqianY. and Bradic, JelenaJ. (2022). High-dimensional semi-supervised learning: in search of optimal inference of the mean. Biometrika 109 387–403
2022
2024
Later among the works it cites.
2024
Later among the works it cites.
Mukherjee, S
2024
Later among the works it cites.
Munos, R
2024
Later among the works it cites.
Rafailov, R
2024
Later among the works it cites.
[author] Ramesh, Shyam SundharS. S., Hu, YifanY., Chaimalas, IasonI., Mehta, VirajV., Sessa, Pier GiuseppeP. G., Bou Ammar, HaithamH. and Bogunovic, IlijaI. (2024). Group robust preference optimization in reward-free rlhf. Advances in Neural Information Processing Systems 37 37100–37137
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
[author] Angelopoulos, Anastasios NA. N., Bates, StephenS., Fannjiang, ClaraC., Jordan, Michael IM. I. and Zrnic, TijanaT. (2023). Prediction-powered inference. Science 382 669–674
2023
Cited alongside, same era.
Bertrand, Q
2023
Cited alongside, same era.
[author] Dubois, YannY., Li, Chen XuechenC. X., Taori, RohanR., Zhang, TianyiT., Gulrajani, IshaanI., Ba, JimmyJ., Guestrin, CarlosC., Liang, Percy SP. S. and Hashimoto, Tatsunori BT. B. (2023). Alpacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems 36 30039–30069
2023
Cited alongside, same era.
2023
Cited alongside, same era.
[author] Mitchell, EricE. (2023). A note on dpo with noisy preferences and relationship to ipo. URL https://ericmitchell. ai/cdpo. pdf
2023
Cited alongside, same era.
Rafailov, R
2023
Cited alongside, same era.
[author] Touvron, HugoH., Lavril, ThibautT., Izacard, GautierG., Martinet, XavierX., Lachaux, Marie-AnneM.-A., Lacroix, TimothéeT., Rozière, BaptisteB., Goyal, NamanN., Hambro, EricE., Azhar, FaisalF., Rodriguez, AurelienA., Joulin, ArmandA., Grave, EdouardE. and Lample, GuillaumeG. (2023). LLaMA: Open and Efficient Foundation Language Models
2023
Cited alongside, same era.
[author] Wang, JiashuoJ., Wang, HaozhaoH., Sun, ShichaoS. and Li, WenjieW. (2023). Aligning language models with human preferences via a bayesian approach. Advances in Neural Information Processing Systems 36 49113–49132
2023
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
Swamy, G
2024
Later among the works it cites.
[author] Team, QwenQ. (2024). Qwen2.5: A Party of Foundation Models
2024
Later among the works it cites.
[author] Wu, YueY., Sun, ZhiqingZ., Yuan, HuizhuoH., Ji, KaixuanK., Yang, YimingY. and Gu, QuanquanQ. (2024b). Self-play preference optimization for language model alignment. NeurIPS (2024) Workshop on Adaptive Foundation Models: Evolving AI for Personalized and Efficient Learning
2024
Later among the works it cites.
Xiong, W
2024
Later among the works it cites.
2024
Later among the works it cites.
[author] Ye, ChenluC., Xiong, WeiW., Zhang, YuhengY., Dong, HanzeH., Jiang, NanN. and Zhang, TongT. (2024). Online iterative reinforcement learning from human feedback with general preference model. Advances in Neural Information Processing Systems 37 81773–81807
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
[author] Lu, ChrisC., Holt, SamuelS., Fanconi, ClaudioC., Chan, AlexA., Foerster, JakobJ., van der Schaar, MihaelaM. and Lange, RobertR. (2025). Discovering preference optimization algorithms with and for large language models. Advances in Neural Information Processing Systems 37 86528–86573
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Zhang, Y
2025
Closest in time.