Fetching the paper…
Reading the bibliography…
Modeling human preferences is crucial for aligning foundation models with human values.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Intransitivity of preferences
Tversky, A · 1969
Earlier work this paper cites.
Mathematical games
Gardner, M · 1970
Earlier work this paper cites.
Social choice theory
Sen, A · 1986
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Freund, Y. and Schapire, R. E · 1999
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Game theory
Owen, G · 2013
Earlier work this paper cites.
Contextual dueling bandits
Dudík, M., Hofmann, K., Schapire, R. E., Slivkins, A., and Zoghi, M · 2015
Earlier work this paper cites.
Stochastic choice and preferences for randomization
Agranov, M. and Ortoleva, P · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
A law of comparative judgment
Thurstone, L. L · 2017
Earlier work this paper cites.
Re-evaluating evaluation
Balduzzi, D., Tuyls, K., Perolat, J., and Graepel, T · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Open-ended learning in symmetric zero-sum games
Balduzzi, D., Garnelo, M., Bachrach, Y., Czarnecki, W., Perolat, J., Jaderberg, M., and Graepel, T · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Cited alongside, same era.
Real world games look like spinning tops
Czarnecki, W. M., Gidel, G., Tracey, B., Tuyls, K., Omidshafiei, S., Balduzzi, D., and Jaderberg, M · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Rlhf workflow: From reward modeling to online rlhf
Dong, H., Xiong, W., Pang, B., Wang, H., Zhao, H., Zhou, Y., Jiang, N., Sahoo, D., Xiong, C., and Zhang, T · 2024
Closest in time.
The llama 3 herd of models, 2024
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., and et al · 2024
Closest in time.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Dubois, Y., Galambosi, B., Liang, P., and Hashimoto, T. B · 2024
Closest in time.
Geometric-averaged preference optimization for soft preference labels
Furuta, H., Lee, K.-H., Gu, S. S., Matsuo, Y., Faust, A., Zen, H., and Gur, I · 2024
Closest in time.
Rebel: Reinforcement learning via regressing relative rewards
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Active ranking without strong stochastic transitivity
Lou, H., Jin, T., Wu, Y., Xu, P., Gu, Q., and Farnoud, F · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R · 2023
Cited alongside, same era.
On the limitations of the elo, real-world games are transitive, not additive
Bertrand, Q., Czarnecki, W. M., and Gidel, G · 2023
Cited alongside, same era.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Jiang, D., Ren, X., and Lin, B. Y · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Cited alongside, same era.
Gao, Z., Chang, J. D., Zhan, W., Oertell, O., Swamy, G., Brantley, K., Joachims, T., Bagnell, J. A., Lee, J. D., and Sun, W · 2024
Closest in time.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Hu, J., Wu, X., Wang, W., Xianyu, Zhang, D., and Cao, Y · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N. A., and Hajishirzi, H · 2024
Closest in time.
Robust preference optimization with provable noise tolerance for llms
Liang, X., Chen, C., Wang, J., Wu, Y., Fu, Z., Shi, Z., Wu, F., and Ye, J · 2024
Closest in time.
Skywork reward model series
Liu, C. Y. and Zeng, L · 2024
Closest in time.
Simpo: Simple preference optimization with a reference-free reward
Meng, Y., Xia, M., and Chen, D · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C., Cheng, C.-A., Mitra, A., Santacroce, M., Awadallah, A., and Xie, T · 2024
Closest in time.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models, 2024
Röttger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Swamy, G., Dann, C., Kidambi, R., Wu, Z. S., and Agarwal, A · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size, 2024
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., Ferret, J., Liu, P., Tafti, P., Friesen, A., Casbon, M., Ramos, S., Kumar, R., Lan, C. L., Jerome, S., Tsitsulin, A., Vieillard, N., Stanczyk, P., Girgin, S., Momchev, N., Hoffman, M., Thakoor, S., Grill, J.-B., Neyshabur, B., Bachem, O., Walton, A., Severyn, A., Parrish, A., Ahmad, A., Hutchison, A., and et al · 2024
Closest in time.
Do-not-answer: Evaluating safeguards in LLMs
Wang, Y., Li, H., Han, X., Nakov, P., and Baldwin, T · 2024
Closest in time.
Evaluating large language models at evaluating instruction following
Zeng, Z., Yu, J., Gao, T., Meng, Y., Goyal, T., and Chen, D · 2024
Closest in time.