Fetching the paper…
Reading the bibliography…
The Bradley-Terry (BT) model is a common and successful practice in reward modeling for Large Language Model (LLM) alignment.
A law of comparative judgment
L. Thurstone · 1927
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Solution of a ranking problem from binary comparisons
L. R. Ford Jr · 1957
Earlier work this paper cites.
Stimulus and response generalization: A stochastic model relating generalization to distance in psychological space
R. N. Shepard · 1957
Earlier work this paper cites.
Individual choice behavior , volume 4
R. D. Luce · 1959
Earlier work this paper cites.
The proposed uscf rating system, its development, theory, and applications
A. E. Elo · 1967
Earlier work this paper cites.
Paradox of nontransitive dice and elusive principle of indifference
M. Gardner · 1970
Earlier work this paper cites.
Response surface fitting using a generalization of the bradley-terry paired comparison model
A. Springall · 1973
Earlier work this paper cites.
A logistic representation of multivariate paired-comparison models
U. Bockenholt · 1988
Earlier work this paper cites.
Advances in prospect theory: Cumulative representation of uncertainty
A. Tversky and D. Kahneman · 1992
Earlier work this paper cites.
Signature verification using a” siamese” time delay neural network
J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah · 1993
Earlier work this paper cites.
A thurstonian pairwise choice model with univariate and multivariate spline transformations
G. De Soete and S. Winsberg · 1993
Earlier work this paper cites.
A comprehensive guide to chess ratings
M. E. Glickman · 1995
Earlier work this paper cites.
Rating the chess rating system
M. E. Glickman and A. C. Jones · 1999
Earlier work this paper cites.
Asymptotics when the number of parameters tends to infinity in the bradley-terry model for paired comparisons
G. Simons and Y.-C. Yao · 1999
Earlier work this paper cites.
Absolute identification by relative judgment
N. Stewart, G. D. Brown, and N. Chater · 2005
Earlier work this paper cites.
Models for paired comparison data: A review with emphasis on dependent data
M. Cattelan · 2012
Earlier work this paper cites.
Estimation of nonparametric conditional moment models with possibly nonsmooth generalized residuals
X. Chen and D. Pouzo · 2012
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire · 2014
Earlier work this paper cites.
Relative judgement is relatively difficult: Evidence against the role of relative judgement in absolute identification
D. Guest, J. S. Adelman, and C. Kent · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
On calibration of modern neural networks
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree
G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Cited alongside, same era.
Asymptotic theory of sparse bradley–terry model
R. Han, R. Ye, C. Tan, and K. Chen · 2020
Cited alongside, same era.
Bandit algorithms
T. Lattimore and C. Szepesvári · 2020
Cited alongside, same era.
Nonparametric regression using deep neural networks with relu activation function
J. Schmidt-Hieber · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Trl: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, and S. Huang · 2020
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Helpsteer: Multi-attribute helpfulness dataset for steerlm
Z. Wang, Y. Dong, J. Zeng, V. Adams, M. N. Sreedhar, D. Egert, O. Delalleau, J. P. Scowcroft, N. Kant, A. Swope, et al · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Z. Yuan, H. Yuan, C. Tan, W. Wang, S. Huang, and F. Huang · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Y. Zhao, R. Joshi, T. Liu, M. Khalman, M. Saleh, and P. J. Liu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A comparative analysis of gradient boosting algorithms
C. Bentéjac, A. Csörgő, and G. Martínez-Muñoz · 2021
Cited alongside, same era.
Convergence rates of deep relu networks for multiclass classification
T. Bos and J. Schmidt-Hieber · 2022
Cited alongside, same era.
Learning from demonstration: Provably efficient adversarial policy imitation with linear function approximation
Z. Liu, Y. Zhang, Z. Fu, Z. Yang, and Z. Wang · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Asymptotic comparison of identifying constraints for bradley-terry models
W. Wu, B. W. Junker, and N. Niezink · 2022
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, M. Rowland, B. Piot, D. Guo, D. Calandriello, M. Valko, and R. Munos · 2023
Cited alongside, same era.
R. Zheng, S. Dou, S. Gao, Y. Hua, W. Shen, B. Wang, Y. Liu, S. Jin, Q. Liu, Y. Zhou, et al · 2023
Later among the works it cites.
Chatbot arena: An open platform for evaluating llms by human preference
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, et al · 2024
Closest in time.
Rlhf workflow: From reward modeling to online rlhf
H. Dong, W. Xiong, B. Pang, H. Wang, H. Zhao, Y. Zhou, N. Jiang, D. Sahoo, C. Xiong, and T. Zhang · 2024
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Y. Dubois, C. X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. S. Liang, and T. B. Hashimoto · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela · 2024
Closest in time.
Direct language model alignment from online ai feedback
S. Guo, B. Zhang, T. Liu, T. Liu, M. Khalman, F. Llinares, A. Rame, T. Mesnard, Y. Zhao, B. Piot, et al · 2024
Closest in time.
Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback
H. Ivison, Y. Wang, J. Liu, Z. Wu, V. Pyatkin, N. Lambert, N. A. Smith, Y. Choi, and H. Hajishirzi · 2024
Closest in time.
Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer
Z. Liu, M. Lu, S. Zhang, B. Liu, H. Guo, Y. Yang, J. Blanchet, and Z. Wang · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
A. Meta · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Inverse-rlignment: Inverse reinforcement learning from demonstrations for llm alignment
H. Sun and M. van der Schaar · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Y. Tang, Z. D. Guo, Z. Zheng, D. Calandriello, R. Munos, M. Rowland, P. H. Richemond, M. Valko, B. Á. Pires, and B. Piot · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Snorkel-mistral-pairrm-dpo, 2024
H. Tran and B. Chris Glaze · 2024
Closest in time.
Is dpo superior to ppo for llm alignment? a comprehensive study
S. Xu, W. Fu, J. Gao, W. Ye, W. Liu, Z. Mei, G. Wang, C. Yu, and Y. Wu · 2024
Closest in time.
Nonparametric logistic regression with deep learning
A. Yara and Y. Terada · 2024
Closest in time.
Y. Yin, Z. Wang, Y. Gu, H. Huang, W. Chen, and M. Zhou · 2024
Closest in time.
Generative verifiers: Reward modeling as next-token prediction
L. Zhang, A. Hosseini, H. Bansal, M. Kazemi, A. Kumar, and R. Agarwal · 2024
Closest in time.
Toward optimal llm alignments using two-player games
R. Zheng, H. Guo, Z. Liu, X. Zhang, Y. Yao, X. Xu, Z. Wang, Z. Xi, T. Gui, Q. Zhang, et al · 2024
Closest in time.