Fetching the paper…
Reading the bibliography…
Building neural reward models from human preferences is a pivotal component in reinforcement learning from human feedback (RLHF) and large language model alignment research.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E. (1952) · 1952
Earlier work this paper cites.
Bayes empirical bayes
Deely, J. and Lindley, D. (1981) · 1981
Earlier work this paper cites.
Bayesian experimental design: A review
Chaloner, K. and Verdinelli, I. (1995) · 1995
Earlier work this paper cites.
Asymptotic statistics
Van der Vaart, A. W. (2000) · 2000
Earlier work this paper cites.
Optimal design of experiments
Pukelsheim, F. (2006) · 2006
Earlier work this paper cites.
Optimum experimental designs, with SAS
Atkinson, A., Donev, A., and Tobias, R. (2007) · 2007
Earlier work this paper cites.
Mathematical statistics
Shao, J. (2008) · 2008
Earlier work this paper cites.
Active learning literature survey
Settles, B. (2009) · 2009
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M. (2011) · 2011
Earlier work this paper cites.
Fisher information distance: A geometrical reading
Costa, S. I., Santos, S. A., and Strapasson, J. E. (2015) · 2015
Earlier work this paper cites.
Coresets for scalable bayesian logistic regression
Huggins, J., Campbell, T., and Broderick, T. (2016) · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Sener, O. and Savarese, S. (2017) · 2017
Earlier work this paper cites.
Bayesian coreset construction via greedy iterative geodesic ascent
Campbell, T. and Broderick, T. (2018) · 2018
Earlier work this paper cites.
On coresets for logistic regression
Munteanu, A., Schwiegelshohn, C., Sohler, C., and Woodruff, D. (2018) · 2018
Cited alongside, same era.
Automated scalable bayesian inference via hilbert coresets
Campbell, T. and Broderick, T. (2019) · 2019
Cited alongside, same era.
Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning
Kirsch, A., Van Amersfoort, J., and Gal, Y. (2019) · 2019
Cited alongside, same era.
Bayesian layers: A module for neural network uncertainty
Tran, D., Dusenberry, M., Van Der Wilk, M., and Hafner, D. (2019) · 2019
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020) · 2020
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Bonbon alignment for large language models and the sweetness of best-of-n sampling
Gui, L., Gârbacea, C., and Veitch, V. (2024) · 2024
Later among the works it cites.
Rewardbench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N. A., and Hajishirzi, H. (2024) · 2024
Later among the works it cites.
Skywork-reward: Bag of tricks for reward modeling in llms
Liu, C. Y., Zeng, L., Liu, J., Yan, R., He, J., Wang, C., Yan, S., Liu, Y., and Zhou, Y. (2024) · 2024
Later among the works it cites.
Uncertainty-aware reward model: Teaching reward models to know what is unknown
Lou, X., Yan, D., Shen, W., Yan, Y., Xie, J., and Zhang, J. (2024) · 2024
Later among the works it cites.
Introducing meta llama 3: The most capable openly available llm to date
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reward model ensembles help mitigate overoptimization
Coste, T., Anwar, U., Kirk, R., and Krueger, D. (2023) · 2023
Cited alongside, same era.
Raft: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T. (2023) · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J. (2023) · 2023
Cited alongside, same era.
Statistical rejection sampling improves preference optimization
Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J. (2023) · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Cited alongside, same era.
Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf
Xiong, W., Dong, H., Ye, C., Zhong, H., Jiang, N., and Zhang, T. (2023) · 2023
Cited alongside, same era.
Meta, A. (2024) · 2024
Later among the works it cites.
Optimal design for human preference elicitation
Mukherjee, S., Lalitha, A., Kalantari, K., Deshmukh, A., Liu, G., Ma, Y., and Kveton, B. (2024) · 2024
Later among the works it cites.
Active preference learning for large language models
Muldrew, W., Hayes, P., Zhang, M., and Barber, D. (2024) · 2024
Later among the works it cites.
Sun, H., Shen, Y., and Ton, J.-F. (2024) · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al. (2024) · 2024
Later among the works it cites.
Snorkel-mistral-pairrm-dpo
Tran, H. and Chris Glaze, B. (2024) · 2024
Later among the works it cites.
Metametrics: Calibrating metrics for generation tasks using human preferences
Winata, G. I., Anugraha, D., Susanto, L., Kuwanto, G., and Wijaya, D. T. (2024) · 2024
Later among the works it cites.
Yang, R., Pan, X., Luo, F., Qiu, S., Zhong, H., Yu, D., and Chen, J. (2024) · 2024
Later among the works it cites.
Yin, Y., Wang, Z., Gu, Y., Huang, H., Chen, W., and Zhou, M. (2024) · 2024
Later among the works it cites.
Zhang, X., Ton, J.-F., Shen, W., Wang, H., and Liu, Y. (2024) · 2024
Later among the works it cites.