Fetching the paper…
Reading the bibliography…
In this report, we introduce a collection of methods to enhance reward modeling for LLMs, focusing specifically on data-centric techniques.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
The Elements of Statistical Learning
J. H. Friedman, T. Hastie, and R. Tibshirani · 2001
Earlier work this paper cites.
Kernel methods for pattern analysis
B. Schölkopf, A. J. Smola, K. R. Müller, P. J. Bartlett, W. S. Davidson, D. C. P. J. M. Williamson, and R. C. R. Schölkopf · 2001
Earlier work this paper cites.
The dangers of inference using the bradley-terry model
C. R. Carvalho, A. D. Polson, and J. G. Scott · 2010
Earlier work this paper cites.
Do user preferences and evaluation measures line up?
M. Sanderson, M. L. Paramita, P. Clough, and E. Kanoulas · 2010
Earlier work this paper cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
Focal loss for dense object detection
T. Lin · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Supervised contrastive learning
P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan · 2020
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information
K. Ethayarajh, Y. Choi, and S. Swayamdipta · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, et al · 2023
Earlier work this paper cites.
Ultrafeedback: Boosting language models with high-quality feedback
G. Cui, L. Yuan, N. Ding, G. Yao, W. Zhu, Y. Ni, G. Xie, Z. Liu, and M. Sun · 2023
Earlier work this paper cites.
Amplify-instruct: Synthetically generated diverse multi-turn conversations for effecient llm training
L. Daniele and Suphavadeeprasit · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2023
Cited alongside, same era.
Camels in a changing climate: Enhancing lm adaptation with tulu 2, 2023
H. Ivison, Y. Wang, V. Pyatkin, N. Lambert, M. Peters, P. Dasigi, J. Jang, D. Wadden, N. A. Smith, I. Beltagy, and H. Hajishirzi · 2023
Cited alongside, same era.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
D. Jiang, X. Ren, and B. Y. Lin · 2023
Cited alongside, same era.
Openorca: An open dataset of gpt augmented flan reasoning traces
W. Lian, B. Goodson, E. Pentland, A. Cook, C. Vong, and "Teknium" · 2023
Cited alongside, same era.
Z. Cai, M. Cao, H. Chen, K. Chen, K. Chen, X. Chen, X. Chen, Z. Chen, Z. Chen, P. Chu, et al · 2024
Closest in time.
Rlhf workflow: From reward modeling to online rlhf
H. Dong, W. Xiong, B. Pang, H. Wang, H. Zhao, Y. Zhou, N. Jiang, D. Sahoo, C. Xiong, and T. Zhang · 2024
Closest in time.
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Closest in time.
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
S. Han, K. Rao, A. Ettinger, L. Jiang, B. Y. Lin, N. Lambert, Y. Choi, and N. Dziri · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
RyokoAI · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Cited alongside, same era.
Helpsteer: Multi-attribute helpfulness dataset for steerlm
Z. Wang, Y. Dong, J. Zeng, V. Adams, M. N. Sreedhar, D. Egert, O. Delalleau, J. P. Scowcroft, N. Kant, A. Swope, et al · 2023
Cited alongside, same era.
Evaluating large language models at evaluating instruction following
Z. Zeng, J. Yu, T. Gao, Y. Meng, T. Goyal, and D. Chen · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Cited alongside, same era.
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang · 2024
Closest in time.
Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models
L. Jiang, K. Rao, S. Han, A. Ettinger, F. Brahman, S. Kumar, N. Mireshghallah, X. Lu, M. Sap, Y. Choi, et al · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
N. Lambert, V. Pyatkin, J. Morrison, L. Miranda, B. Y. Lin, K. Chandu, N. Dziri, S. Kumar, T. Zick, Y. Choi, et al · 2024
Closest in time.
Uncertainty-aware reward model: Teaching reward models to know what is unknown
X. Lou, D. Yan, W. Shen, Y. Yan, J. Xie, and J. Zhang · 2024
Closest in time.
Offsetbias: Leveraging debiased data for tuning evaluators
J. Park, S. Jwa, M. Ren, D. Kim, and S. Choi · 2024
Closest in time.
Metametrics: Calibrating metrics for generation tasks using human preferences
G. I. Winata, D. Anugraha, L. Susanto, G. Kuwanto, and D. T. Wijaya · 2024
Closest in time.
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing
Z. Xu, F. Jiang, L. Niu, Y. Deng, R. Poovendran, Y. Choi, and B. Y. Lin · 2024
Closest in time.
Regularizing hidden states enables learning generalizable reward model for llms
R. Yang, R. Ding, Y. Lin, H. Zhang, and T. Zhang · 2024
Closest in time.
Advancing llm reasoning generalists with preference trees
L. Yuan, G. Cui, H. Wang, N. Ding, X. Wang, J. Deng, B. Shan, H. Chen, R. Xie, Y. Lin, et al · 2024
Closest in time.
General preference modeling with preference representations for aligning language models
Y. Zhang, G. Zhang, Y. Wu, K. Xu, and Q. Gu · 2024
Closest in time.