Fetching the paper…
Reading the bibliography…
Accurately aligning large language models (LLMs) with human preferences is crucial for informing fair, economically sound, and statistically efficient decision-making processes.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
On measures of entropy and information
A. Rényi · 1961
Earlier work this paper cites.
The analysis of permutations
R. L. Plackett · 1975
Earlier work this paper cites.
Perplexity—a measure of the difficulty of speech recognition tasks
F. Jelinek, R. L. Mercer, L. R. Bahl, and J. K. Baker · 1977
Earlier work this paper cites.
Tokenization
G. Grefenstette · 1999
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey, et al · 2008
Earlier work this paper cites.
Social choice and individual values , volume 12
K. J. Arrow · 2012
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
R. D. Luce · 2012
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. D. M.-W. C. Kenton and L. K. Toutanova · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Maximum entropy rl (provably) solves some robust rl problems
B. Eysenbach and S. Levine · 2021
Earlier work this paper cites.
The state of online harassment
E. A. Vogels · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Earlier work this paper cites.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
M. Sap, S. Swayamdipta, L. Vianna, X. Zhou, Y. Choi, and N. A. Smith · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, et al · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2023
Cited alongside, same era.
Mechanism design for large language models
P. Duetting, V. Mirrokni, R. P. Leme, H. Xu, and S. Zuo · 2023
Cited alongside, same era.
Gpts are gpts: An early look at the labor market impact potential of large language models
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, Z. D. Guo, B. Piot, R. Munos, M. Rowland, M. Valko, and D. Calandriello · 2024
Closest in time.
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences
S. Chakraborty, J. Qiu, H. Yuan, A. Koppel, F. Huang, D. Manocha, A. S. Bedi, and M. Wang · 2024
Closest in time.
Dataset reset policy optimization for rlhf
J. D. Chang, W. Shan, O. Oertell, K. Brantley, D. Misra, J. D. Lee, and W. Sun · 2024
Closest in time.
Rlhf workflow: From reward modeling to online rlhf
H. Dong, W. Xiong, B. Pang, H. Wang, H. Zhao, Y. Zhou, N. Jiang, D. Sahoo, C. Xiong, and T. Zhang · 2024
Closest in time.
Learn your reference model for real good alignment
A. Gorbatovski, B. Shaposhnikov, A. Malakhov, N. Surnachev, Y. Aksenov, I. Maksimov, N. Balagansky, and D. Gavrilov · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Eloundou, S. Manning, P. Mishkin, and D. Rock · 2023
Cited alongside, same era.
A survey of reinforcement learning from human feedback, 2023
T. Kaufmann, P. Weng, V. Bengs, and E. Hüllermeier · 2023
Cited alongside, same era.
Statistical rejection sampling improves preference optimization
T. Liu, Y. Zhao, R. Joshi, M. Khalman, M. Saleh, P. J. Liu, and J. Liu · 2023
Cited alongside, same era.
Nash learning from human feedback
R. Munos, M. Valko, D. Calandriello, M. G. Azar, M. Rowland, Z. D. Guo, Y. Tang, M. Geist, T. Mesnard, A. Michi, et al · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model, 2023
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Cited alongside, same era.
Why don’t you do it right? analysing annotators’ disagreement in subjective tasks
M. Sandri, E. Leonardelli, S. Tonelli, and E. Jezek · 2023
Cited alongside, same era.
Whose opinions do language models reflect?
S. Santurkar, E. Durmus, F. Ladhak, C. Lee, P. Liang, and T. Hashimoto · 2023
Cited alongside, same era.
Closest in time.
W. Liang, Z. Izzo, Y. Zhang, H. Lepp, H. Cao, X. Zhao, L. Chen, H. Ye, S. Liu, Z. Huang, et al · 2024
Closest in time.
Disentangling length from quality in direct preference optimization
R. Park, R. Rafailov, S. Ermon, and C. Finn · 2024
Closest in time.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
F. Tajwar, A. Singh, A. Sharma, R. Rafailov, J. Schneider, T. Xie, S. Ermon, C. Finn, and A. Kumar · 2024
Closest in time.
Understanding the performance gap between online and offline alignment algorithms
Y. Tang, D. Z. Guo, Z. Zheng, D. Calandriello, Y. Cao, E. Tarassov, R. Munos, B. Á. Pires, M. Valko, Y. Cheng, et al · 2024
Closest in time.
H. Wang, Y. Lin, W. Xiong, R. Yang, S. Diao, S. Qiu, H. Zhao, and T. Zhang · 2024
Closest in time.
Rlhf and iia: Perverse incentives
W. Xu, S. Dong, X. Lu, G. Lam, Z. Wen, and B. Van Roy · 2024
Closest in time.
Asymptotics of language model alignment
J. Q. Yang, S. Salamatian, Z. Sun, A. T. Suresh, and A. Beirami · 2024
Closest in time.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
C. Ye, W. Xiong, Y. Zhang, N. Jiang, and T. Zhang · 2024
Closest in time.
Provable multi-party reinforcement learning with diverse human feedback
H. Zhong, Z. Deng, W. J. Su, Z. S. Wu, and L. Zhang · 2024
Closest in time.
Fine-tuning attention modules only: Enhancing weight disentanglement in task arithmetic
R. Jin, B. Hou, J. Xiao, W. J. Su, and L. Shen · 2025
Closest in time.
Preserving diversity in supervised fine-tuning of large language models
Z. Li, C. Chen, T. Xu, Z. Qin, J. Xiao, Z.-Q. Luo, and R. Sun · 2025
Closest in time.
K. Liu, Q. Long, Z. Shi, W. J. Su, and J. Xiao · 2025
Closest in time.
Fundamental limits of game-theoretic llm alignment: Smith consistency and preference matching
Z. Shi, K. Liu, Q. Long, W. J. Su, and J. Xiao · 2025
Closest in time.
Magnetic preference optimization: Achieving last-iterate convergence for language model alignment
M. Wang, C. Ma, Q. Chen, L. Meng, Y. Han, J. Xiao, Z. Zhang, J. Huo, W. J. Su, and Y. Yang · 2025
Closest in time.