Fetching the paper…
Reading the bibliography…
Fine-grained control over large language models (LLMs) remains a significant challenge, hindering their adaptability to diverse user needs.
Intransitivity, utility, and the aggregation of preference patterns
K. O. May · 1954
Earlier work this paper cites.
Mathematics and Social Sciences: Proceedings of the Seminars of Menthon-Saint-Bernard, France (1-27 July, 1960) and of Gösing, Austria (3-27 July, 1961) , volume 1
S. H. Sternberg · 1965
Earlier work this paper cites.
Intransitivity of preferences
A. Tversky · 1969
Earlier work this paper cites.
Probabilistic social choice based on simple voting comparisons
P. C. Fishburn · 1984
Earlier work this paper cites.
Multitask learning
R. Caruana · 1997
Earlier work this paper cites.
Handling preferences in evolutionary multiobjective optimization: A survey
C. C. Coello · 2000
Earlier work this paper cites.
Condorcet’s paradox and the likelihood of its occurrence: different perspectives on balanced preferences
W. V. Gehrlein · 2002
Earlier work this paper cites.
A new scalarization method for finding the efficient frontier in non-convex multi-objective problems
A. Ghane-Kanafi and E. Khorram · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Batch active preference-based learning of reward functions
E. Biyik and D. Sadigh · 2018
Earlier work this paper cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
D. Brown, W. Goo, P. Nagarajan, and S. Niekum · 2019
Earlier work this paper cites.
On the weaknesses of reinforcement learning for neural machine translation
L. Choshen, L. Fox, Z. Aizenbud, and O. Abend · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Earlier work this paper cites.
Online learning to rank for sequential music recommendation
B. L. Pereira, A. Ueda, G. Penha, R. L. Santos, and N. Ziviani · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Implementation matters in deep policy gradients: A case study on ppo and trpo
L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, and S. Huang · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, et al · 2021
Earlier work this paper cites.
Fine-tuning language models to find agreement among humans with diverse preferences
M. Bakker, M. Chadwick, H. Sheahan, M. Tessler, L. Campbell-Gillingham, J. Balaguer, N. McAleese, A. Glaese, J. Aslanides, M. Botvinick, et al · 2022
Earlier work this paper cites.
Optimizing prompts for text-to-image generation
Y. Hao, Z. Chi, L. Dong, and F. Wei · 2022
Earlier work this paper cites.
Holistic evaluation of language models
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V. Korthikanti, et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
H. Lee, S. Phatale, H. Mansoor, K. Lu, T. Mesnard, C. Bishop, V. Carbune, and A. Rastogi · 2023
Later among the works it cites.
Alpacaeval: An automatic evaluator of instruction-following models
X. Li, T. Zhang, Y. Dubois, R. Taori, I. Gulrajani, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Y. Lin, L. Tan, H. Lin, Z. Zheng, R. Pi, J. Zhang, S. Diao, H. Wang, H. Zhao, Y. Yao, et al · 2023
Later among the works it cites.
Nash learning from human feedback
R. Munos, M. Valko, D. Calandriello, M. G. Azar, M. Rowland, Z. D. Guo, Y. Tang, M. Geist, T. Mesnard, A. Michi, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Workshop, T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, et al · 2022
Cited alongside, same era.
Introducing claude
Anthropic · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, M. Rowland, B. Piot, D. Guo, D. Calandriello, M. Valko, and R. Munos · 2023
Cited alongside, same era.
Peering through preferences: Unraveling feedback acquisition for aligning large language models
H. Bansal, J. Dang, and A. Grover · 2023
Cited alongside, same era.
Aligning robot and human representations
A. Bobu, A. Peng, P. Agrawal, J. Shah, and A. D. Dragan · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2023
Cited alongside, same era.
OpenAI · 2023
Later among the works it cites.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
A. Pan, J. S. Chan, A. Zou, N. Li, S. Basart, T. Woodside, H. Zhang, S. Emmons, and D. Hendrycks · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Later among the works it cites.
A. Rame, G. Couairon, M. Shukor, C. Dancette, J.-B. Gaya, L. Soulier, and M. Cord · 2023
Later among the works it cites.
Verbosity bias in preference labeling by large language models
K. Saito, A. Wachi, K. Wataoka, and Y. Akimoto · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Large language models in medicine
A. J. Thirunavukarasu, D. S. J. Ting, K. Elangovan, L. Gutierrez, T. F. Tan, and D. S. W. Ting · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
L. Tunstall, E. Beeching, N. Lambert, N. Rajani, K. Rasul, Y. Belkada, S. Huang, L. von Werra, C. Fourrier, N. Habib, et al · 2023
Later among the works it cites.
Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf
W. Xiong, H. Dong, C. Ye, H. Zhong, N. Jiang, and T. Zhang · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Z. Yuan, H. Yuan, C. Tan, W. Wang, S. Huang, and F. Huang · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Y. Zhao, R. Joshi, T. Liu, M. Khalman, M. Saleh, and P. J. Liu · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Later among the works it cites.
Beyond one-preference-for-all: Multi-objective direct preference optimization
Z. Zhou, J. Liu, C. Yang, J. Shao, Y. Liu, X. Yue, W. Ouyang, and Y. Qiao · 2023
Later among the works it cites.
Odin: Disentangled reward mitigates hacking in rlhf, 2024
L. Chen, C. Zhu, D. Soselia, J. Chen, T. Zhou, T. Goldstein, H. Huang, M. Shoeybi, and B. Catanzaro · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
G. Swamy, C. Dann, R. Kidambi, Z. S. Wu, and A. Agarwal · 2024
Closest in time.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
C. Ye, W. Xiong, Y. Zhang, N. Jiang, and T. Zhang · 2024
Closest in time.
Self-rewarding language models
W. Yuan, R. Y. Pang, K. Cho, S. Sukhbaatar, J. Xu, and J. Weston · 2024
Closest in time.