Fetching the paper…
Reading the bibliography…
This paper studies the alignment process of generative models with Reinforcement Learning from Human Feedback (RLHF).
On the weaknesses of reinforcement learning for neural machine translation
Choshen, L., Fox, L., Aizenbud, Z., and Abend, O. (2019) · 1907
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019) · 1909
Earlier work this paper cites.
Various techniques used in connection with random digits
Neumann, V. (1951) · 1951
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E. (1952) · 1952
Earlier work this paper cites.
Convergence analysis of a proximal-like minimization algorithm using bregman functions
Chen, G. and Teboulle, M. (1993) · 1993
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P. (2002) · 2002
Earlier work this paper cites.
Implementation matters in deep policy gradients: A case study on ppo and trpo
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A. (2020) · 2005
Earlier work this paper cites.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, J. and Zhang, T. (2007) · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V., Hayes, T. P., and Kakade, S. M. (2008) · 2008
Earlier work this paper cites.
Linearly parameterized bandits
Rusmevichientong, P. and Tsitsiklis, J. N. (2010) · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
Understanding learned reward functions
Michaud, E. J., Gleave, A., and Russell, S. (2020) · 2012
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y., Broder, J., Kleinberg, R., and Joachims, T. (2012) · 2012
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Russo, D. and Van Roy, B. (2013) · 2013
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., Fürnkranz, J., et al. (2017) · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: Theory and algorithms
Agarwal, A., Jiang, N., Kakade, S. M., and Sun, W. (2019) · 2019
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z. (2020) · 2020
Earlier work this paper cites.
Improved optimistic algorithms for logistic bandits
Faury, L., Abeille, M., Calauzènes, C., and Fercoq, O. (2020) · 2020
Earlier work this paper cites.
Dueling posterior sampling for preference-based reinforcement learning
Novoseller, E., Wei, Y., Sui, Y., Yue, Y., and Burdick, J. (2020) · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., and Huang, S. (2020) · 2020
Earlier work this paper cites.
Preference-based reinforcement learning with finite-time guarantees
Xu, Y., Wang, R., Yang, L., Singh, A., and Dubrawski, A. (2020) · 2020
Earlier work this paper cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2021) · 2021
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. (2021) · 2021
Earlier work this paper cites.
Preference-based online learning with dueling bandits: A survey
Bengs, V., Busa-Fekete, R., El Mesaoudi-Paul, A., and Hüllermeier, E. (2021) · 2021
Earlier work this paper cites.
Towards general function approximation in zero-sum markov games
Huang, B., Lee, J. D., Wang, Z., and Yang, Z. (2021) · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. (2021) · 2021
Earlier work this paper cites.
Dueling rl: reinforcement learning with trajectory preferences
Pacchiano, A., Saha, A., and Lee, J. (2021) · 2021
Cited alongside, same era.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2021) · 2021
Cited alongside, same era.
Optimal algorithms for stochastic contextual preference bandits
Saha, A. (2021) · 2021
Cited alongside, same era.
Batch value-function approximation with only realizability
Xie, T. and Jiang, N. (2021) · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. (2022) · 2022
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M. (2023) · 2023
Closest in time.
Lmflow: An extensible toolkit for finetuning and inference of large foundation models
Diao, S., Pan, R., Dong, H., Shum, K. S., Zhang, J., Xiong, W., and Zhang, T. (2023) · 2023
Closest in time.
RAFT: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., SHUM, K., and Zhang, T. (2023) · 2023
Closest in time.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J. (2023) · 2023
Closest in time.
Openllama: An open reproduction of llama
Geng, X. and Liu, H. (2023) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Chen, X., Zhong, H., Yang, Z., Wang, Z., and Wang, L. (2022) · 2022
Cited alongside, same era.
Fast rates in pool-based batch active learning
Gentile, C., Wang, Z., and Zhang, T. (2022) · 2022
Cited alongside, same era.
Optimizing prompts for text-to-image generation
Hao, Y., Chi, Z., Dong, L., and Wei, F. (2022) · 2022
Cited alongside, same era.
On the sensitivity of reward inference to misspecified human models
Hong, J., Bhatia, K., and Dragan, A. (2022) · 2022
Cited alongside, same era.
Nearly minimax optimal reinforcement learning with linear function approximation
Hu, P., Chen, Y., and Huang, L. (2022) · 2022
Cited alongside, same era.
Provably feedback-efficient reinforcement learning via active reward learning
Kong, D. and Yang, L. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Cited alongside, same era.
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al. (2023) · 2023
Closest in time.
Aligning text-to-image models using human feedback
Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S. (2023) · 2023
Closest in time.
OpenAI (2023) · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. (2023) · 2023
Closest in time.
Tiapkin, D., Belomestny, D., Calandriello, D., Moulines, E., Naumov, A., Perrault, P., Valko, M., and Menard, P. (2023) · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Closest in time.
Zephyr: Direct distillation of lm alignment
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., et al. (2023) · 2023
Closest in time.
Making rl with preference-based feedback efficient via randomization
Wu, R. and Sun, W. (2023) · 2023
Closest in time.
Better aligning text-to-image models with human preference
Wu, X., Sun, K., Zhu, F., Zhao, R., and Li, H. (2023) · 2023
Closest in time.
Corruption-robust algorithms with uncertainty weighting for nonlinear contextual bandits and markov decision processes
Ye, C., Xiong, W., Gu, Q., and Zhang, T. (2023) · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F. (2023) · 2023
Closest in time.
Mathematical Analysis of Machine Learning Algorithms
Zhang, T. (2023) · 2023
Closest in time.
Slic-hf: Sequence likelihood calibration with human feedback
Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J. (2023) · 2023
Closest in time.
Zhong, H. and Zhang, T. (2023) · 2023
Closest in time.
Offline data enhanced on-policy policy gradient with provable guarantees
Zhou, Y., Sekhari, A., Song, Y., and Sun, W. (2023) · 2023
Closest in time.
A novel framework for policy mirror descent with general parameterization and linear convergence
Alfano, C., Yuan, R., and Rebeschini, P. (2024) · 2024
Closest in time.
Theoretical guarantees on the best-of-n alignment policy
Beirami, A., Agarwal, A., Berant, J., D’Amour, A., Eisenstein, J., Nagpal, C., and Suresh, A. T. (2024) · 2024
Closest in time.
Noise contrastive alignment of language models with explicit rewards
Chen, H., He, G., Su, H., and Zhu, J. (2024) · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D. (2024) · 2024
Closest in time.
Snorkel-mistral-pairrm-dpo
Hoang Tran, Chris Glaze, B. H. (2024) · 2024
Closest in time.
Offline minimax soft-q-learning under realizability and partial coverage
Uehara, M., Kallus, N., Lee, J. D., and Sun, W. (2024) · 2024
Closest in time.
Mint: Multi-turn interactive evaluation for tool-augmented llms with language feedback
Wang, X., Wang, Z., Liu, J., Chen, Y., Yuan, L., Peng, H., and Ji, H. (2024) · 2024
Closest in time.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J. (2024) · 2024
Closest in time.