Fetching the paper…
Reading the bibliography…
Reinforcement Learning from Human Feedback (RLHF) has achieved impressive empirical successes while relying on a small amount of human feedback.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2007
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
A tail inequality for quadratic forms of subgaussian random vectors
Hsu, D., Kakade, S., and Zhang, T · 2012
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Tropp, J. A. et al · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S., Du, S., Hu, W., Li, Z., and Wang, R · 2019
Earlier work this paper cites.
Neural temporal-difference learning converges to global optima
Cai, Q., Yang, Z., Lee, J. D., and Wang, Z · 2019
Cited alongside, same era.
Neural trust region/proximal policy optimization attains globally optimal policy
Liu, B., Cai, Q., Yang, Z., and Wang, Z · 2019
Cited alongside, same era.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z · 2019
Cited alongside, same era.
Sample-optimal parametric Q-learning using linearly additive features
Yang, L. and Wang, M · 2019
Cited alongside, same era.
PC-PG: Policy cover directed exploration for provable policy gradient learning
Agarwal, A., Henaff, M., Kakade, S., and Sun, W · 2020
Cited alongside, same era.
A theoretical analysis of deep Q-learning
Fan, J., Wang, Z., Xie, Y., and Yang, Z · 2020
Cautiously optimistic policy optimization and exploration with linear function approximation
Zanette, A., Cheng, C.-A., and Agarwal, A · 2021
Later among the works it cites.
Made: Exploration via maximizing deviation from explored regions
Zhang, T., Rashidinejad, P., Jiao, J., Tian, Y., Gonzalez, J. E., and Russell, S · 2021
Later among the works it cites.
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Chen, X., Zhong, H., Yang, Z., Wang, Z., and Wang, L · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Cited alongside, same era.
Dueling posterior sampling for preference-based reinforcement learning
Novoseller, E., Wei, Y., Sui, Y., Yue, Y., and Burdick, J · 2020
Cited alongside, same era.
Preference-based reinforcement learning with finite-time guarantees
Xu, Y., Wang, R., Yang, L., Singh, A., and Dubrawski, A · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2021
Cited alongside, same era.
Dueling RL: reinforcement learning with trajectory preferences
Pacchiano, A., Saha, A., and Lee, J · 2021
Cited alongside, same era.
CRPO: A new approach for safe reinforcement learning with convergence guarantee
Xu, T., Liang, Y., and Lan, G · 2021
Cited alongside, same era.
Provable benefits of policy learning from human preferences in contextual bandit problems
Ji, X., Wang, H., Chen, M., Zhao, T., and Wang, M · 2023
Later among the works it cites.
A survey of reinforcement learning from human feedback
Kaufmann, T., Weng, P., Bengs, V., and Hüllermeier, E · 2023
Later among the works it cites.
Reinforcement learning with human feedback: Learning dynamic choices via pessimism
Li, Z., Yang, Z., and Wang, M · 2023
Later among the works it cites.
Is RLHF more difficult than standard rl?
Wang, Y., Liu, Q., and Jin, C · 2023
Later among the works it cites.
Making RL with preference-based feedback efficient via randomization
Wu, R. and Sun, W · 2023
Later among the works it cites.
Gibbs sampling from human feedback: A provable KL-constrained framework for RLHF
Xiong, W., Dong, H., Ye, C., Zhong, H., Jiang, N., and Zhang, T · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Zhu, B., Jiao, J., and Jordan, M. I · 2023
Later among the works it cites.