Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) provides a principled framework for aligning AI systems with human preference data.
Rank analysis of incomplete block designs: I. the method of paired comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Robust truncated hinge loss support vector machines
Y. Wu and Y. Liu · 2007
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Learning with noisy labels
N. Natarajan, I. S. Dhillon, P. K. Ravikumar, and A. Tewari · 2013
Earlier work this paper cites.
Perspectives on crowdsourcing annotations for natural language processing
A. Wang, C. D. V. Hoang, and M.-Y. Kan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Robust classification under sample selection bias
A. Liu and B. Ziebart · 2014
Earlier work this paper cites.
Corpus annotation through crowdsourcing: Towards best practice guidelines
M. Sabou, K. Bontcheva, L. Derczynski, and A. Scharl · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Distributionally robust logistic regression
S. Shafieezadeh Abadeh, P. M. Mohajerin Esfahani, and D. Kuhn · 2015
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Adversarial machine learning at scale
A. Kurakin, I. Goodfellow, and S. Bengio · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Tl;dr: Mining reddit to learn automatic summarization
M. Völske, M. Potthast, S. Syed, and B. Stein · 2017
Earlier work this paper cites.
Cognitive biases in crowdsourcing
C. Eickhoff · 2018
Cited alongside, same era.
Co-teaching: Robust training of deep neural networks with extremely noisy labels
B. Han, Q. Yao, X. Yu, G. Niu, M. Xu, W. Hu, I. Tsang, and M. Sugiyama · 2018
Cited alongside, same era.
Virtual adversarial training: a regularization method for supervised and semi-supervised learning
T. Miyato, S.-i. Maeda, M. Koyama, and S. Ishii · 2018
Cited alongside, same era.
A general and adaptive robust loss function
J. T. Barron · 2019
Cited alongside, same era.
Pybullet, a python module for physics simulation for games, robotics and machine learning
E. Coumans and Y. Bai · 2019
Cited alongside, same era.
Theoretically principled trade-off between robustness and accuracy
H. Zhang, Y. Yu, J. Jiao, E. Xing, L. El Ghaoui, and M. Jordan · 2019
Claude, 2023
Anthropic · 2023
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, M. Rowland, B. Piot, D. Guo, D. Calandriello, M. Valko, and R. Munos · 2023
Later among the works it cites.
A universal law of robustness via isoperimetry
S. Bubeck and M. Sellke · 2023
Later among the works it cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, et al · 2023
Later among the works it cites.
Reward model ensembles help mitigate overoptimization
T. Coste, U. Anwar, R. Kirk, and D. Krueger · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Doubly robust off-policy learning on low-dimensional manifolds by deep neural networks
M. Chen, H. Liu, W. Liao, and T. Zhao · 2020
Cited alongside, same era.
Rl baselines3 zoo
A. Raffin · 2020
Cited alongside, same era.
Deep reinforcement learning with robust and smooth policy
Q. Shen, Y. Li, H. Jiang, Z. Wang, and T. Zhao · 2020
Cited alongside, same era.
Learning to summarize from human feedback
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Trl: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, and S. Huang · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Cited alongside, same era.
Y. Dubois, X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Scaling laws for reward model overoptimization
L. Gao, J. Schulman, and J. Hilton · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Later among the works it cites.
Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf
W. Xiong, H. Dong, C. Ye, H. Zhong, N. Jiang, and T. Zhang · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Y. Zhao, R. Joshi, T. Liu, M. Khalman, M. Saleh, and P. J. Liu · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
B. Zhu, M. Jordan, and J. Jiao · 2023
Later among the works it cites.
Robust multi-agent reinforcement learning via adversarial regularization: Theoretical foundation and stable algorithms
A. Bukharin, Y. Li, Y. Yu, Q. Zhang, Z. Chen, S. Zuo, C. Zhang, S. Zhang, and T. Zhao · 2024
Closest in time.
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences
S. Chakraborty, J. Qiu, H. Yuan, A. Koppel, F. Huang, D. Manocha, A. S. Bedi, and M. Wang · 2024
Closest in time.
Rime: Robust preference-based reinforcement learning with noisy preferences
J. Cheng, G. Xiong, X. Dai, Q. Miao, Y. Lv, and F.-Y. Wang · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela · 2024
Closest in time.
Corruption robust offline reinforcement learning with human feedback
D. Mandal, A. Nika, P. Kamalaruban, A. Singla, and G. Radanović · 2024
Closest in time.