Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) has become an essential step in fine-tuning large language models (LLMs) to align them with human preferences.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
T. Roughgarden and O. Schrijvers, “Online prediction with selfish experts,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
M. Völske, M. Potthast, S. Syed, and B. Stein, “Tl; dr: Mining reddit to learn automatic summarization,” in Proceedings of the Workshop on New Frontiers in Summarization , 2017, pp. 59–63
2017
Earlier work this paper cites.
R. Freeman, D. Pennock, C. Podimata, and J. W. Vaughan, “No-regret and incentive-compatible online learning,” in International Conference on Machine Learning . PMLR, 2020, pp. 3270–3279
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
R. Frongillo, R. Gomez, A. Thilagar, and B. Waggoner, “Efficient competitions and online learning with strategic forecasters,” in Proceedings of the 22nd ACM Conference on Economics and Computation , 2021, pp. 479–496
2021
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Asadi, A. Bellet, O.-A. Maillard, and M. Tommasi, “Collaborative algorithms for online personalized mean estimation,” Transactions on Machine Learning Research Journal , 2022
2022
Earlier work this paper cites.
X. Chen, H. Zhong, Z. Yang, Z. Wang, and L. Wang, “Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation,” in International Conference on Machine Learning . PMLR, 2022, pp. 3773–3793
2022
Earlier work this paper cites.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al. , “Open and efficient foundation language models,” Preprint at arXiv. https://doi. org/10.48550/arXiv , vol. 2302, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Havrilla, M. Zhuravinskyi, D. Phung, A. Tiwari, J. Tow, S. Biderman, Q. Anthony, and L. Castricato, “trlx: A framework for large scale reinforcement learning from human feedback,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 8578–8595
2023
Cited alongside, same era.
A. Köpf, Y. Kilcher, D. von Rütte, S. Anagnostidis, Z. R. Tam, K. Stevens, A. Barhoum, D. Nguyen, O. Stanley, R. Nagyfi et al. , “Openassistant conversations-democratizing large language model alignment,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
C. Park, M. Liu, D. Kong, K. Zhang, and A. E. Ozdaglar, “Rlhf from heterogeneous feedback via personalization and preference aggregation,” in ICML 2024 Workshop on Theoretical Foundations of Foundation Models , 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
W. Xiong, H. Dong, C. Ye, Z. Wang, H. Zhong, H. Ji, N. Jiang, and T. Zhang, “Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint,” in Forty-first International Conference on Machine Learning , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
Y. Chen, J. Zhu, and K. Kandasamy, “Mechanism design for collaborative normal mean estimation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
J. Li, M. Li, and H. Chan, “Strategyproof mechanisms for group-fair obnoxious facility location problems,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 9, 2024, pp. 9832–9839
2024
Closest in time.
Y. Wang, H. Zhou, and M. Li, “Positive intra-group externalities in facility location,” in Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems , 2024, pp. 1883–1891
2024
Closest in time.