Fetching the paper…
Reading the bibliography…
This work studies the challenge of aligning large language models (LLMs) with offline preference data.
Emergent tool use from multi-agent autocurricula
Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., and Mordatch, I. (2019) · 1909
Earlier work this paper cites.
Iterative solution of games by fictitious play
Brown, G. W. (1951) · 1951
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E. (1952) · 1952
Earlier work this paper cites.
Empirical processes in m-estimation
van de Geer, S. A. (2000) · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvari, C., and Munos, R. (2006) · 2006
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Syed, U. and Schapire, R. E. (2007) · 2007
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C. (2008) · 2008
Earlier work this paper cites.
Online markov decision processes
Even-Dar, E., Kakade, S. M., and Mansour, Y. (2009) · 2009
Earlier work this paper cites.
Market Structure and Equilibrium
von Stackelberg, H., Bazin, D., Hill, R., and Urch, L. (2010) · 2010
Earlier work this paper cites.
What are the statistical limits of offline rl with linear function approximation?
Wang, R., Foster, D. P., and Kakade, S. M. (2020) · 2010
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K. (2012) · 2012
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Contextual dueling bandits
Dudík, M., Hofmann, K., Schapire, R. E., Slivkins, A., and Zoghi, M. (2015) · 2015
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Heinrich, J., Lanctot, M., and Silver, D. (2015) · 2015
Earlier work this paper cites.
Introduction to online convex optimization
Hazan, E. et al. (2016) · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S. (2016) · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016) · 2016
Earlier work this paper cites.
Emergent complexity via multi-agent competition
Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I. (2017) · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Azuma, K. (2018) · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. (2018) · 2018
Earlier work this paper cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
Vershynin, R. (2018) · 2018
Earlier work this paper cites.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Earlier work this paper cites.
WINOGRANDE: an adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. (2019) · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 2019
Earlier work this paper cites.
Flambe: Structural complexity and representation learning of low rank mdps
Agarwal, A., Kakade, S., Krishnamurthy, A., and Sun, W. (2020) · 2020
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z. (2020) · 2020
Cited alongside, same era.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y., Jia, Z., and Wang, M. (2020) · 2020
Cited alongside, same era.
A theoretical analysis of deep q-learning
Fan, J., Wang, Z., Xie, Y., and Yang, Z. (2020) · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020) · 2020
Cited alongside, same era.
Dueling posterior sampling for preference-based reinforcement learning
Novoseller, E., Wei, Y., Sui, Y., Yue, Y., and Burdick, J. (2020) · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020) · 2020
Open llm leaderboard
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T. (2023) · 2023
Later among the works it cites.
Adversarial model for offline reinforcement learning
Bhardwaj, M., Xie, T., Boots, B., Jiang, N., and Cheng, C.-A. (2023) · 2023
Later among the works it cites.
Ultrafeedback: Boosting language models with high-quality feedback
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M. (2023) · 2023
Later among the works it cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B. (2023) · 2023
Later among the works it cites.
Contrastive prefence learning: Learning from human feedback without rl
Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Preference-based reinforcement learning with finite-time guarantees
Xu, Y., Wang, R., Yang, L., Singh, A., and Dubrawski, A. (2020) · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2021) · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. (2021) · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S. (2021) · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2021) · 2021
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z. (2021) · 2021
Cited alongside, same era.
Li, Y., Bubeck, S., Eldan, R., Del Giorno, A., Gunasekar, S., and Lee, Y. T. (2023) · 2023
Later among the works it cites.
Kernelized offline contextual dueling bandits
Mehta, V., Neopane, O., Das, V., Lin, S., Schneider, J., and Neiswanger, W. (2023) · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Later among the works it cites.
Iterative dpo alignment
Tran, H., Glaze, C., and Hancock, B. (2023) · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment
Tunstall, L., Beeching, E., Lambert, N., Rajani, N., Rasul, K., Belkada, Y., Huang, S., von Werra, L., Fourrier, C., Habib, N., Sarrazin, N., Sanseviero, O., Rush, A. M., and Wolf, T. (2023) · 2023
Later among the works it cites.
Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf
Xiong, W., Dong, H., Ye, C., Zhong, H., Jiang, N., and Zhang, T. (2023) · 2023
Later among the works it cites.
The efficacy of pessimism in asynchronous q-learning
Yan, Y., Li, G., Chen, Y., and Fan, J. (2023) · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Zhu, B., Jiao, J., and Jordan, M. I. (2023) · 2023
Later among the works it cites.
Human alignment of large language models through online preference optimisation
Calandriello, D., Guo, D., Munos, R., Rowland, M., Tang, Y., Pires, B. A., Richemond, P. H., Lan, C. L., Valko, M., Liu, T., et al. (2024) · 2024
Closest in time.
PARL: A unified framework for policy alignment in reinforcement learning from human feedback
Chakraborty, S., Bedi, A., Koppel, A., Wang, H., Manocha, D., Wang, M., and Huang, F. (2024) · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models
Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q. (2024) · 2024
Closest in time.
Efficient exploration for llms
Dwaracherla, V., Asghari, S. M., Hao, B., and Van Roy, B. (2024) · 2024
Closest in time.
Rebel: Reinforcement learning via regressing relative rewards
Gao, Z., Chang, J. D., Zhan, W., Oertell, O., Swamy, G., Brantley, K., Joachims, T., Bagnell, J. A., Lee, J. D., and Sun, W. (2024) · 2024
Closest in time.
Mitigating reward hacking via information-theoretic reward modeling
Miao, Y., Zhang, S., Ding, L., Bao, R., Zhang, L., and Tao, D. (2024) · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. (2024) · 2024
Closest in time.
Warm: On the benefits of weight averaged reward models
Ramé, A., Vieillard, N., Hussenot, L., Dadashi, R., Cideron, G., Bachem, O., and Ferret, J. (2024) · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C., Cheng, C.-A., Mitra, A., Santacroce, M., Awadallah, A., and Xie, T. (2024) · 2024
Closest in time.
Principled penalty-based methods for bilevel reinforcement learning and rlhf
Shen, H., Yang, Z., and Chen, T. (2024) · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Swamy, G., Dann, C., Kidambi, R., Wu, Z. S., and Agarwal, A. (2024) · 2024
Closest in time.
Transforming and combining rewards for aligning large language models
Wang, Z., Nagpal, C., Berant, J., Eisenstein, J., D’Amour, A., Koyejo, S., and Veitch, V. (2024) · 2024
Closest in time.
Xu, H., Sharaf, A., Chen, Y., Tan, W., Shen, L., Van Durme, B., Murray, K., and Kim, Y. J. (2024) · 2024
Closest in time.
Provable offline preference-based reinforcement learning
Zhan, W., Uehara, M., Kallus, N., Lee, J. D., and Sun, W. (2024) · 2024
Closest in time.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D. (2019) · 2062
Closest in time.