Fetching the paper…
Reading the bibliography…
While originally developed for continuous control problems, Proximal Policy Optimization (PPO) has emerged as the work-horse of a variety of reinforcement learning (RL) applications, including the fine-tuning of generative models.
May, K. O. Intransitivity, utility, and the aggregation of preference patterns. Econometrica: Journal of the Econometric Society 1954
1954
Earlier work this paper cites.
Kreweras, G. Aggregation of preference orderings. Mathematics and Social Sciences I: Proceedings of the seminars of Menthon-Saint-Bernard, France (1–27 July 1960) and of Gösing, Austria (3–27 July 1962). 1965; pp 73–79
1965
Earlier work this paper cites.
Tversky, A. Intransitivity of preferences. Psychological review 1969
1969
Earlier work this paper cites.
Simpson, P. B. On Defining Areas of Voter Choice: Professor Tullock on Stable Voting. The Quarterly Journal of Economics 1969
1969
Earlier work this paper cites.
Gardner, M. Mathematical games. 1970; https://www.scientificamerican.com/article/mathematical-games-1970-12/
1970
Earlier work this paper cites.
Kramer, G. H. On a Class of Equilibrium Conditions for Majority Rule. Econometrica 1973
1973
Earlier work this paper cites.
Nemirovskij, A. S.; Yudin, D. B. Problem complexity and method efficiency in optimization. 1983
1983
Earlier work this paper cites.
Fishburn, P. C. Probabilistic social choice based on simple voting comparisons. The Review of Economic Studies 1984
1984
Earlier work this paper cites.
Williams, R. J. Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning. Mach. Learn. 1992
1992
Earlier work this paper cites.
Konda, V.; Tsitsiklis, J. Actor-Critic Algorithms. Advances in Neural Information Processing Systems. 1999
1999
Earlier work this paper cites.
Kakade, S. M. A Natural Policy Gradient. Advances in Neural Information Processing Systems. 2001
2001
Earlier work this paper cites.
Kakade, S.; Langford, J. Approximately optimal approximate reinforcement learning. Proceedings of the Nineteenth International Conference on Machine Learning. 2002; pp 267–274
2002
Earlier work this paper cites.
Bagnell, J. A.; Schneider, J. Covariant policy search. Proceedings of the 18th international joint conference on Artificial intelligence. 2003; pp 1019–1024
2003
Earlier work this paper cites.
Bagnell, J.; Kakade, S. M.; Schneider, J.; Ng, A. Policy search by dynamic programming. Advances in neural information processing systems 2003
2003
Earlier work this paper cites.
Munos, R. Error bounds for approximate policy iteration. ICML. 2003; pp 560–567
2003
Earlier work this paper cites.
Grünwald, P. D.; Dawid, A. P. Game theory, maximum entropy, minimum discrepancy and robust Bayesian decision theory. 2004
2004
Earlier work this paper cites.
Langford, J.; Zhang, T. The epoch-greedy algorithm for multi-armed bandits with side information. Advances in neural information processing systems 2007
2007
Earlier work this paper cites.
Peters, J.; Schaal, S. Reinforcement learning by reward-weighted regression for operational space control. Proceedings of the 24th international conference on Machine learning. 2007; pp 745–750
2007
Earlier work this paper cites.
Ziebart, B. D.; Maas, A. L.; Bagnell, J. A.; Dey, A. K.; others Maximum entropy inverse reinforcement learning. Aaai. 2008; pp 1433–1438
2008
Earlier work this paper cites.
Ross, S.; Gordon, G.; Bagnell, D. A reduction of imitation learning and structured prediction to no-regret online learning. Proceedings of the fourteenth international conference on artificial intelligence and statistics. 2011; pp 627–635
2011
Earlier work this paper cites.
Yue, Y.; Broder, J.; Kleinberg, R.; Joachims, T. The k-armed dueling bandits problem. Journal of Computer and System Sciences 2012
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; Moritz, P. Trust region policy optimization. International conference on machine learning. 2015; pp 1889–1897
2015
Earlier work this paper cites.
Dudík, M.; Hofmann, K.; Schapire, R. E.; Slivkins, A.; Zoghi, M. Contextual dueling bandits. Conference on Learning Theory. 2015; pp 563–587
2015
Earlier work this paper cites.
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. 2017
2017
Earlier work this paper cites.
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; Amodei, D. Deep Reinforcement Learning from Human Preferences. Advances in Neural Information Processing Systems. 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Kalashnikov, D.; Irpan, A.; Pastor, P.; Ibarz, J.; Herzog, A.; Jang, E.; Quillen, D.; Holly, E.; Kalakrishnan, M.; Vanhoucke, V.; others Scalable deep reinforcement learning for vision-based robotic manipulation. Conference on robot learning. 2018; pp 651–673
2018
Earlier work this paper cites.
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; Tafjord, O. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. 2018
2018
Earlier work this paper cites.
Henderson, P.; Islam, R.; Bachman, P.; Pineau, J.; Precup, D.; Meger, D. Deep Reinforcement Learning that Matters. 2019
2019
Earlier work this paper cites.
Kool, W.; van Hoof, H.; Welling, M. Buy 4 REINFORCE Samples, Get a Baseline for Free! DeepRLStructPred@ICLR. 2019
2019
Earlier work this paper cites.
Agarwal, A.; Jiang, N.; Kakade, S. M.; Sun, W. Reinforcement learning: Theory and algorithms. 2019
2019
Earlier work this paper cites.
Sakaguchi, K.; Bras, R. L.; Bhagavatula, C.; Choi, Y. WINOGRANDE: An Adversarial Winograd Schema Challenge at Scale. 2019
2019
Earlier work this paper cites.
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; Choi, Y. HellaSwag: Can a Machine Really Finish Your Sentence? 2019
2019
Earlier work this paper cites.
Peng, X. B.; Kumar, A.; Zhang, G.; Levine, S. Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning. 2019
2019
Earlier work this paper cites.
Jacq, A.; Geist, M.; Paiva, A.; Pietquin, O. Learning from a Learner. Proceedings of the 36th International Conference on Machine Learning. 2019; pp 2990–2999
2019
Earlier work this paper cites.
Ziegler, D. M.; Stiennon, N.; Wu, J.; Brown, T. B.; Radford, A.; Amodei, D.; Christiano, P.; Irving, G. Fine-Tuning Language Models from Human Preferences. 2020
2020
Earlier work this paper cites.
Engstrom, L.; Ilyas, A.; Santurkar, S.; Tsipras, D.; Janoos, F.; Rudolph, L.; Madry, A. Implementation Matters in Deep Policy Gradients: A Case Study on PPO and TRPO. 2020
2020
Earlier work this paper cites.
Hsu, C. C.-Y.; Mendler-Dünner, C.; Hardt, M. Revisiting Design Choices in Proximal Policy Optimization. 2020
2020
Cited alongside, same era.
Shani, L.; Efroni, Y.; Mannor, S. Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps. Proceedings of the AAAI Conference on Artificial Intelligence. 2020; pp 5668–5675
2020
Cited alongside, same era.
Agarwal, A.; Henaff, M.; Kakade, S.; Sun, W. Pc-pg: Policy cover directed exploration for provable policy gradient learning. Advances in neural information processing systems 2020
2020
Cited alongside, same era.
Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; Christiano, P. F. Learning to summarize with human feedback. Advances in Neural Information Processing Systems 2020
2020
Cited alongside, same era.
Watson, J.; Huang, S. H.; Heess, N. Coherent Soft Imitation Learning. 2023
2023
Later among the works it cites.
Lee, K.; Liu, H.; Ryu, M.; Watkins, O.; Du, Y.; Boutilier, C.; Abbeel, P.; Ghavamzadeh, M.; Gu, S. S. Aligning Text-to-Image Models using Human Feedback. 2023
2023
Later among the works it cites.
Azar, M. G.; Rowland, M.; Piot, B.; Guo, D.; Calandriello, D.; Valko, M.; Munos, R. A General Theoretical Paradigm to Understand Learning from Human Preferences. 2023
2023
Later among the works it cites.
Ethayarajh, K.; Xu, W.; Kiela, D. Better, Cheaper, Faster LLM Alignment with KTO. 2023; https://contextual.ai/better-cheaper-faster-llm-alignment-with-kto/
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Richter, L.; Boustati, A.; Nüsken, N.; Ruiz, F.; Akyildiz, O. D. VarGrad: a low-variance gradient estimator for variational inference. Advances in Neural Information Processing Systems 2020
2020
Cited alongside, same era.
Agarwal, A.; Kakade, S. M.; Lee, J. D.; Mahajan, G. On the theory of policy gradient methods: Optimality, approximation, and distribution shift. Journal of Machine Learning Research 2021
2021
Cited alongside, same era.
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; Steinhardt, J. Measuring Massive Multitask Language Understanding. 2021
2021
Cited alongside, same era.
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; Schulman, J. Training Verifiers to Solve Math Word Problems. 2021
2021
Cited alongside, same era.
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. 2021
2021
Cited alongside, same era.
Agarwal, R.; Schwarzer, M.; Castro, P. S.; Courville, A. C.; Bellemare, M. Deep reinforcement learning at the edge of the statistical precipice. Advances in neural information processing systems 2021
2021
Cited alongside, same era.
Swamy, G.; Choudhury, S.; Bagnell, J. A.; Wu, S. Of moments and matching: A game-theoretic framework for closing the imitation gap. International Conference on Machine Learning. 2021; pp 10022–10032
2021
Cited alongside, same era.
Foster, D. J.; Krishnamurthy, A. Efficient First-Order Contextual Bandits: Prediction, Allocation, and Triangular Discrimination. 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
Ball, P. J.; Smith, L.; Kostrikov, I.; Levine, S. Efficient Online Reinforcement Learning with Offline Data. 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Wang, K.; Zhou, K.; Wu, R.; Kallus, N.; Sun, W. The benefits of being distributional: Small-loss bounds for reinforcement learning. Advances in Neural Information Processing Systems 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Wang, Y.; Liu, Q.; Jin, C. Is RLHF More Difficult than Standard RL? A Theoretical Perspective. Thirty-seventh Conference on Neural Information Processing Systems. 2023
2023
Later among the works it cites.
Xiong, W.; Zhong, H.; Shi, C.; Shen, C.; Wang, L.; Zhang, T. Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game. The Eleventh International Conference on Learning Representations. 2023
2023
Later among the works it cites.
Fan, Y.; Watkins, O.; Du, Y.; Liu, H.; Ryu, M.; Boutilier, C.; Abbeel, P.; Ghavamzadeh, M.; Lee, K.; Lee, K. Reinforcement learning for fine-tuning text-to-image diffusion models. Advances in Neural Information Processing Systems 2024
2024
Closest in time.
2024
Closest in time.
Ahmadian, A.; Cremer, C.; Gallé, M.; Fadaee, M.; Kreutzer, J.; Pietquin, O.; Üstün, A.; Hooker, S. Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs. 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Huang, S.; Noukhovitch, M.; Hosseini, A.; Rasul, K.; Wang, W.; Tunstall, L. The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization. 2024
2024
Closest in time.
Wang, G.; Cheng, S.; Zhan, X.; Li, X.; Song, S.; Liu, Y. OpenChat: Advancing Open-source Language Models with Mixed-Quality Data. 2024
2024
Closest in time.
Dubois, Y.; Galambosi, B.; Liang, P.; Hashimoto, T. B. Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. 2024
2024
Closest in time.
Dubois, Y.; Li, X.; Taori, R.; Zhang, T.; Gulrajani, I.; Ba, J.; Guestrin, C.; Liang, P.; Hashimoto, T. B. AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback. 2024
2024
Closest in time.
Meta Introducing Meta Llama 3: The most capable openly available LLM to date. 2024; https://ai.meta.com/blog/meta-llama-3/
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Anthropic Introducing the next generation of Claude. 2024; https://www.anthropic.com/news/claude-3-family
2024
Closest in time.
Lambert, N.; Pyatkin, V.; Morrison, J.; Miranda, L.; Lin, B. Y.; Chandu, K.; Dziri, N.; Kumar, S.; Zick, T.; Choi, Y.; Smith, N. A.; Hajishirzi, H. RewardBench: Evaluating Reward Models for Language Modeling. 2024
2024
Closest in time.
Tajwar, F.; Singh, A.; Sharma, A.; Rafailov, R.; Schneider, J.; Xie, T.; Ermon, S.; Finn, C.; Kumar, A. Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data. 2024
2024
Closest in time.
Xiong, W.; Dong, H.; Ye, C.; Wang, Z.; Zhong, H.; Ji, H.; Jiang, N.; Zhang, T. Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint. 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Farebrother, J.; Orbay, J.; Vuong, Q.; Taïga, A. A.; Chebotar, Y.; Xiao, T.; Irpan, A.; Levine, S.; Castro, P. S.; Faust, A.; Kumar, A.; Agarwal, R. Stop Regressing: Training Value Functions via Classification for Scalable Deep RL. 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Dice, N. E.; Swamy, G.; Choudhury, S.; Sun, W. Efficient Inverse Reinforcement Learning without Compounding Errors. ICML 2024 Workshop on Models of Human Feedback for AI Alignment. 2024
2024
Closest in time.
2024
Closest in time.