Fetching the paper…
Reading the bibliography…
Current reinforcement learning (RL) frameworks for large language models (LLM) post-training typically assume a fixed prompt distribution, which is sub-optimal and bottlenecks scalability.
Some aspects of welfare economics
Pigou, A. C · 1920
Earlier work this paper cites.
A treatise on probability
Keynes, J. M · 1921
Earlier work this paper cites.
Marktform und Gleichgewicht
von Stackelberg, H · 1934
Earlier work this paper cites.
The theory of statistical decision
Savage, L. J · 1951
Earlier work this paper cites.
Social choice and individual values , volume 12
Arrow, K. J · 1952
Earlier work this paper cites.
Minimax theorems
Fan, K · 1953
Earlier work this paper cites.
The Origins of Totalitarianism
Arendt, H · 1958
Earlier work this paper cites.
Some studies in machine learning using the game of checkers
Samuel, A. L · 1959
Earlier work this paper cites.
On pseudo-games
Banos, A · 1968
Earlier work this paper cites.
Evolutionsstrategien für die numerische Optimierung
Schwefel, H.-P · 1977
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J · 1991
Earlier work this paper cites.
The zone of proximal development in Vygotsky’s analysis of learning and instruction
Chaiklin, S. et al · 2003
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Earlier work this paper cites.
Deep learning of representations for unsupervised and transfer learning
Bengio, Y · 2012
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Ross, S. and Bagnell, J. A · 2012
Earlier work this paper cites.
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Learning to optimize via information-directed sampling
Russo, D. and Van Roy, B · 2014
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silve, D · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Double thompson sampling for dueling bandits
Wu, H. and Liu, X · 2016
Earlier work this paper cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Lin, Z., Kostrikov, I., Synnaeve, G., Szlam, A., and Fergus, R · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Earlier work this paper cites.
Adversarial domain randomization
Khirodkar, R. and Kitani, K. M · 2018
Earlier work this paper cites.
On the difficulty of warm-starting neural network training
Ash, J. T. and Adams, R. P · 2019
Earlier work this paper cites.
Accelerating deep learning by focusing on the biggest losers
Jiang, A. H., Wong, D. L.-K., Zhou, G., Andersen, D. G., Dean, J., Ganger, G. R., Joshi, G., Kaminksy, M., Kozuch, M., Lipton, Z. C., et al · 2019
Earlier work this paper cites.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A., Russell, S., Critch, A., and Levine, S · 2020
Earlier work this paper cites.
What is local optimality in nonconvex-nonconcave minimax optimization?
Jin, C., Netrapalli, P., and Jordan, M · 2020
Earlier work this paper cites.
Ordered sgd: A new stochastic optimization framework for empirical risk minimization
Kawaguchi, K. and Lu, H · 2020
Earlier work this paper cites.
Reward is enough
Silver, D., Singh, S., Precup, D., and Sutton, R. S · 2021
Earlier work this paper cites.
Decisions under Ignorance , 2022
Gustafsson, J. E · 2022
Earlier work this paper cites.
Large language models can self-improve
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Earlier work this paper cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Mindermann, S., Brauner, J. M., Razzak, M. T., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., et al · 2022
Earlier work this paper cites.
The primacy bias in deep reinforcement learning
Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P.-L., and Courville, A · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Evolving curricula with regret-based environment design
Parker-Holder, J., Jiang, M., Dennis, M., Samvelyan, M., Foerster, J., Grefenstette, E., and Rocktäschel, T · 2022
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Cited alongside, same era.
Near-optimal local convergence of alternating gradient descent-ascent for minimax optimization
Zhang, G., Wang, Y., Lessard, L., and Grosse, R. B · 2022
Cited alongside, same era.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Dubois, Y., Galambosi, B., Liang, P., and Hashimoto, T. B · 2024
Closest in time.
Efficient exploration for llms
Dwaracherla, V., Asghari, S. M., Hao, B., and Van Roy, B · 2024
Closest in time.
Direct language model alignment from online ai feedback
Guo, S., Zhang, B., Liu, T., Liu, T., Khalman, M., Llinares, F., Rame, A., Mesnard, T., Zhao, Y., Piot, B., et al · 2024
Closest in time.
Orpo: Monolithic preference optimization without reference model
Hong, J., Lee, N., and Thorne, J · 2024
Closest in time.
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abbas, Z., Zhao, R., Modayil, J., White, A., and Machado, M. C · 2023
Cited alongside, same era.
Efficient online reinforcement learning with offline data
Ball, P. J., Smith, L., Kostrikov, I., and Levine, S · 2023
Cited alongside, same era.
Ultrafeedback: Boosting language models with high-quality feedback
Cui, G., Yuan, L., Ding, N., Yao, G., Zhu, W., Ni, Y., Xie, G., Liu, Z., and Sun, M · 2023
Cited alongside, same era.
Maintaining plasticity in deep continual learning
Dohare, S., Hernandez-Garcia, J. F., Rahman, P., Mahmood, A. R., and Sutton, R. S · 2023
Cited alongside, same era.
Raft: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Cited alongside, same era.
Reinforced self-training (rest) for language modeling
Gulcehre, C., Paine, T. L., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., et al · 2023
Cited alongside, same era.
Connecting large language models with evolutionary algorithms yields powerful prompt optimizers
Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., and Yang, Y · 2023
Cited alongside, same era.
Contrastive prefence learning: Learning from human feedback without rl
Hejna, J., Rafailov, R., Sikchi, H., Finn, C., Niekum, S., Knox, W. B., and Sadigh, D · 2023
Cited alongside, same era.
Hu, J., Wu, X., Zhu, Z., Xianyu, Wang, W., Zhang, D., and Cao, Y · 2024
Closest in time.
Open-Endedness is Essential for Artificial Superhuman Intelligence
Hughes, E., Dennis, M., Parker-Holder, J., Behbahani, F., Mavalankar, A., Shi, Y., Schaul, T., and Rocktaschel, T · 2024
Closest in time.
Adaptive data optimization: Dynamic sample selection with scaling laws
Jiang, Y., Zhou, A., Feng, Z., Malladi, S., and Kolter, J. Z · 2024
Closest in time.
Training language models to self-correct via reinforcement learning
Kumar, A., Zhuang, V., Agarwal, R., Su, Y., Co-Reyes, J. D., Singh, A., Baumli, K., Iqbal, S., Bishop, C., Roelofs, R., et al · 2024
Closest in time.
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Li, T., Chiang, W.-L., Frick, E., Dunlap, L., Wu, T., Zhu, B., Gonzalez, J. E., and Stoica, I · 2024
Closest in time.
Skywork Reward Model Series
Liu, C. Y. and Zeng, L · 2024
Closest in time.
STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback
Makar-Limanov, J., Prakash, A., Goktas, D., Greenwald, A., and Ayanian, N · 2024
Closest in time.
SimPO: Simple Preference Optimization with a Reference-Free Reward
Meng, Y., Xia, M., and Chen, D · 2024
Closest in time.
Active Preference Learning for Large Language Models
Muldrew, W., Hayes, P., Zhang, M., and Barber, D · 2024
Closest in time.
The Importance of Directional Feedback for LLM-based Optimizers
Nie, A., Cheng, C.-A., Kolobov, A., and Swaminathan, A · 2024
Closest in time.
Iterative reasoning preference optimization
Pang, R. Y., Yuan, W., Cho, K., He, H., Sukhbaatar, S., and Weston, J · 2024
Closest in time.
Learning Formal Mathematics From Intrinsic Motivation
Poesia, G., Broman, D., Haber, N., and Goodman, N. D · 2024
Closest in time.
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Rosset, C., Cheng, C.-A., Mitra, A., Santacroce, M., Awadallah, A., and Xie, T · 2024
Closest in time.
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
Setlur, A., Garg, S., Geng, X., Garg, N., Smith, V., and Kumar, A · 2024
Closest in time.
The Crucial Role of Samplers in Online Direct Preference Optimization
Shi, R., Zhou, R., and Du, S. S · 2024
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y., and Gan, C · 2024
Closest in time.
Scaling the right thing matters more now than ever , 2024
Sutskever, I · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Tang, Y., Guo, Z. D., Zheng, Z., Calandriello, D., Munos, R., Rowland, M., Richemond, P. H., Valko, M., Pires, B. Á., and Piot, B · 2024
Closest in time.
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Wang, H., Xiong, W., Xie, T., Zhao, H., and Zhang, T · 2024
Closest in time.
Self-play preference optimization for language model alignment
Wu, Y., Sun, Z., Yuan, H., Ji, K., Yang, Y., and Gu, Q · 2024
Closest in time.
Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T · 2024
Closest in time.
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Xu, Z., Jiang, F., Niu, L., Deng, Y., Poovendran, R., Choi, Y., and Lin, B. Y · 2024
Closest in time.
To repeat or not to repeat: Insights from scaling llm under token-crisis
Xue, F., Fu, Y., Zhou, W., Zheng, Z., and You, Y · 2024
Closest in time.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Closest in time.
TextGrad: Automatic” Differentiation” via Text
Yuksekgonul, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., Guestrin, C., and Zou, J · 2024
Closest in time.
Toward Optimal LLM Alignments Using Two-Player Games
Zheng, R., Guo, H., Liu, Z., Zhang, X., Yao, Y., Xu, X., Wang, Z., Xi, Z., Gui, T., Zhang, Q., et al · 2024
Closest in time.
Beyond Preferences in AI Alignment
Zhi-Xuan, T., Carroll, M., Franklin, M., and Ashton, H · 2024
Closest in time.
Take on DeepSeek-R1 , 2025
Nishihara, R · 2025
Closest in time.
Machine Learning Should Maximize Welfare, Not (Only) Accuracy
Rosenfeld, N. and Xu, H · 2025
Closest in time.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Team, D., Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.