Fetching the paper…
Reading the bibliography…
This paper studies reinforcement learning from human feedback (RLHF) for aligning large language models with human preferences.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Intransitivity, utility, and the aggregation of preference patterns
Kenneth O. May · 1954
Earlier work this paper cites.
Intransitivity of preferences
Amos Tversky · 1969
Earlier work this paper cites.
Mathematical games, the paradox of the nontransitive dice and the elusive principle of indifference
M Gardner · 1970
Earlier work this paper cites.
Non-parametric analysis of a generalized regression model: the maximum rank correlation estimator
Aaron K Han · 1987
Earlier work this paper cites.
Semiparametric efficiency bounds
Whitney K Newey · 1990
Earlier work this paper cites.
The limiting distribution of the maximum rank correlation estimator
Robert P Sherman · 1993
Earlier work this paper cites.
Estimation of regression coefficients when some regressors are not always observed
James M. Robins, Andrea Rotnitzky, and Lue Ping Zhao · 1994
Earlier work this paper cites.
Weak convergence
Aad W Van Der Vaart, Jon A Wellner, Aad W van der Vaart, and Jon A Wellner · 1996
Earlier work this paper cites.
Efficient and Adaptive Estimation for Semiparametric Models
Peter J. Bickel, Chris A. J. Klaassen, Ya’acov Ritov, and Jon A. Wellner · 1998
Earlier work this paper cites.
Adjusting for nonignorable drop-out using semiparametric nonresponse models
Daniel O Scharfstein, Andrea Rotnitzky, and James M Robins · 1999
Earlier work this paper cites.
Optimal structural nested models for optimal sequential decisions
James M Robins · 2004
Earlier work this paper cites.
Doubly robust estimation in missing data and causal inference models
Heejung Bang and James M Robins · 2005
Earlier work this paper cites.
Semiparametric Theory and Missing Data
Anastasios A. Tsiatis · 2006
Earlier work this paper cites.
Bounded, efficient and doubly robust estimation with inverse weighting
Zhiqiang Tan · 2010
Earlier work this paper cites.
Improved doubly robust estimation when data are monotonely coarsened, with application to longitudinal studies with dropout
Anastasios A Tsiatis, Marie Davidian, and Weihua Cao · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
A robust method for estimating optimal treatment regimes
Baqun Zhang, Anastasios A Tsiatis, Eric B Laber, and Marie Davidian · 2012
Earlier work this paper cites.
Robust estimation of optimal dynamic treatment regimes for sequential treatment decisions
Baqun Zhang, Anastasios A Tsiatis, Eric B Laber, and Marie Davidian · 2013
Earlier work this paper cites.
Covariate balancing propensity score
Kosuke Imai and Marc Ratkovic · 2014
Earlier work this paper cites.
Doubly Robust Policy Evaluation and Optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li · 2014
Earlier work this paper cites.
Gaussian approximation of suprema of empirical processes
Victor Chernozhukov, Denis Chetverikov, and Kengo Kato · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Stochastic choice and preferences for randomization
Marina Agranov and Pietro Ortoleva · 2015
Earlier work this paper cites.
Bias-reduced doubly robust estimation
Karel Vermeulen and Stijn Vansteelandt · 2015
Earlier work this paper cites.
Q-and a-learning methods for estimating optimal dynamic treatment regimes
Phillip J Schulte, Anastasios A Tsiatis, Eric B Laber, and Marie Davidian · 2015
Earlier work this paper cites.
Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy
Alexander R Luedtke and Mark J Van Der Laan · 2016
Earlier work this paper cites.
Doubly robust off-policy evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip S. Thomas and Emma Brunskill · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Non-parametric methods for doubly robust estimation of continuous treatment effects
Edward H Kennedy, Zongming Ma, Matthew D McHugh, and Dylan S Small · 2017
Earlier work this paper cites.
Minimax estimation of a functional on a structured high-dimensional model
James M Robins, Lingling Li, Rajarshi Mukherjee, Eric Tchetgen Tchetgen, and Aad van der Vaart · 2017
Earlier work this paper cites.
Concordance-assisted learning for estimating optimal individualized treatment regimes
Caiyun Fan, Wenbin Lu, Rui Song, and Yong Zhou · 2017
Earlier work this paper cites.
On estimation of optimal treatment regimes for maximizing t-year survival probability
Runchao Jiang, Wenbin Lu, Rui Song, and Marie Davidian · 2017
Earlier work this paper cites.
Semiparametric single-index model for estimating optimal individualized treatment strategy
Rui Song, Shikai Luo, Donglin Zeng, Hao Helen Zhang, Wenbin Lu, and Zhiguo Li · 2017
Earlier work this paper cites.
TL;DR: Mining reddit to learn automatic summarization
Michael Völske, Maxime Peyrard, Janek Bevendorff, Martin Potthast, and Benno Stein · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton, Andrew G Barto, et al · 2018
Earlier work this paper cites.
Estimation and inference of heterogeneous treatment effects using random forests
Stefan Wager and Susan Athey · 2018
Earlier work this paper cites.
Bounded, efficient and multiply robust estimation of average treatment effects using instrumental variables
Linbo Wang and Eric Tchetgen Tchetgen · 2018
Earlier work this paper cites.
Double/debiased machine learning for treatment and structural parameters
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins · 2018
Earlier work this paper cites.
High-dimensional a-learning for optimal dynamic treatment regimes
Chengchun Shi, Alin Fan, Rui Song, and Wenbin Lu · 2018
Earlier work this paper cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Earlier work this paper cites.
Policy evaluation and optimization with continuous treatments
Nathan Kallus and Angela Zhou · 2018
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Implementation matters in deep rl: A case study on ppo and trpo
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2019
Cited alongside, same era.
Metalearners for estimating heterogeneous treatment effects using machine learning
Sören R Künzel, Jasjeet S Sekhon, Peter J Bickel, and Bin Yu · 2019
Cited alongside, same era.
Orthogonal random forest for causal inference
Miruna Oprescu, Vasilis Syrgkanis, and Zhiwei Steven Wu · 2019
Cited alongside, same era.
Adapting neural networks for the estimation of treatment effects
Claudia Shi, David Blei, and Victor Veitch · 2019
Cited alongside, same era.
Measuring conditional independence by independent residuals for causal discovery
Hao Zhang, Shuigeng Zhou, Jihong Guan, and Jun Huan · 2019
Cited alongside, same era.
More efficient off-policy evaluation through regularized targeted learning
Drcfs: Doubly robust causal feature selection
Francesco Quinzan, Ashkan Soleymani, Patrick Jaillet, Cristian R Rojas, and Stefan Bauer · 2023
Later among the works it cites.
A multiagent reinforcement learning framework for off-policy evaluation in two-sided markets
Chengchun Shi, Runzhe Wan, Ge Song, Shikai Luo, Hongtu Zhu, and Rui Song · 2023
Later among the works it cites.
Semiparametrically efficient off-policy evaluation in linear markov decision processes
Chuhan Xie, Wenhao Yang, and Zhihua Zhang · 2023
Later among the works it cites.
An instrumental variable approach to confounded off-policy evaluation
Yang Xu, Jin Zhu, Chengchun Shi, Shikai Luo, and Rui Song · 2023
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aurelien Bibaut, Ivana Malenica, Nikos Vlassis, and Mark Van Der Laan · 2019
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Understanding learned reward functions
Eric J Michaud, Adam Gleave, and Stuart Russell · 2020
Cited alongside, same era.
Robust inference on population indirect causal effects: the generalized front door criterion
Isabel R Fulcher, Ilya Shpitser, Stella Marealle, and Eric J Tchetgen Tchetgen · 2020
Cited alongside, same era.
The hardness of conditional independence testing and the generalised covariance measure
Rajen D Shah and Jonas Peters · 2020
Cited alongside, same era.
Cassidy Laidlaw, Shivam Singhal, and Anca Dragan · 2024
Later among the works it cites.
The accuracy paradox in rlhf: When better reward models don’t yield better language models
Yanjun Chen, Dawei Zhu, Yirong Sun, Xinghao Chen, Wei Zhang, and Xiaoyu Shen · 2024
Later among the works it cites.
Learn your reference model for real good alignment
Alexey Gorbatovski, Boris Shaposhnikov, Alexey Malakhov, Nikita Surnachev, Yaroslav Aksenov, Ian Maksimov, Nikita Balagansky, and Daniil Gavrilov · 2024
Later among the works it cites.
Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, and Weijie J Su · 2024
Later among the works it cites.
Dense reward for free in reinforcement learning from human feedback
Alex J Chan, Hao Sun, Samuel Holt, and Mihaela Van Der Schaar · 2024
Later among the works it cites.
On designing effective rl reward at training time for llm reasoning
Jiaxuan Gao, Shusheng Xu, Wenjie Ye, Weilin Liu, Chuyi He, Wei Fu, Zhiyu Mei, Guangju Wang, and Yi Wu · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al · 2024
Later among the works it cites.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Later among the works it cites.
Preference ranking optimization for human alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang · 2024
Later among the works it cites.
Generalized preference optimization: A unified approach to offline alignment
Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rémi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo Ávila Pires, and Bilal Piot · 2024
Later among the works it cites.
Provably robust dpo: Aligning language models with noisy feedback
Sayak Ray Chowdhury, Anush Kini, and Nagarajan Natarajan · 2024
Later among the works it cites.
Ropo: Robust preference optimization for large language models
Xize Liang, Chao Chen, Shuang Qiu, Jie Wang, Yue Wu, Zhihang Fu, Zhihao Shi, Feng Wu, and Jieping Ye · 2024
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello · 2024
Later among the works it cites.
Human alignment of large language models through online preference optimisation
Daniele Calandriello, Daniel Guo, Remi Munos, Mark Rowland, Yunhao Tang, Bernardo Avila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, et al · 2024
Later among the works it cites.
Direct nash optimization: Teaching language models to self-improve with general preferences
Corby Rosset, Ching-An Cheng, Arindam Mitra, Michael Santacroce, Ahmed Awadallah, and Tengyang Xie · 2024
Later among the works it cites.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal · 2024
Later among the works it cites.
Online iterative reinforcement learning from human feedback with general preference model
Chenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong, Nan Jiang, and Tong Zhang · 2024
Later among the works it cites.
Semiparametric proximal causal inference
Yifan Cui, Hongming Pu, Xu Shi, Wang Miao, and Eric Tchetgen Tchetgen · 2024
Later among the works it cites.
Debiased inverse propensity score weighting for estimation of average treatment effects with high-dimensional confounders
Yuhao Wang and Rajen D Shah · 2024
Later among the works it cites.
Multiply robust estimation for average treatment effect among treated
Lu Wang and Peisong Han · 2024
Later among the works it cites.
Orthogonalized estimation of difference of q q -functions
Defu Cao and Angela Zhou · 2024
Later among the works it cites.
Combining experimental and historical data for policy evaluation
Ting Li, Chengchun Shi, Qianglin Wen, Yang Sui, Yongli Qin, Chunbo Lai, and Hongtu Zhu · 2024
Later among the works it cites.
Doubly robust interval estimation for optimal policy evaluation in online learning
Ye Shen, Hengrui Cai, and Rui Song · 2024
Later among the works it cites.
Provable multi-party reinforcement learning with diverse human feedback
Huiying Zhong, Zhun Deng, Weijie J Su, Zhiwei Steven Wu, and Linjun Zhang · 2024
Later among the works it cites.
Reward model learning vs. direct policy optimization: A comparative analysis of learning from human preferences
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban, George Tzannetos, Goran Radanović, and Adish Singla · 2024
Later among the works it cites.
The n+ implementation details of rlhf with ppo: A case study on tl; dr summarization
Shengyi Huang, Michael Noukhovitch, Arian Hosseini, Kashif Rasul, Weixun Wang, and Lewis Tunstall · 2024
Later among the works it cites.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Yann Dubois, Balázs Galambosi, Percy Liang, and Tatsunori B Hashimoto · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Reward shaping to mitigate reward hacking in rlhf
Jiayi Fu, Xuandong Zhao, Chengyuan Yao, Heng Wang, Qi Han, and Yanghua Xiao · 2025
Closest in time.
On a connection between imitation learning and rlhf
Teng Xiao, Yige Yuan, Mingxiao Li, Zhengyu Chen, and Vasant G Honavar · 2025
Closest in time.
Robust reinforcement learning from human feedback for large language models fine-tuning
Kai Ye, Hongyi Zhou, Jin Zhu, Francesco Quinzan, and Chengchun Shi · 2025
Closest in time.
Reinforce++: A simple and efficient approach for aligning large language models
Jian Hu · 2025
Closest in time.
Dapo: An open-source llm reinforcement learning system at scale
Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Tiantian Fan, Gaohong Liu, Lingjun Liu, Xin Liu, et al · 2025
Closest in time.
Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks
Yufeng Yuan, Qiying Yu, Xiaochen Zuo, Ruofei Zhu, Wenyuan Xu, Jiaze Chen, Chengyi Wang, TianTian Fan, Zhengyin Du, Xiangpeng Wei, et al · 2025
Closest in time.
Balancing interference and correlation in spatial experimental designs: A causal graph cut approach
Jin Zhu, Jingyi Li, Hongyi Zhou, Yinan Lin, Zhenhua Lin, and Chengchun Shi · 2025
Closest in time.
Value enhancement of reinforcement learning via efficient and robust trust region optimization
Chengchun Shi, Zhengling Qi, Jianing Wang, and Fan Zhou · 2025
Closest in time.
Characterization of efficient influence function for off-policy evaluation under optimal policies
Haoyu Wei · 2025
Closest in time.
Theoretical analysis of kl-regularized rlhf with multiple reference models
Gholamali Aminian, Amir R Asadi, Idan Shenfeld, and Youssef Mroueh · 2025
Closest in time.
A note on DPO with noisy preferences and relationship to IPO
Eric Mitchell · 2025
Closest in time.