Fetching the paper…
Reading the bibliography…
A major challenge in aligning large language models (LLMs) with human preferences is the issue of distribution shift.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Cédric Villani et al · 2009
Earlier work this paper cites.
A tail inequality for quadratic forms of subgaussian random vectors
Daniel Hsu, Sham Kakade, and Tong Zhang · 2012
Earlier work this paper cites.
Concentration Inequalities: A Nonasymptotic Theory of Independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
Kullback-leibler divergence constrained distributionally robust optimization
Zhaolin Hu and L Jeff Hong · 2013
Earlier work this paper cites.
Introduction to nonlinear optimization: Theory, algorithms, and applications with MATLAB
Amir Beck · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Stochastic gradient methods for distributionally robust optimization with f-divergences
Hongseok Namkoong and John C Duchi · 2016
Earlier work this paper cites.
First-order methods in optimization
Amir Beck · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
A robust learning approach for regression models based on distributionally robust optimization
Ruidi Chen and Ioannis Ch Paschalidis · 2018
Earlier work this paper cites.
Data-driven distributionally robust optimization using the wasserstein metric: performance guarantees and tractable reformulations
Peyman Mohajerin Esfahani and Daniel Kuhn · 2018
Earlier work this paper cites.
CARER: Contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen · 2018
Earlier work this paper cites.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun · 2019
Earlier work this paper cites.
Wasserstein distributionally robust optimization: Theory and applications in machine learning
Daniel Kuhn, Peyman Mohajerin Esfahani, Viet Anh Nguyen, and Soroosh Shafieezadeh-Abadeh · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Regularization via mass transportation
Soroosh Shafieezadeh-Abadeh, Daniel Kuhn, and Peyman Mohajerin Esfahani · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Distributionally robust learning
Ruidi Chen, Ioannis Ch Paschalidis, et al · 2020
Earlier work this paper cites.
Large-scale methods for distributionally robust optimization
Daniel Levy, Yair Carmon, John C Duchi, and Aaron Sidford · 2020
Earlier work this paper cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Aditya Modi, Nan Jiang, Ambuj Tewari, and Satinder Singh · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Distributionally robust policy evaluation and learning in offline contextual bandits
Nian Si, Fan Zhang, Zhengyuan Zhou, and Jose Blanchet · 2020
Earlier work this paper cites.
Measuring robustness to natural distribution shifts in image classification
Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt · 2020
Cited alongside, same era.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Cited alongside, same era.
Distributionally robust imitation learning
Mohammad Ali Bashiri, Brian Ziebart, and Xinhua Zhang · 2021
Cited alongside, same era.
Learning models with uniform performance via distributionally robust optimization
John C Duchi and Hongseok Namkoong · 2021
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Cited alongside, same era.
Wilds: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al · 2021
Distributionally robust policy gradient for offline contextual bandits
Zhouhao Yang, Yihong Guo, Pan Xu, Anqi Liu, and Animashree Anandkumar · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Banghua Zhu, Michael Jordan, and Jiantao Jiao · 2023
Later among the works it cites.
Robust reinforcement learning from corrupted human feedback
Alexander Bukharin, Ilgee Hong, Haoming Jiang, Zichong Li, Qingru Zhang, Zixuan Zhang, and Tuo Zhao · 2024
Later among the works it cites.
Maxmin-RLHF: Alignment with diverse human preferences
Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Dinesh Manocha, Furong Huang, Amrit Bedi, and Mengdi Wang · 2024
Later among the works it cites.
Provably robust DPO: Aligning language models with noisy feedback
Sayak Ray Chowdhury, Anush Kini, and Nagarajan Natarajan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
What are the statistical limits of offline RL with linear function approximation?
Ruosong Wang, Dean Foster, and Sham M. Kakade · 2021
Cited alongside, same era.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhengqing Zhou, Qinxun Bai, Zhengyuan Zhou, Linhai Qiu, Jose Blanchet, and Peter Glynn · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
Distributionally robust stochastic optimization with wasserstein distance
Rui Gao and Anton Kleywegt · 2022
Cited alongside, same era.
Wasserstein distributionally robust optimization and variation regularization
Rui Gao, Xi Chen, and Anton J Kleywegt · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nguyen, Thomas Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli · 2024
Later among the works it cites.
Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Jacob Eisenstein, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alexander Nicholas D’Amour, Krishnamurthy Dj Dvijotham, Adam Fisch, Katherine A Heller, Stephen Robert Pfohl, Deepak Ramachandran, Peter Shaw, and Jonathan Berant · 2024
Later among the works it cites.
Open llm leaderboard v2, 2024
Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf · 2024
Later among the works it cites.
Bring your own (non-robust) algorithm to solve robust MDPs by estimating the worst kernel
Uri Gadot, Kaixin Wang, Navdeep Kumar, Kfir Yehuda Levy, and Shie Mannor · 2024
Later among the works it cites.
The language model evaluation harness, 07 2024
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2024
Later among the works it cites.
Understanding the effects of rlhf on llm generalisation and diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu · 2024
Later among the works it cites.
Reward model learning vs. direct policy optimization: A comparative analysis of learning from human preferences
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban, Georgios Tzannetos, Goran Radanovic, and Adish Singla · 2024
Later among the works it cites.
Beyond the binary: Capturing diverse preferences with reward regularization
Vishakh Padmakumar, Chuanyang Jin, Hannah Rose Kirk, and He He · 2024
Later among the works it cites.
Group robust preference optimization in reward-free rlhf
Shyam Sundhar Ramesh, Yifan Hu, Iason Chaimalas, Viraj Mehta, Pier Giuseppe Sessa, Haitham Bou Ammar, and Ilija Bogunovic · 2024
Later among the works it cites.
Distributionally robust model-based offline reinforcement learning with near-optimal sample complexity
Laixi Shi and Yuejie Chi · 2024
Later among the works it cites.
β \beta -dpo: Direct preference optimization with dynamic β \beta
Junkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu, Jinyang Gao, Bolin Ding, Xiang Wang, and Xiangnan He · 2024
Later among the works it cites.
Yuzi Yan, Xingzhou Lou, Jialian Li, Yiping Zhang, Jian Xie, Chao Yu, Yu Wang, Dong Yan, and Yuan Shen · 2024
Later among the works it cites.
Group preference optimization: Few-shot alignment of large language models
Siyan Zhao, John Dang, and Aditya Grover · 2024
Later among the works it cites.
Natural actor-critic for robust reinforcement learning with function approximation
Ruida Zhou, Tao Liu, Min Cheng, Dileep Kalathil, PR Kumar, and Chao Tian · 2024
Later among the works it cites.
Correcting the mythos of KL-regularization: Direct alignment without overoptimization via chi-squared preference optimization
Audrey Huang, Wenhao Zhan, Tengyang Xie, Jason D. Lee, Wen Sun, Akshay Krishnamurthy, and Dylan J Foster · 2025
Closest in time.
Distributionally robust reinforcement learning with human feedback
Debmalya Mandal, Paulius Sasnauskas, and Goran Radanovic · 2025
Closest in time.
Bridging distributionally robust learning and offline rl: An approach to mitigate distribution shift and partial data coverage
Kishan Panaganti, Zaiyan Xu, Dileep Kalathil, and Mohammad Ghavamzadeh · 2025
Closest in time.
Towards robust alignment of language models: Distributionally robustifying direct preference optimization
Junkang Wu, Yuexiang Xie, Zhengyi Yang, Jiancan Wu, Jiawei Chen, Jinyang Gao, Bolin Ding, Xiang Wang, and Xiangnan He · 2025
Closest in time.
Diverging preferences: When do annotators disagree and do models know?
Michael JQ Zhang, Zhilin Wang, Jena D. Hwang, Yi Dong, Olivier Delalleau, Yejin Choi, Eunsol Choi, Xiang Ren, and Valentina Pyatkin · 2025
Closest in time.