Fetching the paper…
Reading the bibliography…
A simple yet effective method for inference-time alignment of generative models is Best-of-$N$ (BoN), where $N$ outcomes are sampled from a reference policy, evaluated using a proxy reward model, and the highest-scoring one is selected.
Statistical theory of extreme values and some practical applications
Emil Julius Gumbel · 1954
Earlier work this paper cites.
Risk-sensitive markov decision processes
Ronald A Howard and James E Matheson · 1972
Earlier work this paper cites.
A new penalty function method for constrained minimization
Barry W Kort and Dimitri P Bertsekas · 1972
Earlier work this paper cites.
Q-learning for risk-sensitive control
Vivek S Borkar · 2002
Earlier work this paper cites.
On solving large-scale finite minimax problems using exponential smoothing
EY Pee and Johannes O Royset · 2011
Earlier work this paper cites.
Robust variable selection with exponential squared loss
Xueqin Wang, Yunlu Jiang, Mian Huang, and Heping Zhang · 2013
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
The master equation and the convergence problem in mean field games:(ams-201)
Pierre Cardaliaguet, François Delarue, Jean-Michel Lasry, and Pierre-Louis Lions · 2019
Earlier work this paper cites.
Deep learning theory review: An optimal control and dynamical systems perspective
Guan-Horng Liu and Evangelos A Theodorou · 2019
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
A short note on an inequality between kl and tv
Clément L Canonne · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Information theory: From coding to learning, 2022
Yury Polyanskiy and Yihong Wu · 2022
Earlier work this paper cites.
Convergence of the inexact langevin algorithm and score-based generative models in kl divergence
Kaylee Yingxi Yang and Andre Wibisono · 2022
Earlier work this paper cites.
Calibrating sequence likelihood improves conditional language generation
Yao Zhao, Mikhail Khalman, Rishabh Joshi, Shashi Narayan, Mohammad Saleh, and Peter J Liu · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al · 2023
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2023
Cited alongside, same era.
Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Jacob Eisenstein, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alex D’Amour, DJ Dvijotham, Adam Fisch, Katherine Heller, Stephen Pfohl, Deepak Ramachandran, et al · 2023
Cited alongside, same era.
Unveiling safety vulnerabilities of large language models
George Kour, Marcel Zalmanovici, Naama Zwerdling, Esther Goldbraich, Ora Nova Fandina, Ateret Anaby-Tavor, Orna Raz, and Eitan Farchi · 2023
Cited alongside, same era.
On tilted losses in machine learning: Theory and applications
Tian Li, Ahmad Beirami, Maziar Sanjabi, and Virginia Smith · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Bond: Aligning llms with best-of-n distillation
Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot, Johan Ferret, Nino Vieillard, Alexandre Ramé, Bobak Shariari, Sarah Perrin, Abe Friesen, Geoffrey Cideron, et al · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
The importance of online data: Understanding preference fine-tuning via coverage
Yuda Song, Gokul Swamy, Aarti Singh, Drew Bagnell, and Wen Sun · 2024
Later among the works it cites.
Inference scaling f-laws: The limits of llm resampling with imperfect verifiers
Benedikt Stroebl, Sayash Kapoor, and Arvind Narayanan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Provable offline preference-based reinforcement learning
Wenhao Zhan, Masatoshi Uehara, Nathan Kallus, Jason D Lee, and Wen Sun · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Cited alongside, same era.
Theoretical guarantees on the best-of-n alignment policy
Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D’Amour, Jacob Eisenstein, Chirag Nagpal, and Ananda Theertha Suresh · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini · 2024
Cited alongside, same era.
Bonbon alignment for large language models and the sweetness of best-of-n sampling
Lin Gui, Cristina Gârbacea, and Victor Veitch · 2024
Cited alongside, same era.
Hanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang, Jiahao Qiu, Ming Yin, Mengdi Wang, Peter Bartlett, and Andrea Zanette · 2024
Later among the works it cites.
Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, and Tong Zhang · 2024
Later among the works it cites.
Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang · 2024
Later among the works it cites.
Asymptotics of language model alignment
Joy Qiping Yang, Salman Salamatian, Ziteng Sun, Ananda Theertha Suresh, and Ahmad Beirami · 2024
Later among the works it cites.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
Chenlu Ye, Wei Xiong, Yuheng Zhang, Nan Jiang, and Tong Zhang · 2024
Later among the works it cites.
Sharp analysis for kl-regularized contextual bandits and rlhf
Heyang Zhao, Chenlu Ye, Quanquan Gu, and Tong Zhang · 2024
Later among the works it cites.
Variational best-of-n alignment
Afra Amini, Tim Vieira, Elliott Ash, and Ryan Cotterell · 2025
Closest in time.
Theoretical analysis of kl-regularized rlhf with multiple reference models
Gholamali Aminian, Amir R Asadi, Idan Shenfeld, and Youssef Mroueh · 2025
Closest in time.
Infalign: Inference-aware language model alignment
Ananth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein, Michael Collins, Adrian Hutter, Jong Lee, Chirag Nagpal, Flavien Prost, Aradhana Sinha, Ananda Theertha Suresh, and Ahmad Beirami · 2025
Closest in time.
Soft best-of-n sampling for model alignment
Mayrink Verdun Claudio, Oesterling Alex, Lakkaraju Himabindu, and P Calmon Flavio · 2025
Closest in time.
Guided speculative inference for efficient test-time alignment of llms
Jonathan Geuter, Youssef Mroueh, and David Alvarez-Melis · 2025
Closest in time.
Measuring goodhart’s law: Towards an evaluation framework for open-ended generative models
J. Hilton, P. Clark, et al · 2025
Closest in time.
Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment
Audrey Huang, Adam Block, Qinghua Liu, Nan Jiang, Dylan J Foster, and Akshay Krishnamurthy · 2025
Closest in time.
Inference-time reward hacking in large language models
Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju, and Flavio du Pin Calmon · 2025
Closest in time.