Fetching the paper…
Reading the bibliography…
A simple and effective method for the inference-time alignment and scaling test-time compute of generative models is best-of-$n$ sampling, where $n$ samples are drawn from a reference policy, ranked based on a reward function, and the highest ranking one is selected.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
FUDGE: Controlled text generation with future discriminators
Yang, K. and Klein, D · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
RL with KL penalties is better viewed as Bayesian inference
Korbak, T., Perez, E., and Buckley, C · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Calibrating sequence likelihood improves conditional language generation
Zhao, Y., Khalman, M., Joshi, R., Narayan, S., Saleh, M., and Liu, P. J · 2022
Earlier work this paper cites.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Earlier work this paper cites.
Compositional preference models for aligning LMs
Go, D., Korbak, T., Kruszewski, G., Rozen, J., and Dymetman, M · 2023
Earlier work this paper cites.
On tilted losses in machine learning: Theory and applications
Li, T., Beirami, A., Sanjabi, M., and Smith, V · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Earlier work this paper cites.
Training language models with language feedback at scale
Scheurer, J., Campos, J. A., Korbak, T., Chan, J. S., Chen, A., Cho, K., and Perez, E · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model, 2023
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Guo, Z. D., Piot, B., Munos, R., Rowland, M., Valko, M., and Calandriello, D · 2024
Cited alongside, same era.
Liar: Leveraging alignment (best-of-n) to jailbreak llms in seconds
Hughes, J., Price, S., Lynch, A., Schaeffer, R., Barez, F., Koyejo, S., Sleight, H., Jones, E., Perez, E., and Sharma, M · 2024
Closest in time.
Rain: Your language models can align themselves without finetuning
Li, Y., Wei, F., Zhao, J., Zhang, C., and Zhang, H · 2024
Closest in time.
Information theoretic guarantees for policy alignment in large language models
Mroueh, Y · 2024
Closest in time.
Controlled decoding from language models
Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., Chen, J., Beutel, A., and Beirami, A · 2024
Closest in time.
Treebon: Enhancing inference-time alignment with speculative tree-search and best-of-n sampling
Qiu, J., Lu, Y., Zeng, Y., Guo, J., Geng, J., Wang, H., Huang, K., Wu, Y., and Wang, M · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beetham, J., Chakraborty, S., Wang, M., Huang, F., Bedi, A. S., and Shah, M · 2024
Cited alongside, same era.
Large language monkeys: Scaling inference compute with repeated sampling
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
Reward model ensembles help mitigate overoptimization
Coste, T., Anwar, U., Kirk, R., and Krueger, D · 2024
Cited alongside, same era.
Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking
Eisenstein, J., Nagpal, C., Agarwal, A., Beirami, A., D’Amour, A., Dvijotham, D., Fisch, A., Heller, K., Pfohl, S., Ramachandran, D., Shaw, P., and Berant, J · 2024
Cited alongside, same era.
Gemma 2: Improving open language models at a practical size
Gemma, T., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Cited alongside, same era.
BoNBoN alignment for large language models and the sweetness of best-of-n sampling
Gui, L., Gârbacea, C., and Veitch, V · 2024
Cited alongside, same era.
Measuring Goodhart’s law, April 2022
Hilton, J. and Gao, L · 2024
Cited alongside, same era.
On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting
Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M
Cited in the paper.
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Closest in time.
Fast best-of-n decoding via speculative rejection
Sun, H., Haider, M., Zhang, R., Yang, H., Qiu, J., Yin, M., Wang, M., Bartlett, P., and Zanette, A · 2024
Closest in time.
Variational best-of-n alignment
Amini, A., Vieira, T., and Cotterell, R · 2025
Closest in time.
InfAlign: Inference-aware language model alignment
Balashankar, A., Sun, Z., Berant, J., Eisenstein, J., Collins, M., Hutter, A., Lee, J., Nagpal, C., Prost, F., Sinha, A., Suresh, A. T., and Beirami, A · 2025
Closest in time.
Guaranteed generation from large language models
Kim, M., Thonet, T., Rozen, J., Lee, H., Jung, K., and Dymetman, M · 2025
Closest in time.
Bond: Aligning llms with best-of-n distillation
Sessa, P. G., Dadashi, R., Hussenot, L., Ferret, J., Vieillard, N., Ramé, A., Shariari, B., Perrin, S., Friesen, A., Cideron, G., et al · 2025
Closest in time.