Fetching the paper…
Reading the bibliography…
Language model alignment is a critical step in training modern generative language models.
Remarks on a multivariate transformation
Rosenblatt, M · 1952
Earlier work this paper cites.
Asymptotic Minimax Character of the Sample Distribution Function and of the Classical Multinomial Estimator
Dvoretzky, A., Kiefer, J., and Wolfowitz, J · 1956
Earlier work this paper cites.
Coarse-to-fine n-best parsing and MaxEnt discriminative reranking
Charniak, E. and Johnson, M · 2005
Earlier work this paper cites.
Discriminative reranking for natural language parsing
Collins, M. and Koo, T · 2005
Earlier work this paper cites.
K-best A* parsing
Pauls, A. and Klein, D · 2009
Earlier work this paper cites.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Tilted empirical risk minimization
Li, T., Beirami, A., Sanjabi, M., and Smith, V · 2021
Earlier work this paper cites.
To beam or not to beam: That is a question of cooperation for language gans
Scialom, T., Dray, P.-A., Staiano, J., Lamprier, S., and Piwowarski, B · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
PPL-MCTS: Constrained textual generation through discriminator-guided MCTS decoding
Chaffin, A., Claveau, V., and Kijak, E · 2022
Earlier work this paper cites.
RL with KL penalties is better viewed as Bayesian inference
Korbak, T., Perez, E., and Buckley, C. L · 2022
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback, 2022
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
Pan, A., Bhatia, K., and Steinhardt, J · 2022
Earlier work this paper cites.
Reward gaming in conditional text generation
Pang, R. Y., Padmakumar, V., Sellam, T., Parikh, A. P., and He, H · 2022
Earlier work this paper cites.
Defining and characterizing reward gaming
Skalse, J., Howe, N., Krasheninnikov, D., and Krueger, D · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Cited alongside, same era.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Rowland, M., Piot, B., Guo, D., Calandriello, D., Valko, M., and Munos, R · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2024
Closest in time.
Information theoretic guarantees for policy alignment in large language models
Mroueh, Y · 2024
Closest in time.
Controlled decoding from language models
Mudgal, S., Lee, J., Ganapathy, H., Li, Y., Wang, T., Huang, Y., Chen, Z., Cheng, H.-T., Collins, M., Strohman, T., Chen, J., Beutel, A., and Beirami, A · 2024
Closest in time.
Learning to reason with llms
OpenAI · 2024
Closest in time.
Treebon: Enhancing inference-time alignment with speculative tree-search and best-of-n sampling
Qiu, J., Lu, Y., Zeng, Y., Guo, J., Geng, J., Wang, H., Huang, K., Wu, Y., and Wang, M · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Coste, T., Anwar, U., Kirk, R., and Krueger, D · 2023
Cited alongside, same era.
Scaling laws for reward model overoptimization
Gao, L., Schulman, J., and Hilton, J · 2023
Cited alongside, same era.
On tilted losses in machine learning: Theory and applications
Li, T., Beirami, A., Sanjabi, M., and Smith, V · 2023
Cited alongside, same era.
Flirt: Feedback loop in-context red teaming
Mehrabi, N., Goyal, P., Dupuy, C., Hu, Q., Ghosh, S., Zemel, R., Chang, K.-W., Galstyan, A., and Gupta, R · 2023
Cited alongside, same era.
Confronting reward model overoptimization with constrained rlhf
Moskovitz, T., Singh, A. K., Strouse, D., Sandholm, T., Salakhutdinov, R., Dragan, A. D., and McAleer, S · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Cited alongside, same era.
Variational best-of-n alignment, 2024
Amini, A., Vieira, T., and Cotterell, R · 2024
Cited alongside, same era.
Liar: Leveraging alignment (best-of-n) to jailbreak llms in seconds
Beetham, J., Chakraborty, S., Wang, M., Huang, F., Bedi, A. S., and Shah, M · 2024
Cited alongside, same era.
Sessa, P. G., Dadashi, R., Hussenot, L., Ferret, J., Vieillard, N., Ramé, A., Shariari, B., Perrin, S., Friesen, A., Cideron, G., Girgin, S., Stanczyk, P., Michi, A., Sinopalnikov, D., Ramos, S., Héliou, A., Severyn, A., Hoffman, M., Momchev, N., and Bachem, O · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Closest in time.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Closest in time.
A strongreject for empty jailbreaks
Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., et al · 2024
Closest in time.
Transforming and combining rewards for aligning large language models
Wang, Z., Nagpal, C., Berant, J., Eisenstein, J., D’Amour, A., Koyejo, S., and Veitch, V · 2024
Closest in time.
From decoding to meta-generation: Inference-time algorithms for large language models
Welleck, S., Bertsch, A., Finlayson, M., Schoelkopf, H., Xie, A., Neubig, G., Kulikov, I., and Harchaoui, Z · 2024
Closest in time.
Asymptotics of language model alignment
Yang, J. Q., Salamatian, S., Sun, Z., Suresh, A. T., and Beirami, A · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Closest in time.
International Scientific Report on the Safety of Advanced AI
Yohsua, B., Daniel, P., Tamay, B., Rishi, B., Stephen, C., Yejin, C., Danielle, G., Hoda, H., Leila, K., Shayne, L., et al · 2024
Closest in time.
Backtracking improves generation safety, 2024
Zhang, Y., Chi, J., Nguyen, H., Upasani, K., Bikel, D. M., Weston, J., and Smith, E. M · 2024
Closest in time.
Probabilistic inference in language models via twisted sequential monte carlo
Zhao, S., Brekelmans, R., Makhzani, A., and Grosse, R. B · 2024
Closest in time.
Calibrating sequence likelihood improves conditional language generation
Zhao, Y., Khalman, M., Joshi, R., Narayan, S., Saleh, M., and Liu, P. J · 2024
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
Guaranteed generation from large language models
Kim, M., Thonet, T., Rozen, J., Lee, H., Jung, K., and Dymetman, M · 2025
Closest in time.
RewardBench: Evaluating reward models for language modeling
Lambert, N., Pyatkin, V., Morrison, J., Miranda, L., Lin, B. Y., Chandu, K., Dziri, N., Kumar, S., Zick, T., Choi, Y., Smith, N. A., and Hajishirzi, H · 2025
Closest in time.