Fetching the paper…
Reading the bibliography…
We present a theoretical framework showing that popular LLM alignment methods, including RLHF and its variants, can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less preferred) distributions.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time, ii
Donsker, M. D. and Varadhan, S · 1975
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Mine: mutual information neural estimation
Belghazi, M. I., Baratin, A., Rajeswar, S., Ozair, S., Bengio, Y., Courville, A., and Hjelm, R. D · 2018
Earlier work this paper cites.
Improved adam optimizer for deep neural networks
Zhang, Z · 2018
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., and Kamar, E · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Earlier work this paper cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Earlier work this paper cites.
Spear phishing with large language models
Hazell, J · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Earlier work this paper cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Liu, X., Xu, N., Chen, M., and Xiao, C · 2023
Earlier work this paper cites.
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Is rlhf more difficult than standard rl? a theoretical perspective
Wang, Y., Liu, Q., and Jin, C · 2023
Cited alongside, same era.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Cited alongside, same era.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., Yu, Z., and Xing, X · 2023
Cited alongside, same era.
Adding conditional control to text-to-image diffusion models
Zhang, L., Rao, A., and Agrawala, M · 2023
Binary classifier optimization for large language model alignment
Jung, S., Han, G., Nam, D. W., and On, K.-W · 2024
Later among the works it cites.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M., and Hashimoto, T · 2024
Later among the works it cites.
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Li, L., Dong, B., Wang, R., Hu, X., Zuo, W., Lin, D., Qiao, Y., and Shao, J · 2024
Later among the works it cites.
Towards understanding jailbreak attacks in LLMs: A representation space analysis
Lin, Y., He, P., Xu, H., Xing, Y., Yamada, M., Liu, H., and Tang, J · 2024
Later among the works it cites.
Dual active learning for reinforcement learning from human feedback
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Synthetic lies: Understanding ai-generated misinformation and evaluating algorithmic and human solutions
Zhou, J., Zhang, Y., Luo, Q., Parker, A. G., and De Choudhury, M · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Direct preference optimization with an offset
Amini, A., Vieira, T., and Cotterell, R · 2024
Cited alongside, same era.
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Koh, P. W. W., Ippolito, D., Tramer, F., and Schmidt, L · 2024
Cited alongside, same era.
Exploration-driven policy optimization in rlhf: Theoretical insights on efficient data utilization
Du, Y., Winnicki, A., Dalal, G., Mannor, S., and Srikant, R · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Dubois, Y., Galambosi, B., Liang, P., and Hashimoto, T. B · 2024
Cited alongside, same era.
Liu, P., Shi, C., and Sun, W. W · 2024
Later among the works it cites.
Online merging optimizers for boosting rewards and mitigating tax in alignment
Lu, K., Yu, B., Huang, F., Fan, Y., Lin, R., and Zhou, C · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2024
Later among the works it cites.
Xiao, J., Li, Z., Xie, X., Getzen, E., Fang, C., Long, Q., and Su, W. J · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Entropy law: The story behind data compression and llm performance
Yin, M., Wu, C., Wang, Y., Wang, H., Guo, W., Wang, Y., Liu, Y., Tang, R., Lian, D., and Chen, E · 2024
Later among the works it cites.
Advancing llm reasoning generalists with preference trees
Yuan, L., Cui, G., Wang, H., Ding, N., Wang, X., Deng, J., Shan, B., Chen, H., Xie, R., Lin, Y., et al · 2024
Later among the works it cites.
Self-exploring language models: Active preference elicitation for online alignment
Zhang, S., Yu, D., Sharma, H., Zhong, H., Liu, Z., Yang, Z., Wang, S., Hassan, H., and Wang, Z · 2024
Later among the works it cites.
On prompt-driven safeguarding for large language models
Zheng, C., Yin, F., Zhou, H., Meng, F., Zhou, J., Chang, K.-W., Huang, M., and Peng, N · 2024
Later among the works it cites.
T-reg: Preference optimization with token-level reward regularization
Zhou, W., Zhang, S., Zhao, L., and Meng, T · 2024
Later among the works it cites.