Fetching the paper…
Reading the bibliography…
Safety alignment is an important procedure before the official deployment of a Large Language Model (LLM).
Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing
Zhou, Z., Chen, X., Li, E., Zeng, L., Luo, K., and Zhang, J · 1905
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Solving math word problems with process-and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Safe rlhf: Safe reinforcement learning from human feedback
Dai, J., Pan, X., Sun, R., Ji, J., Xu, X., Liu, M., Wang, Y., and Yang, Y · 2023
Earlier work this paper cites.
Raft: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Earlier work this paper cites.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Sun, R., Wang, Y., and Yang, Y · 2023
Earlier work this paper cites.
Let’s verify step by step
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
Training socially aligned language models in simulated human society
Liu, R., Yang, R., Jia, C., Zhang, G., Zhou, D., Dai, A. M., Yang, D., and Vosoughi, S · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Earlier work this paper cites.
Math-shepherd: A label-free step-by-step verifier for llms in mathematical reasoning
Wang, P., Li, L., Shao, Z., Xu, R., Dai, D., Li, Y., Chen, D., Wu, Y., and Sui, Z · 2023
Earlier work this paper cites.
Pairwise proximal policy optimization: Harnessing relative feedback for llm alignment
Wu, T., Zhu, B., Zhang, R., Wen, Z., Ramchandran, K., and Jiao, J · 2023
Cited alongside, same era.
Selfee: Iterative self-revising llm empowered by self-feedback generation
Ye, S., Jo, Y., Kim, D., Kim, S., Hwang, H., and Seo, M · 2023
Cited alongside, same era.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 07 2024
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2024
Cited alongside, same era.
Harmful fine-tuning attacks and defenses for large language models: A survey
Are smarter llms safer? exploring safety-reasoning trade-offs in prompting and fine-tuning
Li, A., Mo, Y., Li, M., Wang, Y., and Wang, Y · 2025
Closest in time.
There may not be aha moment in r1-zero-like training — a pilot study
Liu, Z., Chen, C., Li, W., Pang, T., Du, C., and Lin, M · 2025
Closest in time.
O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning
Luo, H., Shen, L., He, H., Wang, Y., Liu, S., Li, W., Tan, N., Cao, X., and Tao, D · 2025
Closest in time.
Cot-valve: Length-compressible chain-of-thought tuning
Ma, X., Wan, G., Yu, R., Fang, G., and Wang, X · 2025
Closest in time.
Muennighoff, N., Yang, Z., Shi, W., Li, X. L., Fei-Fei, L., Hajishirzi, H., Zettlemoyer, L., Liang, P., Candès, E., and Hashimoto, T · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huang, T., Hu, S., Ilhan, F., Tekin, S. F., and Liu, L · 2024
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2024
Cited alongside, same era.
Representation noising effectively prevents harmful fine-tuning on llms
Rosati, D., Wehner, J., Williams, K., Bartoszcze, Ł., Atanasov, D., Gonzales, R., Majumdar, S., Maple, C., Sajjad, H., and Rudzicz, F · 2024
Cited alongside, same era.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al · 2024
Cited alongside, same era.
Hˆ 3 fusion: Helpful, harmless, honest fusion of aligned llms
Tekin, S. F., Ilhan, F., Huang, T., Hu, S., Yahn, Z., and Liu, L · 2024
Cited alongside, same era.
Monte carlo tree search boosts reasoning via iterative preference learning
Xie, Y., Goyal, A., Zheng, W., Kan, M.-Y., Lillicrap, T. P., Kawaguchi, K., and Shieh, M · 2024
Cited alongside, same era.
Improving alignment and robustness with circuit breakers
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Kolter, J. Z., Fredrikson, M., and Hendrycks, D · 2024
Cited alongside, same era.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Cited alongside, same era.
Closest in time.
Tinyzero
Pan, J., Zhang, J., Wang, X., Yuan, L., Peng, H., and Suhr, A · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Team, K., Du, A., Gao, B., Xing, B., Jiang, C., Chen, C., Li, C., Xiao, C., Du, C., Liao, C., et al · 2025
Closest in time.
Xu, Z., Gardiner, J., and Belguith, S · 2025
Closest in time.
Yang, J., Jin, D., Tang, A., Shen, L., Zhu, D., Chen, Z., Wang, D., Cui, Q., Zhang, Z., Zhou, J., et al · 2025
Closest in time.
Limo: Less is more for reasoning
Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., and Liu, P · 2025
Closest in time.
7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient
Zeng, W., Huang, Y., Liu, W., He, K., Liu, Q., Ma, Z., and He, J · 2025
Closest in time.
The hidden risks of large reasoning models: A safety assessment of r1
Zhou, K., Liu, C., Zhao, X., Jangam, S., Srinivasa, J., Liu, G., Song, D., and Wang, X. E · 2025
Closest in time.
Bot: Breaking long thought processes of o1-like large language models through backdoor attack
Zhu, Z., Zhang, H., Zhang, M., Wang, R., Wu, G., Xu, K., and Wu, B · 2025
Closest in time.