Fetching the paper…
Reading the bibliography…
Large Reasoning Models (LRMs), such as OpenAI o1 and DeepSeek-R1, have been rapidly progressing and achieving breakthrough performance on complex reasoning tasks such as mathematics and coding.
Rusu, A. A., Colmenarejo, S. G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R. (2015) · 2015
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O. (2021) · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. (2022) · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. (2022) · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023) · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. (2023) · 2023
Earlier work this paper cites.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K. (2023) · 2023
Earlier work this paper cites.
Liu, W., Zeng, W., He, K., Jiang, Y., and He, J. (2023) · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. (2023) · 2023
Earlier work this paper cites.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Röttger, P., Kirk, H. R., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D. (2023) · 2023
Earlier work this paper cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Earlier work this paper cites.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., Truong, S., Arora, S., Mazeika, M., Hendrycks, D., Lin, Z., Cheng, Y., Koyejo, S., Song, D., and Li, B. (2023) · 2023
Cited alongside, same era.
Tree of thoughts: deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. (2023) · 2023
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. (2024) · 2024
Cited alongside, same era.
Length-controlled alpacaeval: A simple debiasing of automatic evaluators
Dubois, Y., Liang, P., and Hashimoto, T. (2024) · 2024
Cited alongside, same era.
Deliberative alignment: Reasoning enables safer language models
Guan, M. Y., Joglekar, M., Wallace, E., Jain, S., Barak, B., Heylar, A., Dias, R., Vallone, A., Ren, H., Wei, J., et al. (2024) · 2024
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R. (2024) · 2024
Later among the works it cites.
A strongreject for empty jailbreaks
Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., et al. (2024) · 2024
Later among the works it cites.
Challenges and barriers of using large language models (llm) such as chatgpt for diagnostic medicine with a focus on digital pathology–a recent scoping review
Ullah, E., Parwani, A., Baig, M. M., and Singh, R. (2024) · 2024
Later among the works it cites.
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W. (2024) · 2024
Later among the works it cites.
Wildchat: 1m chatgpt interaction logs in the wild
Zhao, W., Ren, X., Hessel, J., Cardie, C., Choi, Y., and Deng, Y. (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al. (2024) · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I. (2024) · 2024
Cited alongside, same era.
Pku-saferlhf: A safety alignment preference dataset for llama family models
Ji, J., Hong, D., Zhang, B., Chen, B., Dai, J., Zheng, B., Qiu, T., Li, B., and Yang, Y. (2024) · 2024
Cited alongside, same era.
Luo, W., Ma, S., Liu, X., Guo, X., and Xiao, C. (2024) · 2024
Cited alongside, same era.
Rule based rewards for language model safety
Mu, T., Helyar, A., Heidecke, J., Achiam, J., Vallone, A., Kivlichan, I., Lin, M., Beutel, A., Schulman, J., and Weng, L. (2024) · 2024
Cited alongside, same era.
Using an llm to help with code understanding
Nam, D., Macvean, A., Hellendoorn, V., Vasilescu, B., and Myers, B. (2024) · 2024
Cited alongside, same era.
Rethinking legal judgement prediction in a realistic scenario in the era of large language models
Nigam, S. K., Deroy, A., Maity, S., and Bhattacharya, A. (2024) · 2024
Cited alongside, same era.
Later among the works it cites.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Zheng, Y., Zhang, R., Zhang, J., YeYanhan, Y., and Luo, Z. (2024) · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. (2025) · 2025
Closest in time.
Safety tax: Safety alignment makes your large reasoning models less reasonable
Huang, T., Hu, S., Ilhan, F., Tekin, S. F., Yahn, Z., Xu, Y., and Liu, L. (2025) · 2025
Closest in time.
Safechain: Safety of language models with long chain-of-thought reasoning capabilities
Jiang, F., Xu, Z., Li, Y., Niu, L., Xiang, Z., Li, B., Lin, B. Y., and Poovendran, R. (2025) · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning
Team, Q. (2025) · 2025
Closest in time.
Star-1: Safer alignment of reasoning llms with 1k data
Wang, Z., Tu, H., Wang, Y., Wu, J., Mei, J., Bartoldson, B. R., Kailkhura, B., and Xie, C. (2025) · 2025
Closest in time.
The hidden risks of large reasoning models: A safety assessment of r1
Zhou, K., Liu, C., Zhao, X., Jangam, S., Srinivasa, J., Liu, G., Song, D., and Wang, X. E. (2025) · 2025
Closest in time.