Fetching the paper…
Reading the bibliography…
Large reasoning models (LRMs) have emerged as a significant advancement in artificial intelligence, representing a specialized class of large language models (LLMs) designed to tackle complex reasoning tasks.
Fine-pruning: Defending against backdooring attacks on deep neural networks
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2018
Earlier work this paper cites.
Weight poisoning attacks on pretrained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, et al · 2021
Earlier work this paper cites.
Sponge examples: Energy-latency attacks on neural networks
Ilia Shumailov, Yiren Zhao, Daniel Bates, et al · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, et al · 2021
Earlier work this paper cites.
Be careful about poisoned word embeddings: Exploring the vulnerability of the embedding layers in NLP models
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He · 2021
Earlier work this paper cites.
Backdoor learning: A survey
Yiming Li, Yong Jiang, Zhifeng Li, et al · 2022
Earlier work this paper cites.
Hidden trigger backdoor attack on NLP models via linguistic style manipulation
Xudong Pan, Mi Zhang, Beina Sheng, Jiaming Zhu, and Min Yang · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, et al · 2022
Earlier work this paper cites.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, et al · 2023
Earlier work this paper cites.
Let’s verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, et al · 2023
Earlier work this paper cites.
Trojtext: Test-time invisible textual trojan insertion
Qian Lou, Yepeng Liu, and Bo Feng · 2023
Earlier work this paper cites.
On the exploitability of instruction tuning
Manli Shu, Jiongxiao Wang, Chen Zhu, et al · 2023
Earlier work this paper cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, et al · 2023
Cited alongside, same era.
BITE: textual backdoor attacks with iterative trigger injection
Jun Yan, Vansh Gupta, and Xiang Ren · 2023
Cited alongside, same era.
Automatic chain of thought prompting in large language models
Zhuosheng Zhang, Aston Zhang, Mu Li, et al · 2023
Cited alongside, same era.
Stealthy and persistent unalignment on large language models via backdoor injections
Yuanpu Cao, Bochuan Cao, and Jinghui Chen · 2024
Cited alongside, same era.
Do NOT think that much for 2+3=? on the overthinking of o1-like llms
Xingyu Chen, Jiahao Xu, Tian Liang, et al · 2024
Cited alongside, same era.
Universal jailbreak backdoors from poisoned human feedback
Javier Rando and Florian Tramèr · 2024
Later among the works it cites.
Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Qwen Team · 2024
Later among the works it cites.
Badchain: Backdoor chain-of-thought prompting for large language models
Zhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian, Radha Poovendran, and Bo Li · 2024
Later among the works it cites.
BEEAR: embedding-based adversarial removal of safety backdoors in instruction-tuned language models
Yi Zeng, Weiyu Sun, Tran Ngoc Huynh, et al · 2024
Later among the works it cites.
Marco-o1: Towards open reasoning models for open-ended solutions
Yu Zhao, Huifeng Yin, Bo Zeng, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jianshuo Dong, Ziyuan Zhang, Qingjie Zhang, et al · 2024
Cited alongside, same era.
Denial-of-service poisoning attacks against large language models
Kuofeng Gao, Tianyu Pang, Chao Du, et al · 2024
Cited alongside, same era.
Coercing llms to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu, et al · 2024
Cited alongside, same era.
Exploring backdoor vulnerabilities of chat models
Yunzhuo Hao, Wenkai Yang, and Yankai Lin · 2024
Cited alongside, same era.
Aaron Jaech, Adam Kalai, Adam Lerer, et al · 2024
Cited alongside, same era.
Backdoorllm: A comprehensive benchmark for backdoor attacks on large language models
Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun Sun · 2024
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, et al · 2024
Cited alongside, same era.
The danger of overthinking: Examining the reasoning-action dilemma in agentic tasks
Alejandro Cuadron, Dacheng Li, Wenjie Ma, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Overthink: Slowdown attacks on reasoning llms
Abhinav Kumar, Jaechul Roh, Ali Naseh, et al · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning
QwenTeam · 2025
Closest in time.
An Yang, Baosong Yang, Beichen Zhang, et al · 2025
Closest in time.
Probe before you talk: Towards black-box defense against backdoor unalignment for large language models
Biao Yi, Tiansheng Huang, Sishuo Chen, Tong Li, Zheli Liu, Zhixuan Chu, and Yiming Li · 2025
Closest in time.
Bot: Breaking long thought processes of o1-like large language models through backdoor attack
Zihao Zhu, Hongbao Zhang, Mingda Zhang, et al · 2025
Closest in time.