Fetching the paper…
Reading the bibliography…
Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions.
Probable inference, the law of succession, and statistical inference
E. B. Wilson · 1927
Earlier work this paper cites.
Building blocks for assurance cases
R. Bloomfield and K. Netkachova · 2014
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant · 2019
Earlier work this paper cites.
Initial analysis of underhanded source code
D. Wheeler · 2020
Earlier work this paper cites.
Is power-seeking AI an existential risk?
J. Carlsmith · 2022
Earlier work this paper cites.
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover, 2022
A. Cotra · 2022
Earlier work this paper cites.
Understanding strategic deception and deceptive alignment
Apollo · 2023
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in LLMs
L. Berglund, A. C. Stickland, M. Balesni, M. Kaufmann, M. Tong, T. Korbak, D. Kokotajlo, and O. Evans · 2023
Earlier work this paper cites.
Scheming AIs: Will AIs fake alignment during training in order to get power?
J. Carlsmith · 2023
Earlier work this paper cites.
AI deception: A survey of examples, risks, and potential solutions
P. S. Park, S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks · 2023
Earlier work this paper cites.
Model evaluation for extreme risks
T. Shevlane, S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal, M. Anderljung, N. Kolt, L. Ho, D. Siddarth, S. Avin, W. Hawkins, B. Kim, I. Gabriel, V. Bolina, J. Clark, Y. Bengio, P. Christiano, and A. Dafoe · 2023
Earlier work this paper cites.
Responsible scaling policy, 2024
Anthropic · 2024
Cited alongside, same era.
Towards evaluations-based safety cases for AI scheming
M. Balesni, M. Hobbhahn, D. Lindner, A. Meinke, T. Korbak, J. Clymer, B. Shlegeris, J. Scheurer, C. Stix, R. Shah, N. Goldowsky-Dill, D. Braun, B. Chughtai, O. Evans, D. Kokotajlo, and L. Bushnaq · 2024
Cited alongside, same era.
Sabotage evaluations for frontier models
J. Benton, M. Wagner, E. Christiansen, C. Anil, E. Perez, J. Srivastav, E. Durmus, D. Ganguli, S. Kravec, B. Shlegeris, J. Kaplan, H. Karnofsky, E. Hubinger, R. Grosse, S. R. Bowman, and D. Duvenaud · 2024
Cited alongside, same era.
M. D. Buhl, G. Sett, L. Koessler, J. Schuett, and M. Anderljung · 2024
Cited alongside, same era.
Frontier models are capable of in-context scheming
A. Meinke, B. Schoen, J. Scheurer, M. Balesni, R. Shah, and M. Hobbhahn · 2024
Later among the works it cites.
The alignment problem from a deep learning perspective
R. Ngo, L. Chan, and S. Mindermann · 2024
Later among the works it cites.
Evaluating frontier models for dangerous capabilities
M. Phuong, M. Aitchison, E. Catt, S. Cogan, A. Kaskasoli, V. Krakovna, D. Lindner, M. Rahtz, Y. Assael, S. Hodkinson, et al · 2024
Later among the works it cites.
Large language models can strategically deceive their users when put under pressure
J. Scheurer, M. Balesni, and M. Hobbhahn · 2024
Later among the works it cites.
Automated researchers can subtly sandbag
J. Gasteiger, A. Khan, S. Bowman, V. Mikulik, E. Perez, and F. Roger · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Clymer, N. Gabrieli, D. Krueger, and T. Larsen · 2024
Cited alongside, same era.
Sycophancy to subterfuge: Investigating reward-tampering in large language models
C. Denison, M. MacDiarmid, F. Barez, D. Duvenaud, S. Kravec, S. Marks, N. Schiefer, R. Soklaski, A. Tamkin, J. Kaplan, B. Shlegeris, S. R. Bowman, E. Perez, and E. Hubinger · 2024
Cited alongside, same era.
MISR: Measuring instrumental self-reasoning in frontier models
K. Fronsdal and D. Lindner · 2024
Cited alongside, same era.
Safety case template for frontier AI: A cyber inability argument
A. Goemans, M. D. Buhl, J. Schuett, T. Korbak, J. Wang, B. Hilton, and G. Irving · 2024
Cited alongside, same era.
The case for ensuring that powerful AIs are controlled, 2024
R. Greenblatt and B. Shlegeris · 2024
Cited alongside, same era.
Safety cases at AISI
G. Irving · 2024
Cited alongside, same era.
Me, myself, and AI: The situational awareness dataset (SAD) for LLMs
R. Laine, B. Chughtai, J. Betley, K. Hariharan, M. Balesni, J. Scheurer, M. Hobbhahn, A. Meinke, and O. Evans · 2024
Cited alongside, same era.
Alignment faking in large language models
R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, et al
Cited in the paper.
Updating the Frontier Safety Framework
Google DeepMind · 2025
Closest in time.
A. Mallen, C. Griffin, M. Wagner, A. Abate, and B. Shlegeris · 2025
Closest in time.
An approach to technical AGI safety and security
R. Shah, A. Irpan, A. M. Turner, A. Wang, A. Conmy, D. Lindner, J. Brown-Cohen, L. Ho, N. Nanda, R. A. Popa, R. Jain, R. Greig, S. Albanie, S. Emmons, S. Farquhar, S. Krier, S. Rajamanoharan, S. Bridgers, T. Ijitoye, T. Everitt, V. Krakovna, V. Varma, V. Mikulik, Z. Kenton, D. Orr, S. Legg, N. Goodman, A. Dafoe, F. Flynn, and A. Dragan · 2025
Closest in time.
AI sandbagging: Language models can strategically underperform on evaluations
T. van der Weij, F. Hofstätter, O. Jaffe, S. F. Brown, and F. R. Ward · 2025
Closest in time.
Adaptive deployment of untrusted LLMs reduces distributed threats
J. Wen, V. Hebbar, C. Larson, A. Bhatt, A. Radhakrishnan, M. Sharma, H. Sleight, S. Feng, H. He, E. Perez, B. Shlegeris, and A. Khan · 2025
Closest in time.