Fetching the paper…
Reading the bibliography…
While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on "unfaithfulness" has cast doubt on its reliability.
Sympy: symbolic computing in python
A. Meurer, C. P. Smith, M. Paprocki, O. Čertík, S. B. Kirpichev, M. Rocklin, A. Kumar, S. Ivanov, J. K. Moore, S. Singh, T. Rathnayake, S. Vig, B. E. Granger, R. P. Muller, F. Bonazzi, H. Gupta, S. Vats, F. Johansson, F. Pedregosa, M. J. Curry, A. R. Terrel, v. Roučka, A. Saboo, I. Fernando, S. Kulal, R. Cimrman, and A. Scopatz · 2017
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
D. Saxton, E. Grefenstette, F. Hill, and P. Kohli · 2019
Earlier work this paper cites.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?
A. Jacovi and Y. Goldberg · 2020
Earlier work this paper cites.
By Default, GPTs Think In Plain Sight
F. Roger · 2022
Earlier work this paper cites.
Scheming ais: Will ais fake alignment during training in order to get power?, 2023
J. Carlsmith · 2023
Earlier work this paper cites.
Measuring faithfulness in chain-of-thought reasoning, 2023
T. Lanham, A. Chen, A. Radhakrishnan, B. Steiner, C. Denison, D. Hernandez, D. Li, E. Durmus, E. Hubinger, J. Kernion, K. Lukošiūtė, K. Nguyen, N. Cheng, N. Joseph, N. Schiefer, O. Rausch, R. Larson, S. McCandlish, S. Kundu, S. Kadavath, S. Yang, T. Henighan, T. Maxwell, T. Telleen-Lawton, T. Hume, Z. Hatfield-Dodds, J. Kaplan, J. Brauner, S. R. Bowman, and E. Perez · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark, 2023
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2023
Earlier work this paper cites.
The Translucent Thoughts Hypotheses and Their Implications
F. Roger · 2023
Earlier work this paper cites.
Preventing language models from hiding their reasoning, 2023
F. Roger and R. Greenblatt · 2023
Earlier work this paper cites.
LLMs are (mostly) not helped by filler tokens
K. Sachan · 2023
Earlier work this paper cites.
M. Turpin, J. Michael, E. Perez, and S. R. Bowman · 2023
Earlier work this paper cites.
Towards evaluations-based safety cases for ai scheming, 2024
M. Balesni, M. Hobbhahn, D. Lindner, A. Meinke, T. Korbak, J. Clymer, B. Shlegeris, J. Scheurer, C. Stix, R. Shah, N. Goldowsky-Dill, D. Braun, B. Chughtai, O. Evans, D. Kokotajlo, and L. Bushnaq · 2024
Earlier work this paper cites.
Sabotage evaluations for frontier models, 2024
J. Benton, M. Wagner, E. Christiansen, C. Anil, E. Perez, J. Srivastav, E. Durmus, D. Ganguli, S. Kravec, B. Shlegeris, J. Kaplan, H. Karnofsky, E. Hubinger, R. Grosse, S. R. Bowman, and D. Duvenaud · 2024
Cited alongside, same era.
Safety cases: How to justify the safety of advanced ai systems, 2024
J. Clymer, N. Gabrieli, D. Krueger, and T. Larsen · 2024
Cited alongside, same era.
Ai control: Improving safety despite intentional subversion, 2024
R. Greenblatt, B. Shlegeris, K. Sachan, and F. Roger · 2024
Cited alongside, same era.
Training large language models to reason in a continuous latent space, 2024
S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian · 2024
Cited alongside, same era.
Cot red-handed: Stress testing chain-of-thought monitoring, 2025
B. Arnav, P. Bernabeu-Pérez, N. Helm-Burger, T. Kostolansky, H. Whittingham, and M. Phuong · 2025
Closest in time.
Detecting misbehavior in frontier reasoning models
B. Baker, J. Huizinga, L. Gao, Z. Dou, M. Y. Guan, A. Madry, W. Zaremba, J. Pachocki, and D. Farhi · 2025
Closest in time.
Reasoning models don’t always say what they think
Y. Chen, J. Benton, A. Radhakrishnan, J. Uesato, C. Denison, J. Schulman, A. Somani, P. Hase, M. Wagner, F. Roger, V. Mikulik, S. Bowman, J. Leike, J. Kaplan, and E. Perez · 2025
Closest in time.
Are deepseek r1 and other reasoning models more faithful?, 2025
J. Chua and O. Evans · 2025
Closest in time.
Mona: Myopic optimization with non-myopic approval can mitigate multi-step reward hacking, 2025
S. Farquhar, V. Varma, D. Lindner, D. Elson, C. Biddulph, I. Goodfellow, and R. Shah · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Li, H. Liu, D. Zhou, and T. Ma · 2024
Cited alongside, same era.
Deliberation in latent space via differentiable cache augmentation, 2024
L. Liu, J. Pfeiffer, J. Wu, J. Xie, and A. Szlam · 2024
Cited alongside, same era.
Hidden in plain text: Emergence & mitigation of steganographic collusion in llms, 2024
Y. Mathew, O. Matthews, R. McCarthy, J. Velja, C. S. de Witt, D. Cope, and N. Schoots · 2024
Cited alongside, same era.
the case for CoT unfaithfulness is overstated
nostalgebraist · 2024
Cited alongside, same era.
Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning, 2024
D. Paul, R. West, A. Bosselut, and B. Faltings · 2024
Cited alongside, same era.
Let’s think dot by dot: Hidden computation in transformer language models, 2024
J. Pfau, W. Merrill, and S. R. Bowman · 2024
Cited alongside, same era.
Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases
Anthropic · 2025
Cited alongside, same era.
Chain-of-thought reasoning in the wild is not always faithful, 2025
I. Arcuschin, J. Janiak, R. Krzyzanowski, S. Rajamanoharan, N. Nanda, and A. Conmy · 2025
Cited alongside, same era.
Closest in time.
Scaling up test-time compute with latent reasoning: A recurrent depth approach, 2025
J. Geiping, S. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein · 2025
Closest in time.
Safety cases: A scalable approach to frontier ai safety, 2025
B. Hilton, M. D. Buhl, T. Korbak, and G. Irving · 2025
Closest in time.
Robustly improving llm fairness in realistic settings via interpretability, 2025
A. Karvonen and S. Marks · 2025
Closest in time.
Secret collusion among generative ai agents: Multi-agent deception via steganography, 2025
S. R. Motwani, M. Baranchuk, M. Strohmeier, V. Bolina, P. H. S. Torr, L. Hammond, and C. S. de Witt · 2025
Closest in time.
Evaluating frontier models for stealth and situational awareness, 2025
M. Phuong, R. S. Zimmermann, Z. Wang, D. Lindner, V. Krakovna, S. Cogan, A. Dafoe, L. Ho, and R. Shah · 2025
Closest in time.
An approach to technical agi safety and security, 2025
R. Shah, A. Irpan, A. M. Turner, A. Wang, A. Conmy, D. Lindner, J. Brown-Cohen, L. Ho, N. Nanda, R. A. Popa, R. Jain, R. Greig, S. Albanie, S. Emmons, S. Farquhar, S. Krier, S. Rajamanoharan, S. Bridgers, T. Ijitoye, T. Everitt, V. Krakovna, V. Varma, V. Mikulik, Z. Kenton, D. Orr, S. Legg, N. Goodman, A. Dafoe, F. Flynn, and A. Dragan · 2025
Closest in time.