Fetching the paper…
Reading the bibliography…
Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A Olshausen and David J Field · 1997
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomáš Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment, 2017
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy, 2018
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2018
Earlier work this paper cites.
GPT-2 Output Dataset
OpenAI · 2019
Earlier work this paper cites.
Investigating learning dynamics of bert fine-tuning
Yaru Hao, Li Dong, Furu Wei, and Ke Xu · 2020
Earlier work this paper cites.
Interpreting gpt: The logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Burns, Saurav Kadavath, Akul Arora, Steven Basart, Dawn Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Toy models of superposition, 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in llms, 2023
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans · 2023
Earlier work this paper cites.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning, 2023
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, et al · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable features in language models, 2023
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Earlier work this paper cites.
Toxicity in chatgpt: Analyzing persona-assigned language models, 2023
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Earlier work this paper cites.
Finding neurons in a haystack: Case studies with sparse probing
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Earlier work this paper cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation, 2023
Rusheb Shah, Quentin Feuillade-Montixi, Soroush Pour, Arush Tagade, Stephen Casper, and Javier Rando · 2023
Earlier work this paper cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Earlier work this paper cites.
Linear representations of sentiment in large language models, 2023
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda · 2023
Earlier work this paper cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Earlier work this paper cites.
LMSYS-chat-1m: A large-scale real-world LLM conversation dataset
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P Xing, et al · 2023
Earlier work this paper cites.
Refusal in language models is mediated by a single direction, 2024
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda · 2024
Cited alongside, same era.
Using dictionary learning features as classifiers, October 2024a
Trenton Bricken, Jonathan Marcus, Siddharth Mishra-Sharma, Meg Tong, Ethan Perez, Mrinank Sharma, Kelley Rivoire, Thomas Henighan, and Adam Jermyn · 2024
Cited alongside, same era.
Stage–wise model diffing, December 2024b
Trenton Bricken, Siddharth Mishra-Sharma, Jonathan Marcus, Adam Jermyn, Christopher Olah, Kelley Rivoire, and Thomas Henighan · 2024
Cited alongside, same era.
Improving steering vectors by targeting sparse autoencoder features, 2024
Sviatoslav Chalnev, Matthew Siu, and Arthur Conmy · 2024
Cited alongside, same era.
Sycophancy to subterfuge: Investigating reward-tampering in large language models, 2024
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet, 2024
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, et al · 2024
Later among the works it cites.
Steering language models with activation engineering, 2024
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid · 2024
Later among the works it cites.
Circuit tracing: Revealing computational graphs in language models
Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Nicholas L. Turner, Brian Chen, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Rivoire, Thomas Conerly, Chris Olah, and Joshua Batson · 2025
Closest in time.
System card: Claude opus 4 & claude sonnet 4
Anthropic · 2025
Closest in time.
Saes are good for steering – if you select the right features, 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, and Evan Hubinger · 2024
Cited alongside, same era.
Evaluating feature steering: A case study in mitigating social biases, 2024
Esin Durmus, Alex Tamkin, Jack Clark, Jerry Wei, et al · 2024
Cited alongside, same era.
Scaling and evaluating sparse autoencoders, 2024
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Cited alongside, same era.
Who’s asking? user personas and the mechanics of latent misalignment, 2024
Asma Ghandeharioun, Ann Yuan, Marius Guerard, Emily Reif, Michael A. Lepori, and Lucas Dixon · 2024
Cited alongside, same era.
What is in your safe data? identifying benign data that breaks safety, 2024
Luxi He, Mengzhou Xia, and Peter Henderson · 2024
Cited alongside, same era.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al · 2024
Cited alongside, same era.
Personas as a way to model truthfulness in language models, 2024
Nitish Joshi, Javier Rando, Abulhair Saparov, Najoung Kim, and He He · 2024
Cited alongside, same era.
Sieve: SAEs beat baselines on a real-world task (a code generation case study), 2024
Adam Karvonen, Dhruv Pai, Mason Wang, and Ben Keigwin · 2024
Cited alongside, same era.
Dana Arad, Aaron Mueller, and Yonatan Belinkov · 2025
Closest in time.
Monitoring reasoning models for misbehavior and the risks of promoting obfuscation
Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi · 2025
Closest in time.
Investigating truthfulness in a pre-release o3 model
Neil Chowdhury, Daniel Johnson, Vincent Huang, Jacob Steinhardt, and Sarah Schwettmann · 2025
Closest in time.
Thought crime: Backdoors and emergent misalignment in reasoning models, 2025
James Chua, Jan Betley, Mia Taylor, and Owain Evans · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Interim research report: Mechanisms of awareness, 2025
Josh Engels, Neel Nanda, and Senthooran Rajamanoharan · 2025
Closest in time.
Training on documents about reward hacking induces reward hacking, 2025
Evan Hubinger and Nathan Hu · 2025
Closest in time.
Codenn: A neural network model for source code summarization
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer · 2025
Closest in time.
When bad data leads to good models, 2025
Kenneth Li, Yida Chen, Fernanda Viégas, and Martin Wattenberg · 2025
Closest in time.
Auditing language models for hidden objectives, 2025
Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra-Sharma, Daniel Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, Samuel R. Bowman, Shan Carter, Brian Chen, Hoagy Cunningham, Carson Denison, Florian Dietz, Satvik Golechha, Akbir Khan, Jan Kirchner, Jan Leike, Austin Meek, Kei Nishimura-Gasparian, Euan Ong, Christopher Olah, Adam Pearce, Fabien Roger, Jeanne Salle, Andy Shih, Meg Tong, Drake Thomas, Kelley Rivoire, Adam Jermyn, Monte MacDiarmid, Tom Henighan, and Evan Hubinger · 2025
Closest in time.
Robustly identifying concepts introduced during chat fine-tuning using crosscoders, 2025
Julian Minder, Clement Dumas, Caden Juang, Bilal Chugtai, and Neel Nanda · 2025
Closest in time.
Sycophancy in gpt-4o
OpenAI · 2025
Closest in time.
Llms know more than they show: On the intrinsic representation of llm hallucinations, 2025
Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, and Yonatan Belinkov · 2025
Closest in time.
Convergent linear representations of emergent misalignment, 2025
Anna Soligo, Edward Turner, Senthooran Rajamanoharan, and Neel Nanda · 2025
Closest in time.
Investigating task-specific prompts and sparse autoencoders for activation monitoring
Henk Tillman and Dan Mossing · 2025
Closest in time.
Model organisms for emergent misalignment, 2025
Edward Turner, Anna Soligo, Mia Taylor, Senthooran Rajamanoharan, and Neel Nanda · 2025
Closest in time.
Towards system 2 reasoning in llms: Learning how to think with meta chain-of-thought
Violet Xiang, Charlie Snell, Kanishk Gandhi, Alon Albalak, Anikait Singh, Chase Blagden, Duy Phung, Rafael Rafailov, Nathan Lile, Dakota Mahan, Louis Castricato, Jan‑Philipp Fränken, Nick Haber, and Chelsea Finn · 2025
Closest in time.
Representation engineering: A top-down approach to ai transparency, 2025
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2025
Closest in time.