Fetching the paper…
Reading the bibliography…
Language models can behave in unexpected and unsafe ways, and so it is valuable to monitor their outputs.
Seeing stars: exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, November 2018
Guillaume Alain and Yoshua Bengio · 2018
Earlier work this paper cites.
Interpretability and Analysis in Neural NLP
Yonatan Belinkov, Sebastian Gehrmann, and Ellie Pavlick · 2020
Earlier work this paper cites.
The Internal State of an LLM Knows When It’s Lying, October 2023
Amos Azaria and Tom Mitchell · 2023
Earlier work this paper cites.
Finding Neurons in a Haystack: Case Studies with Sparse Probing, June 2023
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas · 2023
Earlier work this paper cites.
Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, and Jacob Andreas · 2023
Earlier work this paper cites.
Linear Representations of Sentiment in Large Language Models, October 2023
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda · 2023
Earlier work this paper cites.
Representation Engineering: A Top-Down Approach to AI Transparency, October 2023
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2023
Cited alongside, same era.
Using Dictionary Learning Features as Classifiers, October 2024
Trenton Bricken, Jonathan Marcus, Siddharth Mishra-Sharma, Meg Tong, Ethan Perez, Mrinank Sharma, Kelley Rivoire, and Thomas Henighan · 2024
Cited alongside, same era.
Scaling and evaluating sparse autoencoders, June 2024
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu · 2024
Cited alongside, same era.
Estimating Knowledge in Large Language Models Without Generating a Single Token, October 2024
Daniela Gottesman and Mor Geva · 2024
Cited alongside, same era.
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders, August 2024
Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda · 2024
Later among the works it cites.
Measuring short-form factuality in large language models, November 2024
Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus · 2024
Later among the works it cites.
Sparse Autoencoder Features for Classifications and Transferability, February 2025
Jack Gallifant, Shan Chen, Kuleen Sasse, Hugo Aerts, Thomas Hartvigsen, and Danielle S. Bitterman · 2025
Closest in time.
Detecting Strategic Deception Using Linear Probes, February 2025
Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, and Marius Hobbhahn · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saes (usually) transfer between base and chat models
Connor Kissane, Robert Krzyzanowski, Arthur Conmy, and Neel Nanda · 2024
Cited alongside, same era.
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations, October 2024
Hadas Orgad, Michael Toker, Zorik Gekhman, Roi Reichart, Idan Szpektor, Hadas Kotek, and Yonatan Belinkov · 2024
Cited alongside, same era.
LatentQA: Teaching LLMs to Decode Activations Into Natural Language, December 2024
Alexander Pan, Lijie Chen, and Jacob Steinhardt · 2024
Cited alongside, same era.
Subhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark, and Neel Nanda · 2025
Closest in time.
Negative Results for Sparse Autoencoders On Downstream Tasks and Deprioritising SAE Research (Mechanistic Interpretability Team Progress Update), March 2025
Lewis Smith, Sen Rajamanoharan, Arthur Conmy, Callum McDougall, Janos Kramar, Tom Lieberum, Rohin Shah, and Neel Nanda · 2025
Closest in time.