Fetching the paper…
Reading the bibliography…
Representation Engineering (RepE) is a novel paradigm for controlling the behavior of LLMs.
Estimating causal effects of treatments in randomized and nonrandomized studies
Donald B Rubin · 1975
Earlier work this paper cites.
The superintelligent will: Motivation and instrumental rationality in advanced artificial agents
Nick Bostrom · 2012
Earlier work this paper cites.
Probabilistic reasoning in intelligent systems: networks of plausible inference
Judea Pearl · 2014
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang · 2015
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks, 2016
Alec Radford, Luke Metz, and Soumith Chintala · 2016
Earlier work this paper cites.
Deep feature interpolation for image content changes
Paul Upchurch, Jacob Gardner, Geoff Pleiss, Robert Pless, Noah Snavely, Kavita Bala, and Kilian Weinberger · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, 2018
Guillaume Alain and Yoshua Bengio · 2018
Earlier work this paper cites.
Gabriel Grand, Idan Asher Blank, Francisco Pereira, and Evelina Fedorenko · 2018
Earlier work this paper cites.
Ares: a framework for quantifying the resilience of deep neural networks
Brandon Reagen, Udit Gupta, Lillian Pentecost, Paul Whatmough, Sae Kyu Lee, Niamh Mulholland, David Brooks, and Gu-Yeon Wei · 2018
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
Editing in style: Uncovering the local semantics of gans
Edo Collins, Raja Bala, Bob Price, and Sabine Susstrunk · 2020
Earlier work this paper cites.
GoEmotions: A dataset of fine-grained emotions
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
vaderSentiment
C. J. Hutto · 2020
Earlier work this paper cites.
On the "steerability" of generative adversarial networks
Ali Jahanian, Lucy Chai, and Phillip Isola · 2020
Earlier work this paper cites.
Interpreting the latent space of gans for semantic face editing
Yujun Shen, Jinjin Gu, Xiaoou Tang, and Bolei Zhou · 2020
Earlier work this paper cites.
Unsupervised discovery of interpretable directions in the GAN latent space
Andrey Voynov and Artem Babenko · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush · 2020
Earlier work this paper cites.
Can language models encode perceptual structure without grounding? a case study in color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard · 2021
Earlier work this paper cites.
Hatebert: Retraining bert for abusive language detection in english
Tommaso Caselli, Valerio Basile, Jelena Mitrović, and Michael Granitzer · 2021
Earlier work this paper cites.
Bold: Dataset and metrics for measuring biases in open-ended language generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Gan-control: Explicitly controllable gans
Alon Shoshan, Nadav Bhonker, Igor Kviatkovsky, and Gerard Medioni · 2021
Earlier work this paper cites.
Hijack-gan: Unintended-use of pretrained, black-box gans
Hui-Po Wang, Ning Yu, and Mario Fritz · 2021
Earlier work this paper cites.
Stylespace analysis: Disentangled controls for stylegan image generation
Zongze Wu, Dani Lischinski, and Eli Shechtman · 2021
Earlier work this paper cites.
Toy models of superposition, September 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Earlier work this paper cites.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman · 2022
Earlier work this paper cites.
Extracting latent steering vectors from pretrained language models
Nishant Subramani, Nivedita Suresh, and Matthew Peters · 2022
Earlier work this paper cites.
Red-teaming large language models using chain of utterances for safety-alignment, 2023
Rishabh Bhardwaj and Soujanya Poria · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning, 2023
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Chris Olah · 2023
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2023
Earlier work this paper cites.
Detecting text formality: A study of text classification approaches
Daryna Dementieva, Nikolay Babakov, and Alexander Panchenko · 2023
Earlier work this paper cites.
Verb conjugation in transformers is determined by linear encodings of subject number
Sophie Hao and Tal Linzen · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa · 2023
Earlier work this paper cites.
Self-detoxifying language models via toxification reversal
Chak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang, and Wenjie Li · 2023
Earlier work this paper cites.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis · 2023
Earlier work this paper cites.
Aligning large language models with human preferences through representation engineering
Wenhao Liu, Xiaohua Wang, Muling Wu, Tianlong Li, Changze Lv, Zixuan Ling, Jianhao Zhu, Cenyuan Zhang, Xiaoqing Zheng, and Xuanjing Huang · 2023
Earlier work this paper cites.
The geometry of truth: Emergent linear structure in large language model representations of true/false datasets
Samuel Marks and Max Tegmark · 2023
Earlier work this paper cites.
Understanding and controlling a maze-solving policy network, 2023
Ulisse Mini, Peli Grietzer, Mrinank Sharma, Austin Meek, Monte MacDiarmid, and Alexander Matt Turner · 2023
Earlier work this paper cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Earlier work this paper cites.
Interventional probing in high dimensions: An NLI case study
Julia Rozanova, Marco Valentino, Lucas Cordeiro, and André Freitas · 2023
Earlier work this paper cites.
Haoran Wang and Kai Shu · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica · 2023
Earlier work this paper cites.
Representation tuning
Christopher Ackerman · 2024
Earlier work this paper cites.
Can language models safeguard themselves, instantly and for free?
Dyah Adila, Changho Shin, Yijing Zhang, and Frederic Sala · 2024
Cited alongside, same era.
Refusal in language models is mediated by a single direction, 2024
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda · 2024
Cited alongside, same era.
Causalgym: Benchmarking causal interpretability methods on linguistic tasks
Aryaman Arora, Dan Jurafsky, and Christopher Potts · 2024
Cited alongside, same era.
Intervention lens: from representation surgery to string counterfactuals, 2024
Matan Avitan, Ryan Cotterell, Yoav Goldberg, and Shauli Ravfogel · 2024
Cited alongside, same era.
Understanding jailbreak success: A study of latent space dynamics in large language models, 2024
Future events as backdoor triggers: Investigating temporal vulnerabilities in llms
Sara Price, Arjun Panickssery, Sam Bowman, and Asa Cooper Stickland · 2024
Later among the works it cites.
Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu, Zhenfei Yin, Yu Qiao, Yong Liu, and Jing Shao · 2024
Later among the works it cites.
Spectral editing of activations for large language model alignment, 2024
Yifu Qiu, Zhao, Yftah Ziser, Anna Korhonen, Edoardo M. Ponti, and Shay B. Cohen · 2024
Later among the works it cites.
Controlling large language model agents with entropic activation steering
Nate Rahn, Pierluca D’Oro, and Marc G Bellemare · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sarah Ball, Frauke Kreuter, and Nina Panickssery · 2024
Cited alongside, same era.
Safeinfer: Context adaptive decoding time safety alignment for large language models, 2024
Somnath Banerjee, Soham Tripathy, Sayan Layek, Shanu Kumar, Animesh Mukherjee, and Rima Hazra · 2024
Cited alongside, same era.
Mechanistic interpretability for AI safety - a review
Leonard Bereska and Stratis Gavves · 2024
Cited alongside, same era.
Stylitgan: Image-based relighting via latent control
Anand Bhattad, James Soole, and D.A. Forsyth · 2024
Cited alongside, same era.
Benchmarking mental state representations in language models
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, and Andreas Bulling · 2024
Cited alongside, same era.
Comparing bottom-up and top-down steering approaches on in-context learning tasks
Madeline Brumley, Joe Kwon, David Krueger, Dmitrii Krasheninnikov, and Usman Anwar · 2024
Cited alongside, same era.
Self-control of LLM behaviors by compressing suffix gradient into prefix controller
Min Cai, Yuchen Zhang, Shichang Zhang, Fan Yin, Difan Zou, Yisong Yue, and Ziniu Hu · 2024
Cited alongside, same era.
Improving steering vectors by targeting sparse autoencoder features, 2024
Sviatoslav Chalnev, Matthew Siu, and Arthur Conmy · 2024
Cited alongside, same era.
Goutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf, and Pradeep Ravikumar · 2024
Later among the works it cites.
Representation noising: A defence mechanism against harmful finetuning
Domenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze, Robie Gonzales, carsten maple, Subhabrata Majumdar, Hassan Sajjad, and Frank Rudzicz · 2024
Later among the works it cites.
Multi-property steering of large language models with dynamic activation composition
Daniel Scalena, Gabriele Sarti, and Malvina Nissim · 2024
Later among the works it cites.
Constructing benchmarks and interventions for combating hallucinations in llms
Adi Simhi, Jonathan Herzig, Idan Szpektor, and Yonatan Belinkov · 2024
Later among the works it cites.
Representation surgery: Theory and practice of affine steering
Shashwat Singh, Shauli Ravfogel, Jonathan Herzig, Roee Aharoni, Ryan Cotterell, and Ponnurangam Kumaraguru · 2024
Later among the works it cites.
Steering without side effects: Improving post-deployment control of language models, 2024
Asa Cooper Stickland, Alexander Lyzhov, Jacob Pfau, Salsabila Mahdi, and Samuel R. Bowman · 2024
Later among the works it cites.
Analyzing the generalization and reliability of steering vectors
Daniel Chee Hian Tan, David Chanin, Aengus Lynch, Adrià Garriga-Alonso, Dimitrios Kanoulas, Brooks Paige, and Robert Kirk · 2024
Later among the works it cites.
On the difficulty of faithful chain-of-thought reasoning in large language models
Sree Harsha Tanneru, Dan Ley, Chirag Agarwal, and Himabindu Lakkaraju · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Later among the works it cites.
Exploring and steering the moral compass of large language models
Alejandro Tlaie · 2024
Later among the works it cites.
Function vectors in large language models
Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau · 2024
Later among the works it cites.
Initial response selection for prompt jailbreaking using model steering
Thien Q. Tran, Koki Wataoka, and Tsubasa Takahashi · 2024
Later among the works it cites.
Steering language models with activation engineering, 2024
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid · 2024
Later among the works it cites.
Extending activation steering to broad skills and multiple behaviours
Teun van der Weij, Massimo Poesio, and Nandi Schoots · 2024
Later among the works it cites.
A language model’s guide through latent space
Dimitri von Rütte, Sotiris Anagnostidis, Gregor Bachmann, and Thomas Hofmann · 2024
Later among the works it cites.
Does editing provide evidence for localization?
Zihao Wang and Victor Veitch · 2024
Later among the works it cites.
Controllm: Crafting diverse personalities for language models
Yixuan Weng, Shizhu He, Kang Liu, Shengping Liu, and Jun Zhao · 2024
Later among the works it cites.
Tradeoffs between alignment and helpfulness in language models with representation engineering, 2024
Yotam Wolf, Noam Wies, Dorin Shteyman, Binyamin Rothberg, Yoav Levine, and Amnon Shashua · 2024
Later among the works it cites.
Mitigating privacy seesaw in large language models: Augmented privacy neuron editing via activation patching
Xinwei Wu, Weilong Dong, Shaoyang Xu, and Deyi Xiong · 2024
Later among the works it cites.
Enhancing multiple dimensions of trustworthiness in llms via sparse activation control
Yuxin Xiao, Wan Chaoqun, Yonggang Zhang, Wenxiao Wang, Binbin Lin, Xiaofei He, Xu Shen, and Jieping Ye · 2024
Later among the works it cites.
Enhancing semantic consistency of large language models through model editing: An interpretability-oriented approach
Jingyuan Yang, Dapeng Chen, Yajing Sun, Rongjun Li, Zhiyong Feng, and Wei Peng · 2024
Later among the works it cites.
Jailbreak attacks and defenses against large language models: A survey, 2024
Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong, Xinlei He, Jiaxing Song, Ke Xu, and Qi Li · 2024
Later among the works it cites.
Lofit: Localized fine-tuning on llm representations
Fangcong Yin, Xi Ye, and Greg Durrett · 2024
Later among the works it cites.
Robust llm safeguarding via refusal feature adversarial training, 2024
Lei Yu, Virginie Do, Karen Hambardzumyan, and Nicola Cancedda · 2024
Later among the works it cites.
Ziqian Zeng, Jianwei Wang, Junyao Yang, Zhengdong Lu, Huiping Zhuang, and Cen Chen · 2024
Later among the works it cites.
Uncovering latent chain of thought vectors in language models, 2024
Jason Zhang and Scott Viteri · 2024
Later among the works it cites.
Steering knowledge selection behaviours in llms via sae-based representation engineering, 2024
Yu Zhao, Alessio Devoto, Giwon Hong, Xiaotang Du, Aryo Pradipta Gema, Hongru Wang, Xuanli He, Kam-Fai Wong, and Pasquale Minervini · 2024
Later among the works it cites.
Language models represent beliefs of self and others
Wentao Zhu, Zhining Zhang, and Yizhou Wang · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, J Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Later among the works it cites.
Daniel Beaglehole, Adityanarayanan Radhakrishnan, Enric Boix-Adserà, and Mikhail Belkin · 2025
Closest in time.
Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs
Jan Betley, Daniel Chee Hian Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, and Owain Evans · 2025
Closest in time.
Understanding (un)reliability of steering vectors in language models
Joschka Braun, Carsten Eickhoff, David Krueger, Seyed Ali Bahrainian, and Dmitrii Krasheninnikov · 2025
Closest in time.
Rethinking the reliability of representation engineering in large language models, 2025
Zhihong Deng, Jing Jiang, Guodong Long, and Chengqi Zhang · 2025
Closest in time.
Not all language model features are linear
Joshua Engels, Eric J Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark · 2025
Closest in time.
I found >800 orthogonal "write code" steering vectors, July 2024
Jacob Goldman-Wetzler and Alexander Matt Turner · 2025
Closest in time.
Detecting strategic deception using linear probes, 2025
Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, and Marius Hobbhahn · 2025
Closest in time.
Improving reasoning performance in large language models via representation engineering
Bertram Højer, Oliver Simon Jarvis, and Stefan Heinrich · 2025
Closest in time.
A unified understanding and evaluation of steering methods, 2025
Shawn Im and Yixuan Li · 2025
Closest in time.
DRESSing up LLM: Efficient stylized question-answering via style subspace editing
Xinyu Ma, Yifeng Xu, Yang Lin, Tianlong Wang, Xu Chu, Xin Gao, Junfeng Zhao, and Yasha Wang · 2025
Closest in time.
Simple probes can catch sleeper agents
Monte MacDiarmid, Timothy Maxwell, Nicholas Schiefer, Jesse Mu, Jared Kaplan, David Duvenaud, Sam Bowman, Alex Tamkin, Ethan Perez, Mrinank Sharma, Carson Denison, and Evan Hubinger · 2025
Closest in time.
Evaluating sparse autoencoders for controlling open-ended text generation
Aleksandar Makelo, Nathaniel Monson, and Julius Adebayo · 2025
Closest in time.
Exploring Gemma Scope, 2025
Neuronpedia · 2025
Closest in time.
Multi-attribute steering of language models via targeted intervention, 2025
Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, and Mohit Bansal · 2025
Closest in time.
Convergent linear representations of emergent misalignment, 2025
Anna Soligo, Edward Turner, Senthooran Rajamanoharan, and Neel Nanda · 2025
Closest in time.
Axbench: Steering llms? even simple baselines outperform sparse autoencoders, 2025
Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D. Manning, and Christopher Potts · 2025
Closest in time.
Data-centric artificial intelligence: A survey
Daochen Zha, Zaid Pervaiz Bhat, Kwei-Herng Lai, Fan Yang, Zhimeng Jiang, Shaochen Zhong, and Xia Hu · 2025
Closest in time.