Fetching the paper…
Reading the bibliography…
Fine-tuning large language models (LLMs) can lead to unintended out-of-distribution generalization.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A Olshausen and David J Field · 1997
Earlier work this paper cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals, 2021
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2006
Earlier work this paper cites.
Efficient sparse coding algorithms
Honglak Lee, Alexis Battle, Rajat Raina, and Andrew Ng · 2006
Earlier work this paper cites.
Coping with label shift via distributionally robust optimisation, 2021
Jingzhao Zhang, Aditya Menon, Andreas Veit, Srinadh Bhojanapalli, Sanjiv Kumar, and Suvrit Sra · 2010
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories, 2021
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality, 2013
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Controllable invariance through adversarial feature learning
Qizhe Xie, Zihang Dai, Yulun Du, Eduard Hovy, and Graham Neubig · 2017
Earlier work this paper cites.
Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study
John R Zech, Marcus A Badgeley, Manway Liu, Anthony B Costa, Joseph J Titano, and Eric Karl Oermann · 2018
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
Learning not to learn: Training deep neural networks with biased data, 2019
Byungju Kim, Hyunwoo Kim, Kyungsu Kim, Sungjin Kim, and Junmo Kim · 2019
Earlier work this paper cites.
Distributionally robust language modeling
Yonatan Oren, Shiori Sagawa, Tatsunori B. Hashimoto, and Percy Liang · 2019
Earlier work this paper cites.
Understanding black-box predictions via influence functions, 2020
Pang Wei Koh and Percy Liang · 2020
Earlier work this paper cites.
Learning from failure: De-biasing classifier from biased classifier
Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin · 2020
Earlier work this paper cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Earlier work this paper cites.
Distributionally robust neural networks
Shiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, and Percy Liang · 2020
Earlier work this paper cites.
Towards debiasing NLU models from unknown biases
Prasetya Ajie Utama, Nafise Sadat Moosavi, and Iryna Gurevych · 2020
Earlier work this paper cites.
Double-hard debias: Tailoring word embeddings for gender bias mitigation
Tianlu Wang, Xi Victoria Lin, Nazneen Fatema Rajani, Bryan McCann, Vicente Ordonez, and Caiming Xiong · 2020
Earlier work this paper cites.
Editing factual knowledge in language models, 2021
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning, 2021
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Just train twice: Improving group robustness without training group information
Evan Z Liu, Behzad Haghgoo, Annie S Chen, Aditi Raghunathan, Pang Wei Koh, Shiori Sagawa, Percy Liang, and Chelsea Finn · 2021
Earlier work this paper cites.
Increasing robustness to spurious correlations using forgettable examples
Yadollah Yaghoobzadeh, Soroush Mehri, Remi Tachet des Combes, T. J. Hazen, and Alessandro Sordoni · 2021
Earlier work this paper cites.
Masktune: Mitigating spurious correlations by forcing to explore
Saeid Asgari, Aliasghar Khani, Fereshte Khani, Ali Gholami, Linh Tran, Ali Mahdavi Amiri, and Ghassan Hamarneh · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models, 2022
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Knowledge neurons in pretrained transformers, 2022
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Earlier work this paper cites.
Toy models of superposition, 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah · 2022
Earlier work this paper cites.
Ensembles and cocktails: Robust finetuning for natural language generation, 2022
John Hewitt, Xiang Lisa Li, Sang Michael Xie, Benjamin Newman, and Percy Liang · 2022
Earlier work this paper cites.
Datamodels: Predicting predictions from training data, 2022
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Cited alongside, same era.
Fine-tuning can distort pretrained features and underperform out-of-distribution, 2022
Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Fast model editing at scale, 2022
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning · 2022
Cited alongside, same era.
Spread spurious attribute: Improving worst-group accuracy with spurious attribute estimation, 2022
Batchtopk: A simple improvement for topk-saes, 2024
Bart Bussmann, Patrick Leask, and Neel Nanda · 2024
Later among the works it cites.
Sycophancy to subterfuge: Investigating reward-tampering in large language models, 2024
Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, and Evan Hubinger · 2024
Later among the works it cites.
Applying sparse autoencoders to unlearn knowledge in language models, 2024
Eoin Farrell, Yeu-Tong Lau, and Arthur Conmy · 2024
Later among the works it cites.
Mechanistic unlearning: Robust knowledge unlearning and editing via mechanistic localization, 2024
Phillip Guo, Aaquib Syed, Abhay Sheshadri, Aidan Ewart, and Gintare Karolina Dziugaite · 2024
Later among the works it cites.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junhyun Nam, Jaehyung Kim, Jaeho Lee, and Jinwoo Shin · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations, 2022
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Cited alongside, same era.
Adversarial concept erasure in kernel space
Shauli Ravfogel, Francisco Vargas, Yoav Goldberg, and Ryan Cotterell · 2022
Cited alongside, same era.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals, 2022
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Cited alongside, same era.
Barack: Partially supervised group robustness with guarantees, 2022
Nimit S. Sohoni, Maziar Sanjabi, Nicolas Ballas, Aditya Grover, Shaoliang Nie, Hamed Firooz, and Christopher Ré · 2022
Cited alongside, same era.
Leace: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning, 2023
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nicholas L Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Chris Olah · 2023
Cited alongside, same era.
Generalization analogies: A testbed for generalizing ai oversight to hard-to-measure domains, 2023
Joshua Clymer, Garrett Baker, Rohan Subramani, and Sam Wang · 2023
Cited alongside, same era.
Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka, Edward Grefenstette, Tim Rocktäschel, and David Scott Krueger · 2024
Later among the works it cites.
Open source automated interpretability for sparse autoencoder features
Caden Juang, Gonçalo Paulo, Jacob Drori, and Nora Belrose · 2024
Later among the works it cites.
On scalable oversight with weak llms judging strong llms, 2024
Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, and Rohin Shah · 2024
Later among the works it cites.
Clarify: Improving model robustness with natural language corrections
Yoonho Lee, Michelle S Lam, Helena Vasconcelos, Michael S Bernstein, and Chelsea Finn · 2024
Later among the works it cites.
Beyond accuracy: Ensuring correct predictions with correct rationales
Tang Li, Mengmeng Ma, and Xi Peng · 2024
Later among the works it cites.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda · 2024
Later among the works it cites.
Samuel Marks and Max Tegmark · 2024
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models, 2024
Kiho Park, Yo Joong Choe, and Victor Veitch · 2024
Later among the works it cites.
Automatically interpreting millions of features in large language models, 2024
Gonçalo Paulo, Alex Mallen, Caden Juang, and Nora Belrose · 2024
Later among the works it cites.
Fine-tuning enhances existing mechanisms: A case study on entity tracking, 2024
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau · 2024
Later among the works it cites.
Guardrail baselines for unlearning in llms, 2024
Pratiksha Thaker, Yash Maurya, Shengyuan Hu, Zhiwei Steven Wu, and Virginia Smith · 2024
Later among the works it cites.
Function vectors in large language models, 2024
Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau · 2024
Later among the works it cites.
Lmsys-chat-1m: A large-scale real-world llm conversation dataset, 2024
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P. Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang · 2024
Later among the works it cites.
Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025
Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, and Owain Evans · 2025
Closest in time.
Are sparse autoencoders useful? a case study in sparse probing, 2025
Subhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark, and Neel Nanda · 2025
Closest in time.
Saebench: A comprehensive benchmark for sparse autoencoders in language model interpretability, 2025
Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Matthew Wearden, Arthur Conmy, Samuel Marks, and Neel Nanda · 2025
Closest in time.
Aly M. Kassem, Zhuan Shi, Negar Rostamzadeh, and Golnoosh Farnadi · 2025
Closest in time.
Analyzing (in)abilities of saes via formal languages, 2025
Abhinav Menon, Manish Shrivastava, David Krueger, and Ekdeep Singh Lubana · 2025
Closest in time.
Sparse autoencoders for hypothesis generation, 2025
Rajiv Movva, Kenny Peng, Nikhil Garg, Jon Kleinberg, and Emma Pierson · 2025
Closest in time.
How new data permeates llm knowledge and how to dilute it, 2025
Chen Sun, Renat Aksitov, Andrey Zhmoginov, Nolan Andrew Miller, Max Vladymyrov, Ulrich Rueckert, Been Kim, and Mark Sandler · 2025
Closest in time.
Model organisms for emergent misalignment, 2025
Edward Turner, Anna Soligo, Mia Taylor, Senthooran Rajamanoharan, and Neel Nanda · 2025
Closest in time.
Persona features control emergent misalignment, 2025
Miles Wang, Tom Dupré la Tour, Olivia Watkins, Alex Makelov, Ryan A. Chi, Samuel Miserendino, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing · 2025
Closest in time.
Axbench: Steering llms? even simple baselines outperform sparse autoencoders, 2025
Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang, Jing Huang, Dan Jurafsky, Christopher D. Manning, and Christopher Potts · 2025
Closest in time.