Fetching the paper…
Reading the bibliography…
Despite investments in improving model safety, studies show that misaligned capabilities remain latent in safety-tuned models.
Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them
Hila Gonen and Yoav Goldberg · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
Elena Voita, Rico Sennrich, and Ivan Titov · 2019
Earlier work this paper cites.
What happens to bert embeddings during fine-tuning?
Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney · 2020
Earlier work this paper cites.
The right tool for the job: Matching model and instance complexities
Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A Smith · 2020
Earlier work this paper cites.
Similarity analysis of contextual word representation models
John Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2020
Earlier work this paper cites.
Process for adapting language models to society (palms) with values-targeted datasets
Irene Solaiman and Christy Dennison · 2021
Earlier work this paper cites.
Language models as agent models
Jacob Andreas · 2022
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Confident adaptive language modeling
Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Tran, Yi Tay, and Donald Metzler · 2022
Earlier work this paper cites.
Extracting latent steering vectors from pretrained language models
Nishant Subramani, Nivedita Suresh, and Matthew E Peters · 2022
Earlier work this paper cites.
Problems with cosine as a measure of embedding similarity for high frequency words
Kaitlyn Zhou, Kawin Ethayarajh, Dallas Card, and Dan Jurafsky · 2022
Earlier work this paper cites.
Using large language models to simulate multiple humans and replicate human subject studies
Gati Aher, Rosa I. Arriaga, and Adam T. Kalai · 2023
Earlier work this paper cites.
The reversal curse: LLMs trained on “A is B” fail to learn “B is A”
Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans · 2023
Earlier work this paper cites.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei Koh, Daphne Ippolito, Florian Tramèr, and Ludwig Schmidt · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong · 2023
Cited alongside, same era.
Marked personas: Using natural language prompts to measure stereotypes in language models
Myra Cheng, Esin Durmus, and Dan Jurafsky · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Cited alongside, same era.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan · 2023
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau · 2023
Later among the works it cites.
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2023
Later among the works it cites.
The system model and the user model: Exploring AI dashboard design
Fernanda Viégas and Martin Wattenberg · 2023
Later among the works it cites.
Jailbroken: How does LLM safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Later among the works it cites.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Jump to conclusions: Short-cutting transformers with linear transformations
Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva · 2023
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson · 2023
Cited alongside, same era.
Bias runs deep: Implicit reasoning biases in persona-assigned LLMs
Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot · 2023
Cited alongside, same era.
In-context learning creates task vectors
Roee Hendel, Mor Geva, and Amir Globerson · 2023
Cited alongside, same era.
Personas as a way to model truthfulness in language models
Nitish Joshi, Javier Rando, Abulhair Saparov, Najoung Kim, and He He · 2023
Cited alongside, same era.
Out of one, many: Using language models to simulate human samples
Nancy Fulda Joshua R. Gubler Christopher Rytting Lisa P. Argyle, Ethan C. Busby and David Wingate · 2023
Cited alongside, same era.
Sheng Liu, Lei Xing, and James Zou · 2023
Cited alongside, same era.
Closest in time.
Refusal mechanisms: initial experiments with llama-2-7b-chat
Andy Arditi and Oscar Obeso · 2024
Closest in time.
Refusal in LLMs is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Wes Gurnee, and Neel Nanda · 2024
Closest in time.
Selfie: Self-interpretation of large language model embeddings
Haozhe Chen, Carl Vondrick, and Chengzhi Mao · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva · 2024
Closest in time.
Linearity of relation decoding in transformer language models
Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K Kummerfeld, and Rada Mihalcea · 2024
Closest in time.
Mechanistically eliciting latent behaviors in language models
Andrew Mack and Alex Turner · 2024
Closest in time.
Is cosine-similarity of embeddings really about similarity?
Harald Steck, Chaitanya Ekanadham, and Nathan Kallus · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Closest in time.