Fetching the paper…
Reading the bibliography…
Recent work in activation steering has demonstrated the potential to better control the outputs of Large Language Models (LLMs), but it involves finding steering vectors.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2020 · 1910
Earlier work this paper cites.
Linguistic Regularities in Continuous Space Word Representations
Mikolov, T.; Yih, W.-t.; and Zweig, G. 2013 · 2013
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Pennington, J.; Socher, R.; and Manning, C. 2014 · 2014
Earlier work this paper cites.
Toxic Comment Classification Challenge
Adams, C.; Sorensen, J.; Elliott, J.; Dixon, L.; McDonald, M.; nithum; and Cukierski, W. 2017 · 2017
Earlier work this paper cites.
All-but-the-Top: Simple and Effective Postprocessing for Word Representations
Mu, J.; and Viswanath, P. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Earlier work this paper cites.
Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
Borkan, D.; Dixon, L.; Sorensen, J.; Thain, N.; and Vasserman, L. 2019 · 2019
Earlier work this paper cites.
How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings
Ethayarajh, K. 2019 · 2019
Earlier work this paper cites.
OpenWebText Corpus
Gokaslan, A.; and Cohen, V. 2019 · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Earlier work this paper cites.
Interpreting GPT: The Logit Lens
nostalgebrist. 2020 · 2020
Cited alongside, same era.
Persistent Anti-Muslim Bias in Large Language Models
Abid, A.; Farooqi, M.; and Zou, J. 2021 · 2021
Cited alongside, same era.
Isotropy in the Contextual Embedding Space: Clusters and Manifolds
Cai, X.; Huang, J.; Bian, Y.; and Church, K. 2021 · 2021
Cited alongside, same era.
SkoltechNLP at SemEval-2021 Task 5: Leveraging Sentence-level Pre-training for Toxic Span Detection
Dale, D.; Markov, I.; Logacheva, V.; Kozlova, O.; Semenov, N.; and Panchenko, A. 2021 · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Black, S.; Biderman, S.; Hallahan, E.; Anthony, Q.; Gao, L.; Golding, L.; He, H.; Leahy, C.; McDonell, K.; Phang, J.; Pieler, M.; Prashanth, U. S.; Purohit, S.; Reynolds, L.; Tow, J.; Wang, B.; and Weinbach, S. 2022 · 2022
Sparse Autoencoders Find Highly Interpretable Features in Language Models
Cunningham, H.; Ewart, A.; Riggs, L.; Huben, R.; and Sharkey, L. 2023 · 2023
Closest in time.
Language Models Represent Space and Time
Gurnee, W.; and Tegmark, M. 2023 · 2023
Closest in time.
Editing models with task arithmetic
Ilharco, G.; Ribeiro, M. T.; Wortsman, M.; Schmidt, L.; Hajishirzi, H.; and Farhadi, A. 2023 · 2023
Closest in time.
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
Li, K.; Patel, O.; Viégas, F.; Pfister, H.; and Wattenberg, M. 2023 · 2023
Closest in time.
Emergent Linear Representations in World Models of Self-Supervised Sequence Models
Nanda, N.; Lee, A.; and Wattenberg, M. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Toy Models of Superposition
Elhage, N.; Hume, T.; Olsson, C.; Schiefer, N.; Henighan, T.; Kravec, S.; Hatfield-Dodds, Z.; Lasenby, R.; Drain, D.; Chen, C.; Grosse, R.; McCandlish, S.; Kaplan, J.; Amodei, D.; Wattenberg, M.; and Olah, C. 2022 · 2022
Cited alongside, same era.
Locating and Editing Factual Associations in GPT
Meng, K.; Bau, D.; Andonian, A. J.; and Belinkov, Y. 2022 · 2022
Cited alongside, same era.
Extracting Latent Steering Vectors from Pretrained Language Models
Subramani, N.; Suresh, N.; and Peters, M. 2022 · 2022
Cited alongside, same era.
LEACE: Perfect linear concept erasure in closed form
Belrose, N.; Schneider-Joseph, D.; Ravfogel, S.; Cotterell, R.; Raff, E.; and Biderman, S. 2023 · 2023
Cited alongside, same era.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Bricken, T.; Templeton, A.; Batson, J.; Chen, B.; Jermyn, A.; Conerly, T.; Turner, N.; Anil, C.; Denison, C.; Askell, A.; Lasenby, R.; Wu, Y.; Kravec, S.; Schiefer, N.; Maxwell, T.; Joseph, N.; Hatfield-Dodds, Z.; Tamkin, A.; Nguyen, K.; McLean, B.; Burke, J. E.; Hume, T.; Carter, S.; Henighan, T.; and Olah, C. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023 · 2023
Closest in time.
Red-teaming language models via activation engineering
Rimsky, N. 2023 · 2023
Closest in time.
Function Vectors in Large Language Models
Todd, E.; Li, M. L.; Sharma, A. S.; Mueller, A.; Wallace, B. C.; and Bau, D. 2023 · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; Bikel, D.; Blecher, L.; Ferrer, C. C.; Chen, M.; Cucurull, G.; Esiobu, D.; Fernandes, J.; Fu, J.; Fu, W.; Fuller, B.; Gao, C.; Goswami, V.; Goyal, N.; Hartshorn, A.; Hosseini, S.; Hou, R.; Inan, H.; Kardas, M.; Kerkez, V.; Khabsa, M.; Kloumann, I.; Korenev, A.; Koura, P. S.; Lachaux, M.-A.; Lavril, T.; Lee, J.; Liskovich, D.; Lu, Y.; Mao, Y.; Martinet, X.; Mihaylov, T.; Mishra, P.; Molybog, I.; Nie, Y.; Poulton, A.; Reizenstein, J.; Rungta, R.; Saladi, K.; Schelten, A.; Silva, R.; Smith, E. M.; Subramanian, R.; Tan, X. E.; Tang, B.; Taylor, R.; Williams, A.; Kuan, J. X.; Xu, P.; Yan, Z.; Zarov, I.; Zhang, Y.; Fan, A.; Kambadur, M.; Narang, S.; Rodriguez, A.; Stojnic, R.; Edunov, S.; and Scialom, T. 2023 · 2023
Closest in time.
Activation Addition: Steering Language Models Without Optimization
Turner, A. M.; Thiergart, L.; Udell, D.; Leech, G.; Mini, U.; and MacDiarmid, M. 2023 · 2023
Closest in time.
Representation Engineering: A Top-Down Approach to AI Transparency
Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Pan, A.; Yin, X.; Mazeika, M.; Dombrowski, A.-K.; Goel, S.; Li, N.; Byun, M. J.; Wang, Z.; Mallen, A.; Basart, S.; Koyejo, S.; Song, D.; Fredrikson, M.; Kolter, J. Z.; and Hendrycks, D. 2023 · 2023
Closest in time.