Fetching the paper…
Reading the bibliography…
Despite the recent successes of large, pretrained neural language models (LLMs), comparatively little is known about the representations of linguistic structure they learn during pretraining, which can lead to unexpected behaviors in response to prompt variation or distribution shift.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“Aspects of the Theory of Syntax”
Noam Chomsky · 1965
Earlier work this paper cites.
“Semantics: Volume 2”
John Lyons · 1977
Earlier work this paper cites.
“Lexical Competence”, A Bradford book
D. Marconi · 1997
Earlier work this paper cites.
“The Prague School and North American functionalist approaches to syntax”
Frederick Newmeyer · 2001
Earlier work this paper cites.
“Causal inference in statistics: An overview”
Judea Pearl · 2009
Earlier work this paper cites.
“Performance-Compatible Competence Grammar”
Ivan Sag and Thomas Wasow · 2011
Earlier work this paper cites.
“Convex optimization: Algorithms and complexity”
Sébastien Bubeck · 2015
Earlier work this paper cites.
“Explaining and Harnessing Adversarial Examples”
Ian Goodfellow, Jonathon Shlens and Christian Szegedy · 2015
Earlier work this paper cites.
“Autoencoding beyond pixels using a learned similarity metric”
Anders Larsen, Søren Sønderby, Hugo Larochelle and Ole Winther · 2016
Earlier work this paper cites.
“Causal inference by using invariant prediction: identification and confidence intervals”
Jonas Peters, Peter Bühlmann and Nicolai Meinshausen · 2016
Earlier work this paper cites.
“Towards deep learning models resistant to adversarial attacks”
Aleksander Madry et al · 2017
Earlier work this paper cites.
“An overview of multi-task learning in deep neural networks”
Sebastian Ruder · 2017
Earlier work this paper cites.
“ConceptNet 5.5: An Open Multilingual Graph of General Knowledge”
Robyn Speer, Joshua Chin and Catherine Havasi · 2017
Earlier work this paper cites.
“Adversarial examples for generative models”
Jernej Kos, Ian Fischer and Dawn Song · 2018
Earlier work this paper cites.
“Spine: Sparse interpretable neural embeddings”
Anant Subramanian et al · 2018
Earlier work this paper cites.
“Characterizing and learning equivalence classes of causal DAGs under interventions”
Karren Yang, Abigail Katcoff and Caroline Uhler · 2018
Earlier work this paper cites.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani and David Lopez-Paz · 2019
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Scalable verified training for provably robust image classification”
Sven Gowal et al · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach”
Yinhan Liu et al · 2019
Earlier work this paper cites.
“Language Models as Knowledge Bases?”
Fabio Petroni et al · 2019
Earlier work this paper cites.
“Huggingface’s transformers: State-of-the-art natural language processing”
Thomas Wolf et al · 2019
Earlier work this paper cites.
“Invariance, causality and robustness”
Peter Bühlmann · 2020
Earlier work this paper cites.
“What BERT Is Not: Lessons from a New Suite of Psycholinguistic Diagnostics for Language Models”
Allyson Ettinger · 2020
Earlier work this paper cites.
“Neural Natural Language Inference Models Partially Embed Theories of Lexical Entailment and Negation”
Atticus Geiger, Kyle Richardson and Christopher Potts · 2020
Earlier work this paper cites.
“Semantic competence”
Diego Marconi · 2020
Earlier work this paper cites.
“Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection”
Shauli Ravfogel et al · 2020
Earlier work this paper cites.
“On the Systematicity of Probing Contextualized Word Representations: The Case of Hypernymy in BERT”
Abhilasha Ravichander et al · 2020
Earlier work this paper cites.
“A Primer in BERTology: What We Know About How BERT Works”
Anna Rogers, Olga Kovaleva and Anna Rumshisky · 2020
Cited alongside, same era.
“The risks of invariant risk minimization”
Elan Rosenfeld, Pradeep Ravikumar and Andrej Risteski · 2020
Cited alongside, same era.
“Exploring the Linear Subspace Hypothesis in Gender Bias Mitigation”
Francisco Vargas and Ryan Cotterell · 2020
Cited alongside, same era.
“Foundations of structural causal models with cycles and latent variables”
Stephan Bongers, Patrick Forré, Jonas Peters and Joris Mooij · 2021
Cited alongside, same era.
“Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals”
Yanai Elazar, Shauli Ravfogel, Alon Jacovi and Yoav Goldberg · 2021
Cited alongside, same era.
“Measuring and Improving Consistency in Pretrained Language Models”
“Fundamental limits and tradeoffs in invariant representation learning”
Han Zhao et al · 2022
Later among the works it cites.
“Language models can explain neurons in language models”
Steven Bills et al · 2023
Closest in time.
“Towards Monosemanticity: Decomposing Language Models With Dictionary Learning” https://transformer-circuits.pub/2023/monosemantic-features/index.html
Trenton Bricken et al · 2023
Closest in time.
“Towards Automated Circuit Discovery for Mechanistic Interpretability”
Arthur Conmy et al · 2023
Closest in time.
“Sparse autoencoders find highly interpretable features in language models”
Hoagy Cunningham et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yanai Elazar et al · 2021
Cited alongside, same era.
“A mathematical framework for transformer circuits”
Nelson Elhage et al · 2021
Cited alongside, same era.
“Analyzing BERT’s Knowledge of Hypernymy via Prompting”
Michael Hanna and David Mareček · 2021
Cited alongside, same era.
“Selecting data augmentation for simulating interventions”
Maximilian Ilse, Jakub Tomczak and Patrick Forré · 2021
Cited alongside, same era.
“Probing Across Time: What Does RoBERTa Know and When?”
Zeyu Liu et al · 2021
Cited alongside, same era.
“Evaluating the robustness of neural language models to input perturbations”
Milad Moradi and Matthias Samwald · 2021
Cited alongside, same era.
“AI and the everything in the whole wide world benchmark”
Inioluwa Raji et al · 2021
Cited alongside, same era.
Atticus Geiger, Chris Potts and Thomas Icard · 2023
Closest in time.
“AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap”
Q Liao and Jennifer Vaughan · 2023
Closest in time.
“Pre-Train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing”
Pengfei Liu et al · 2023
Closest in time.
“Dissociating language and thought in large language models: a cognitive perspective”
Kyle Mahowald et al · 2023
Closest in time.
“Symbols and grounding in large language models”
Ellie Pavlick · 2023
Closest in time.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron et al · 2023
Closest in time.
“Llama: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Closest in time.
“On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective”
Jindong Wang et al · 2023
Closest in time.
“Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small”
Kevin Wang et al · 2023
Closest in time.
“Interpretability at Scale: Identifying Causal Mechanisms in Alpaca”
Zhengxuan Wu, Atticus Geiger, Christopher Potts and Noah Goodman · 2023
Closest in time.
“GLUE-X: Evaluating Natural Language Understanding Models from an Out-of-Distribution Generalization Perspective”
Linyi Yang et al · 2023
Closest in time.
“The clock and the pizza: Two stories in mechanistic explanation of neural networks”
Ziqian Zhong, Ziming Liu, Max Tegmark and Jacob Andreas · 2023
Closest in time.
“Representation engineering: A top-down approach to ai transparency”
Andy Zou et al · 2023
Closest in time.
“Foundational challenges in assuring alignment and safety of large language models”
Usman Anwar et al · 2024
Closest in time.
“Leace: Perfect linear concept erasure in closed form”
Nora Belrose et al · 2024
Closest in time.
“Mechanistic Interpretability for AI Safety–A Review”
Leonard Bereska and Efstratios Gavves · 2024
Closest in time.
Marc Canby, Adam Davies, Chirag Rastogi and Julia Hockenmaier · 2024
Closest in time.
Adam Davies and Ashkan Khakzar · 2024
Closest in time.
Abhimanyu Dubey et al · 2024
Closest in time.
“OLMo: Accelerating the Science of Language Models”
Dirk Groeneveld et al · 2024
Closest in time.
“State of what art? a call for multi-prompt llm evaluation”
Moran Mizrahi et al · 2024
Closest in time.
“Automatically Interpreting Millions of Features in Large Language Models”
Gonçalo Paulo, Alex Mallen, Caden Juang and Nora Belrose · 2024
Closest in time.
“Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks”
Charlotte Siska, Katerina Marazopoulou, Melissa Ailem and James Bono · 2024
Closest in time.
“Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet”
Adly Templeton et al · 2024
Closest in time.