Fetching the paper…
Reading the bibliography…
Concept-based interpretability methods offer a lens into the internals of foundation models by decomposing their embeddings into high-level concepts.
Silhouettes: a graphical aid to the interpretation and validation of cluster analysis
Rousseeuw, P. J · 1987
Earlier work this paper cites.
Exploiting context to identify lexical atoms–a statistical view of linguistic context
Zhai, C · 1997
Earlier work this paper cites.
Twenty Newsgroups
Mitchell, T · 1999
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D., Ng, A., and Jordan, M · 2001
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S · 2011
Earlier work this paper cites.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Towards more human-like concept learning in machines: Compositionality, causality, and learning-to-learn
Lake, B. M · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T · 2016
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., and Girshick, R · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Caron, M., Bojanowski, P., Joulin, A., and Douze, M · 2018
Earlier work this paper cites.
Learning to make analogies by contrasting abstract relational structure
Hill, F., Santoro, A., Barrett, D., Morcos, A., and Lillicrap, T · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al · 2018
Earlier work this paper cites.
The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions
Tschandl, P., Rosendahl, C., and Kittler, H · 2018
Earlier work this paper cites.
Interpretable basis decomposition for visual explanation
Zhou, B., Sun, Y., Bau, D., and Torralba, A · 2018
Earlier work this paper cites.
Measuring compositionality in representation learning
Andreas, J · 2019
Cited alongside, same era.
Towards automatic concept-based explanations
Ghorbani, A., Wexler, J., Zou, J. Y., and Kim, B · 2019
Cited alongside, same era.
Concept whitening for interpretable image recognition
Chen, Z., Bei, Y., and Rudin, C · 2020
Cited alongside, same era.
Concepts and compositionality: in search of the brain’s language of thought
Frankland, S. M. and Greene, J. D · 2020
Cited alongside, same era.
Concept bottleneck models
Koh, P. W., Nguyen, T., Tang, Y. S., Mussmann, S., Pierson, E., Kim, B., and Liang, P · 2020
Cited alongside, same era.
On completeness-aware concept-based explanations in deep neural networks
Yeh, C.-K., Kim, B., Arik, S., Li, C.-L., Pfister, T., and Ravikumar, P · 2020
Cited alongside, same era.
The internal state of an llm knows when its lying
Azaria, A. and Mitchell, T · 2023
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Bricken, T., Templeton, A., Batson, J., Chen, B., Jermyn, A., Conerly, T., Turner, N., Anil, C., Denison, C., Askell, A., Lasenby, R., Wu, Y., Kravec, S., Schiefer, N., Maxwell, T., Joseph, N., Hatfield-Dodds, Z., Tamkin, A., Nguyen, K., McLean, B., Burke, J. E., Hume, T., Carter, S., Henighan, T., and Olah, C · 2023
Later among the works it cites.
Craft: Concept recursive activation factorization for explainability
Fel, T., Picard, A., Bethune, L., Boissin, T., Vigouroux, D., Colin, J., Cadène, R., and Serre, T · 2023
Later among the works it cites.
Topex: Topic-based explanations for model comparison
Havaldar, S., Stein, A., Wong, E., and Ungar, L. H · 2023
Later among the works it cites.
Diffusion models already have a semantic latent space
Kwon, M., Jeong, J., and Uh, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Editing a classifier by rewriting its prediction rules
Santurkar, S., Tsipras, D., Elango, M., Bau, D., Torralba, A., and Madry, A · 2021
Cited alongside, same era.
Lecture notes on high-dimensional data
Wegner, S.-A · 2021
Cited alongside, same era.
Leveraging sparse linear layers for debuggable deep networks
Wong, E., Santurkar, S., and Madry, A · 2021
Cited alongside, same era.
Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors
Yun, Z., Chen, Y., Olshausen, B., and Lecun, Y · 2021
Cited alongside, same era.
Meaningfully debugging model mistakes using conceptual counterfactual explanations
Abid, A., Yuksekgonul, M., and Zou, J · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2023
Later among the works it cites.
Rectifying group irregularities in explanations for distribution shift
Stein, A., Wu, Y., Wong, E., and Naik, M · 2023
Later among the works it cites.
Evaluating and mitigating discrimination in language model decisions
Tamkin, A., Askell, A., Lovitt, L., Durmus, E., Joseph, N., Kravec, S., Nguyen, K., Kaplan, J., and Ganguli, D · 2023
Later among the works it cites.
Function vectors in large language models
Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Linear spaces of meanings: compositional structures in vision-language models
Trager, M., Perera, P., Zancato, L., Achille, A., Bhatia, P., and Soatto, S · 2023
Later among the works it cites.
Concept algebra for (score-based) text-controlled generative models
Wang, Z., Gui, L., Negrea, J., and Veitch, V · 2023
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in alpaca
Wu, Z., Geiger, A., Icard, T., Potts, C., and Goodman, N · 2023
Later among the works it cites.
Language in a bottle: Language model guided concept bottlenecks for interpretable image classification
Yang, Y., Panagopoulou, A., Zhou, S., Jin, D., Callison-Burch, C., and Yatskar, M · 2023
Later among the works it cites.
Post-hoc concept bottleneck models
Yuksekgonul, M., Wang, M., and Zou, J · 2023
Later among the works it cites.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2024
Closest in time.
Function vectors in large language models
Todd, E., Li, M., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Turpin, M., Michael, J., Perez, E., and Bowman, S · 2024
Closest in time.