Decoderlens: Layerwise interpretation of encoder-decoder transformers
Original
Langedijk, A., Mohebbi, H., Sarti, G., Zuidema, W., and Jumelet, J · 2023
Later among the works it cites.
Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla
Original
Lieberum, T., Rahtz, M., Kramár, J., Irving, G., Shah, R., and Mikulik, V · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., and Clark, P · 2023
Later among the works it cites.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Mallen, A., Asai, A., Zhong, V., Das, R., Khashabi, D., and Hajishirzi, H · 2023
Later among the works it cites.
A mechanism for solving relational tasks in transformer language models
Original
Merullo, J., Eickhoff, C., and Pavlick, E · 2023
Later among the works it cites.
Can llms facilitate interpretation of pre-trained language models?
Original
Mousi, B., Durrani, N., and Dalvi, F · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Nanda, N., Lee, A., and Wattenberg, M · 2023
Later among the works it cites.
Future lens: Anticipating subsequent tokens from a single hidden state
Pal, K., Sun, J., Yuan, A., Wallace, B. C., and Bau, D · 2023
Later among the works it cites.
Multimodal neurons in pretrained text-only transformers
Schwettmann, S., Chowdhury, N., Klein, S., Bau, D., and Torralba, A · 2023
Later among the works it cites.
Explaining black box text modules in natural language with language models
Singh, C., Hsu, A., Antonello, R., Jain, S., Huth, A., Yu, B., and Gao, J · 2023
Later among the works it cites.
The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models
Slobodkin, A., Goldman, O., Caciularu, A., Dagan, I., and Ravfogel, S · 2023
Later among the works it cites.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis
Stolfo, A., Belinkov, Y., and Sachan, M · 2023
Later among the works it cites.
Analyzing vision transformers for image classification in class embedding space
Original
Vilas, M. G., Schaumlöffel, T., and Roig, G · 2023
Later among the works it cites.
Gaussian Process Probes (GPP) for uncertainty-aware probing
Original
Wang, Z., Ku, A., Baldridge, J., Griffiths, T. L., and Kim, B · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. R · 2023
Later among the works it cites.
Give me the facts! a survey on factual knowledge probing in pre-trained language models
Youssef, P., Koraş, O., Li, M., Schlötterer, J., and Seifert, C · 2023
Later among the works it cites.
Towards best practices of activation patching in language models: Metrics and methods
Original
Zhang, F. and Nanda, N · 2023
Later among the works it cites.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Original
Zhong, Z., Wu, Z., Manning, C. D., Potts, C., and Chen, D · 2023
Later among the works it cites.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Gianinazzi, L., Gajda, J., Lehmann, T., Podstawski, M., Niewiadomski, H., Nyczyk, P., et al · 2024
Closest in time.
Understanding and patching compositional reasoning in llms
Original
Li, Z., Jiang, G., Xie, H., Song, L., Lian, D., and Wei, Y · 2024
Closest in time.
The expressive power of transformers with chain of thought
Merrill, W. and Sabharwal, A · 2024
Closest in time.