Fetching the paper…
Reading the bibliography…
Information flows by routes inside the network via mechanisms implemented in the model.
Visualizing data using t-sne
van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
A universal part-of-speech tagset
Petrov, S., Das, D., and McDonald, R · 2012
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Bau, A., Belinkov, Y., Sajjad, H., Durrani, N., Dalvi, F., and Glass, J · 2019
Earlier work this paper cites.
Adaptively sparse transformers
Correia, G. M., Niculae, V., and Martins, A. F. T · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I · 2019
Earlier work this paper cites.
The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?
Bastings, J. and Filippova, K · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Neural natural language inference models partially embed theories of lexical entailment and negation
Geiger, A., Richardson, K., and Potts, C · 2020
Earlier work this paper cites.
Attention is not only a weight: Analyzing transformers with vector norms
Kobayashi, G., Kuribayashi, T., Yokoi, S., and Inui, K · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Earlier work this paper cites.
Opt: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways, 2022
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Cited alongside, same era.
Knowledge neurons in pretrained transformers
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F · 2022
Cited alongside, same era.
A circuit for python docstrings in a 4-layer attention-only transformer, 2023
Heimersheim, S. and Janiak, J · 2023
Later among the works it cites.
In-context learning creates task vectors, 2023
Hendel, R., Geva, M., and Globerson, A · 2023
Later among the works it cites.
The hydra effect: Emergent self-repair in language model computations, 2023
McGrath, T., Rahtz, M., Kramar, J., Mikulik, V., and Legg, S · 2023
Later among the works it cites.
Traveling words: A geometric interpretation of transformers, 2023
Molina, R · 2023
Later among the works it cites.
Understanding arithmetic reasoning in language models using causal mediation analysis, 2023
Stolfo, A., Belinkov, Y., and Sachan, M · 2023
Later among the works it cites.
Attribution patching outperforms automated circuit discovery, 2023
Syed, A., Rager, C., and Conmy, A · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring the mixing of contextual information in the transformer
Ferrando, J., Gállego, G. I., and Costa-jussà, M. R · 2022
Cited alongside, same era.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Goyal, N., Gao, C., Chaudhary, V., Chen, P.-J., Wenzek, G., Ju, D., Krishnan, S., Ranzato, M., Guzmán, F., and Fan, A · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Cited alongside, same era.
The singular value decompositions of transformer weight matrices are highly interpretable, 2022
Millidge, B. and Black, S · 2022
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation, 2022
NLLB, T., Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Gonzalez, G. M., Hansanti, P., Hoffman, J., Jarrett, S., Sadagopan, K. R., Rowe, D., Spruit, S., Tran, C., Andrews, P., Ayan, N. F., Bhosale, S., Edunov, S., Fan, A., Gao, C., Goswami, V., Guzmán, F., Koehn, P., Mourachko, A., Ropers, C., Saleem, S., Schwenk, H., and Wang, J · 2022
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Cited alongside, same era.
Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity Even better
Dale, D., Voita, E., Barrault, L., and Costa-jussà, M. R · 2023
Cited alongside, same era.
Dale, D., Voita, E., Lam, J., Hansanti, P., Ropers, C., Kalbassi, E., Gao, C., Barrault, L., and Costa-jussà, M. R · 2023
Cited alongside, same era.
Later among the works it cites.
Function vectors in large language models, 2023
Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2023
Later among the works it cites.
Neurons in large language models: Dead, n-gram, positional, 2023
Voita, E., Ferrando, J., and Nalmpantis, C · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2023
Later among the works it cites.
Spectral filters, dark signals, and attention sinks, 2024
Cancedda, N · 2024
Closest in time.
Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms, 2024
Hanna, M., Pezzelle, S., and Belinkov, Y · 2024
Closest in time.
Atp*: An efficient and scalable method for localizing llm behaviour to components, 2024
Kramár, J., Lieberum, T., Shah, R., and Nanda, N · 2024
Closest in time.
Explorations of self-repair in language models, 2024
Rushing, C. and Nanda, N · 2024
Closest in time.