Fetching the paper…
Reading the bibliography…
Since the advent of ChatGPT, Large Language Models (LLMs) have excelled in various tasks but remain as black-box systems.
(No Title) ( 381)
Tulving, E. (1972). “episodic and semantic memory,” in organization of memory · 1972
Earlier work this paper cites.
Mental models. Towards a cognitive science of language, inference, and consciousness
Johnson-Laird, P · 1983
Earlier work this paper cites.
Psychological review 95 , 163
Kintsch, W. (1988). The role of knowledge in discourse comprehension: a construction-integration model · 1988
Earlier work this paper cites.
Psychological review 99 , 195
Squire, L. R. (1992). Memory and the hippocampus: a synthesis from findings with rats, monkeys, and humans · 1992
Earlier work this paper cites.
Semantic structures vol. 18
Jackendoff, R. S · 1992
Earlier work this paper cites.
Trends in cognitive sciences 3 , 223–232
Levelt, W. J. (1999). Models of word production · 1999
Earlier work this paper cites.
Preprint at arXiv
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020). Scaling laws for neural language models · 2001
Earlier work this paper cites.
Preprint at arXiv
Shazeer, N. (2020). Glu variants improve transformer · 2002
Earlier work this paper cites.
Cerebral Cortex 12 , 818–830
Krichmar, J. L., and Edelman, G. M. (2002). Machine psychology: autonomous behavior, perceptual categorization and conditioning in a brain-based device · 2002
Earlier work this paper cites.
Annual review of psychology 54 , 115–144
Staddon, J. E., and Cerutti, D. T. (2003). Operant conditioning · 2003
Earlier work this paper cites.
International Journal of Cognitive Informatics and Natural Intelligence (IJCINI) 1 , 66–77
Wang, Y. (2007). The oar model of neural informatics for internal knowledge representation in the brain · 2007
Earlier work this paper cites.
Cognitive systems research 11 , 81–92
Wang, Y., and Chiew, V. (2010). On the cognitive process of human problem solving · 2010
Earlier work this paper cites.
The number sense: How the mind creates mathematics
Dehaene, S · 2011
Earlier work this paper cites.
arXiv preprint arXiv:1306.0125
Whitehill, J. (2013). Understanding act-r-an outsider’s perspective · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. (2013) · 2013
Earlier work this paper cites.
Rules of the mind
Anderson, J. R · 2014
Earlier work this paper cites.
Aspects of the Theory of Syntax
Chomsky, N · 2014
Earlier work this paper cites.
Science 350 , 1332–1338
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B. (2015). Human-level concept learning through probabilistic program induction · 2015
Earlier work this paper cites.
Advances in neural information processing systems 28
Zhang, X., Zhao, J., and LeCun, Y. (2015). Character-level convolutional networks for text classification · 2015
Earlier work this paper cites.
Advances in neural information processing systems 30
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L. (2017) · 2017
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Gilpin, L. H., Bau, D., Yuan, B. Z., Bajwa, A., Specter, M., and Kagal, L. (2018) · 2018
Earlier work this paper cites.
Queue 16 , 31–57
Lipton, Z. C. (2018). The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery · 2018
Earlier work this paper cites.
Digital signal processing 73 , 1–15
Montavon, G., Samek, W., and Müller, K.-R. (2018). Methods for interpreting and understanding deep neural networks · 2018
Earlier work this paper cites.
Preprint at arXiv
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding · 2018
Earlier work this paper cites.
Personality neuroscience 1 , e17
Rouault, M., McWilliams, A., Allen, M. G., and Fleming, S. M. (2018). Human metacognition across domains: insights from individual differences and neuroimaging · 2018
Earlier work this paper cites.
An analysis of encoder representations in transformer-based machine translation
Raganato, A., and Tiedemann, J. (2018) · 2018
Earlier work this paper cites.
arXiv preprint arXiv:1808.08745
Narayan, S., Cohen, S. B., and Lapata, M. (2018). Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization · 2018
Earlier work this paper cites.
OpenAI blog 1 , 9
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I. et al. (2019). Language models are unsupervised multitask learners · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
Kovaleva, O., Romanov, A., Rogers, A., and Rumshisky, A. (2019) · 2019
Earlier work this paper cites.
BERT has a mouth, and it must speak: BERT as a Markov random field language model
Wang, A., and Cho, K. (2019) · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I. (2019) · 2019
Earlier work this paper cites.
Adaptively sparse transformers
Correia, G. M., Niculae, V., and Martins, A. F. T. (2019) · 2019
Earlier work this paper cites.
Text Generation from Knowledge Graphs with Graph Transformers
Koncel-Kedziorski, R., Bekal, D., Luan, Y., Lapata, M., and Hajishirzi, H. (2019) · 2019
Earlier work this paper cites.
Understanding the difficulty of training transformers
Liu, L., Liu, X., Gao, J., Chen, W., and Han, J. (2020a) · 2020
Earlier work this paper cites.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T. (2020) · 2020
Earlier work this paper cites.
Distill 5 , e00024–001
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. (2020). Zoom in: An introduction to circuits · 2020
Earlier work this paper cites.
Advances in neural information processing systems 33 , 12388–12401
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S. (2020). Investigating gender bias in language models using causal mediation analysis · 2020
Earlier work this paper cites.
Preprint at arXiv
Mollas, I., Chrysopoulou, Z., Karlos, S., and Tsoumakas, G. (2020). Ethos: an online hate speech detection dataset · 2020
Earlier work this paper cites.
The heads hypothesis: A unifying statistical approach towards understanding multi-headed attention in bert
Pande, M., Budhraja, A., Nema, P., Kumar, P., and Khapra, M. M. (2021) · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
Geiger, A., Lu, H., Icard, T., and Potts, C. (2021) · 2021
Earlier work this paper cites.
Transformer Circuits Thread
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T. et al. (2021). A mathematical framework for transformer circuits · 2021
Cited alongside, same era.
Preprint at arXiv
Santana, A., and Colombini, E. (2021). Neural attention models in deep learning: Survey and taxonomy · 2021
Cited alongside, same era.
ACM Transactions on Intelligent Systems and Technology (TIST) 12 , 1–32
Chaudhari, S., Mithal, V., Polatkan, G., and Ramanath, R. (2021). An attentive survey of attention models · 2021
Cited alongside, same era.
IEEE Transactions on Knowledge and Data Engineering 35 , 3279–3298
Brauwers, G., and Frasincar, F. (2021). A general survey on attention mechanisms in deep learning · 2021
Cited alongside, same era.
Proceedings of the National Academy of Sciences 118 , e2105646118
Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., and Fedorenko, E. (2021). The neural architecture of language: Integrative modeling converges on predictive processing · 2021
Cited alongside, same era.
Preprint at arXiv
Luo, H., and Specia, L. (2024). From understanding to utilization: A survey on explainability for large language models · 2024
Closest in time.
Scientific Reports 14 , 21445
Marjieh, R., Sucholutsky, I., van Rijn, P., Jacoby, N., and Griffiths, T. L. (2024). Large language models predict human sensory judgments across six modalities · 2024
Closest in time.
arXiv preprint arXiv:2401.17671
Mischler, G., Li, Y. A., Bickel, S., Mehta, A. D., and Mesgarani, N. (2024). Contextual feature extraction hierarchies converge in large language models and the brain · 2024
Closest in time.
Advances in Neural Information Processing Systems 36
Bietti, A., Cabannes, V., Bouchacourt, D., Jegou, H., and Bottou, L. (2024). Birth of a transformer: A memory viewpoint · 2024
Closest in time.
arXiv preprint arXiv:2411.10115
Dana, L., Pydi, M. S., and Chevaleyre, Y. (2024). Memorization in attention-only transformers · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Millidge, B., Seth, A., and Buckley, C. L. (2021). Predictive coding: a theoretical and experimental review · 2021
Cited alongside, same era.
Preprint at arXiv
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E. et al. (2021). On the opportunities and risks of foundation models · 2021
Cited alongside, same era.
Language models are few-shot multilingual learners
Winata, G. I., Madotto, A., Lin, Z., Liu, R., Yosinski, J., and Fung, P. (2021) · 2021
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O. (2021) · 2021
Cited alongside, same era.
arXiv preprint arXiv:2110.14168
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R. et al. (2021). Training verifiers to solve math word problems · 2021
Cited alongside, same era.
Journal of Machine Learning Research 23 , 1–39
Fedus, W., Zoph, B., and Shazeer, N. (2022). Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity · 2022
Cited alongside, same era.
Neel Nanda’s Blog 1
Nanda, N. (2022). A comprehensive mechanistic interpretability explainer & glossary · 2022
Cited alongside, same era.
Cutting off the head ends the conflict: A mechanism for interpreting and mitigating knowledge conflicts in language models
Jin, Z., Cao, P., Yuan, H., Chen, Y., Xu, J., Li, H., Jiang, X., Liu, K., and Zhao, J. (2024a) · 2024
Closest in time.
Preprint at arXiv
Ferrando, J., and Voita, E. (2024). Information flow routes: Automatically interpreting language models at scale · 2024
Closest in time.
Preprint at arXiv
Wu, W., Wang, Y., Xiao, G., Peng, H., and Fu, Y. (2024). Retrieval head mechanistically explains long-context factuality · 2024
Closest in time.
Preprint at arXiv
Tang, H., Lin, Y., Lin, J., Han, Q., Hong, S., Yao, Y., and Wang, G. (2024). Razorattention: Efficient kv cache compression through retrieval heads · 2024
Closest in time.
How does gpt-2 predict acronyms? extracting and understanding a circuit via mechanistic interpretability
García-Carrasco, J., Maté, A., and Trujillo, J. C. (2024) · 2024
Closest in time.
Circuit component reuse across tasks in transformer language models
Merullo, J., Eickhoff, C., and Pavlick, E. (2024) · 2024
Closest in time.
Identifying semantic induction heads to understand in-context learning
Ren, J., Guo, Q., Yan, H., Liu, D., Zhang, Q., Qiu, X., and Lin, D. (2024) · 2024
Closest in time.
Preprint at arXiv
Chughtai, B., Cooney, A., and Nanda, N. (2024). Summing up the facts: Additive mechanisms behind factual recall in llms · 2024
Closest in time.
Function vectors in large language models
Todd, E., Li, M., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D. (2024) · 2024
Closest in time.
Preprint at arXiv
Edelman, B. L., Edelman, E., Goel, S., Malach, E., and Tsilivis, N. (2024). The evolution of statistical induction heads: In-context learning markov chains · 2024
Closest in time.
What needs to go right for an induction head? a mechanistic study of in-context learning circuits and their formation
Singh, A. K., Moskovitz, T., Hill, F., Chan, S. C., and Saxe, A. M. (2024) · 2024
Closest in time.
Preprint at arXiv
Ji-An, L., Zhou, C. Y., Benna, M. K., and Mattar, M. G. (2024). Linking in-context learning in transformers to human episodic memory · 2024
Closest in time.
Preprint at arXiv
Crosbie, J. (2024). Induction heads as an essential mechanism for pattern matching in in-context learning · 2024
Closest in time.
The mechanistic basis of data dependence and abrupt learning in an in-context classification task
Reddy, G. (2024) · 2024
Closest in time.
In-context language learning: Architectures and algorithms
Akyürek, E., Wang, B., Kim, Y., and Andreas, J. (2024) · 2024
Closest in time.
Preprint at arXiv
Hoscilowicz, J., Wiacek, A., Chojnacki, J., Cieslak, A., Michon, L., Urbanevych, V., and Janicki, A. (2024). Nl-iti: Optimizing probing and intervention for improvement of iti method · 2024
Closest in time.
Steering large language models for cross-lingual information retrieval
Guo, P., Ren, Y., Hu, Y., Cao, Y., Li, Y., and Huang, H. (2024) · 2024
Closest in time.
Enhancing semantic consistency of large language models through model editing: An interpretability-oriented approach
Yang, J., Chen, D., Sun, Y., Li, R., Feng, Z., and Peng, W. (2024) · 2024
Closest in time.
Detecting and understanding vulnerabilities in language models via mechanistic interpretability
García-Carrasco, J., Maté, A., and Trujillo, J. (2024) · 2024
Closest in time.
Iteration head: A mechanistic study of chain-of-thought
Cabannes, V., Arnal, C., Bouaziz, W., Yang, X. A., Charton, F., and Kempe, J. (2024) · 2024
Closest in time.
Successor heads: Recurring, interpretable attention heads in the wild
Gould, R., Ong, E., Ogden, G., and Conmy, A. (2024) · 2024
Closest in time.
Preprint at arXiv
Kim, G., Valentino, M., and Freitas, A. (2024). A mechanistic interpretation of syllogistic reasoning in auto-regressive language models · 2024
Closest in time.
Preprint at arXiv
Wiegreffe, S., Tafjord, O., Belinkov, Y., Hajishirzi, H., and Sabharwal, A. (2024). Answer, assemble, ace: Understanding how transformers answer multiple choice questions · 2024
Closest in time.
On the difficulty of faithful chain-of-thought reasoning in large language models
Tanneru, S. H., Ley, D., Agarwal, C., and Lakkaraju, H. (2024) · 2024
Closest in time.
Preprint at arXiv
Fu, T., Huang, H., Ning, X., Zhang, G., Chen, B., Wu, T., Wang, H., Huang, Z., Li, S., Yan, S. et al. (2024). Moa: Mixture of sparse attention for automatic large language model compression · 2024
Closest in time.
Preprint at arXiv
Yu, Z., and Ananiadou, S. (2024). How do large language models learn in-context? query and key matrices of in-context heads are two towers for metric learning · 2024
Closest in time.
Competition of mechanisms: Tracing how language models handle facts and counterfactuals
Ortu, F., Jin, Z., Doimo, D., Sachan, M., Cazzaniga, A., and Schölkopf, B. (2024) · 2024
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Geiger, A., Wu, Z., Potts, C., Icard, T., and Goodman, N. (2024) · 2024
Closest in time.
The linear representation hypothesis and the geometry of large language models
Park, K., Choe, Y. J., and Veitch, V. (2024) · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods
Zhang, F., and Nanda, N. (2024) · 2024
Closest in time.
Linearity of relation decoding in transformer language models
Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D. (2024) · 2024
Closest in time.
Neurons in large language models: Dead, n-gram, positional
Voita, E., Ferrando, J., and Nalmpantis, C. (2024) · 2024
Closest in time.
Preprint at arXiv
Lv, A., Zhang, K., Chen, Y., Wang, Y., Liu, L., Wen, J.-R., Xie, J., and Yan, R. (2024). Interpreting key mechanisms of factual recall in transformer-based language models · 2024
Closest in time.
Preprint at arXiv
Johansson, R., Hammer, P., and Lofthouse, T. (2024). Functional equivalence with nars · 2024
Closest in time.
NewsBench: A systematic evaluation framework for assessing editorial capabilities of large language models in Chinese journalism
Li, M., Chen, M.-B., Tang, B., ShengbinHou, S., Wang, P., Deng, H., Li, Z., Xiong, F., Mao, K., Peng, C., and Luo, Y. (2024d) · 2024
Closest in time.
T-eval: Evaluating the tool utilization capability of large language models step by step
Chen, Z., Du, W., Zhang, W., Liu, K., Liu, J., Zheng, M., Zhuo, J., Zhang, S., Lin, D., Chen, K., and Zhao, F. (2024b) · 2024
Closest in time.