Fetching the paper…
Reading the bibliography…
Context utilisation, the ability of Language Models (LMs) to incorporate relevant information from the provided context when generating responses, remains largely opaque to users, who cannot determine whether models draw from parametric memory or provided context, nor identify which specific context pieces inform the response.
Keeping the neural networks simple by minimizing the description length of the weights
Hinton, Geoffrey E. and Drew van Camp. 1993 · 1993
Earlier work this paper cites.
The magical number seven, plus or minus two: Some limits on our capacity for processing information
Baddeley, Alan, Richard M Shiffrin, Robert M Nosofsky, and George A Miller. 1994 · 1994
Earlier work this paper cites.
The minimum description length principle
Grünwald, Peter D. 2007 · 2007
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, Matthew D and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Li, Jiwei, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
Axiomatic Attribution for Deep Networks
Sundararajan, Mukund, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Towards better understanding of gradient-based attribution methods for deep neural networks
Ancona, Marco, Enea Ceolini, Cengiz Öztireli, and Markus Gross. 2018 · 2018
Earlier work this paper cites.
The description length of deep learning models
Blier, Léonard and Yann Ollivier. 2018 · 2018
Earlier work this paper cites.
Sharp nearby, fuzzy far away: How neural language models use context
Khandelwal, Urvashi, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Earlier work this paper cites.
A benchmark for interpretability methods in deep neural networks
Hooker, Sara, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. 2019 · 2019
Earlier work this paper cites.
Attention is not Explanation
Jain, Sarthak and Byron C. Wallace. 2019a · 2019
Earlier work this paper cites.
Attention is not explanation
Jain, Sarthak and Byron C. Wallace. 2019b · 2019
Earlier work this paper cites.
The (un) reliability of saliency methods
Kindermans, Pieter-Jan, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. 2019 · 2019
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Miller, Tim. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, Alec, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Is attention interpretable?
Serrano, Sofia and Noah A. Smith. 2019 · 2019
Earlier work this paper cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Tenney, Ian, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Earlier work this paper cites.
Attention is not not explanation
Wiegreffe, Sarah and Yuval Pinter. 2019a · 2019
Cited alongside, same era.
Attention is not not explanation
Wiegreffe, Sarah and Yuval Pinter. 2019b · 2019
Cited alongside, same era.
Quantifying attention flow in transformers
Abnar, Sara and Willem Zuidema. 2020 · 2020
Cited alongside, same era.
A diagnostic study of explainability techniques for text classification
Atanasova, Pepa, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 · 2020
Cited alongside, same era.
ERASER: A Benchmark to Evaluate Rationalized NLP Models
DeYoung, Jay, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Cited alongside, same era.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Jacovi, Alon and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
Detecting knowledge conflicts in large language models via representation patching
Wang, Xinyu, Hang Zhou, Jiacheng Liu, and Maosong Sun. 2023 · 2023
Later among the works it cites.
Characterizing mechanisms for factual recall in language models
Yu, Qinan, Jack Merullo, and Ellie Pavlick. 2023 · 2023
Later among the works it cites.
Understanding the interplay between parametric and contextual knowledge for large language models
Cheng, Sitao, Liangming Pan, Xunjian Yin, Xinyi Wang, and William Yang Wang. 2024 · 2024
Later among the works it cites.
Selfcite: Self-supervised alignment for context attribution in large language models
Chuang, Yung-Sung, Benjamin Cohen-Wang, Zejiang Shen, Zhaofeng Wu, Hu Xu, Xi Victoria Lin, James R Glass, Shang-Wen Li, and Wen-tau Yih. 2024 · 2024
Later among the works it cites.
Contextcite: Attributing model generation to context
Cohen-Wang, Benjamin, Harshay Shah, Kristian Georgiev, and Aleksander Madry. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Information-theoretic probing with minimum description length
Voita, Elena and Ivan Titov. 2020 · 2020
Cited alongside, same era.
QED: A framework and dataset for explanations in question answering
Lamm, Matthew, Jennimaria Palomaki, Chris Alberti, Daniel Andor, Eunsol Choi, Livio Baldini Soares, and Michael Collins. 2021 · 2021
Cited alongside, same era.
Discretized integrated gradients for explaining language models
Sanyal, Soumya and Xiang Ren. 2021 · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, Jianlin, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2021 · 2021
Cited alongside, same era.
Diagnostics-Guided Explanation Generation
Atanasova, Pepa, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2022 · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Meng, Kevin, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Cutting off the head ends the conflict: A mechanism for interpreting and mitigating knowledge conflicts in language models
Jin, Zhuoran, Pengfei Cao, Hongbang Yuan, Yubo Chen, Jiexin Xu, Huaijun Li, Xiaojian Jiang, Kang Liu, and Jun Zhao. 2024 · 2024
Later among the works it cites.
Lost in the middle: How language models use long contexts
Liu, Nelson F., Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024 · 2024
Later among the works it cites.
Towards faithful model explanation in NLP: A survey
Lyu, Qing, Marianna Apidianaki, and Chris Callison-Burch. 2024 · 2024
Later among the works it cites.
A glitch in the matrix? locating and detecting language model grounding with fakepedia
Monea, Giovanni, Maxime Peyrard, Martin Josifoski, Vishrav Chaudhary, Jason Eisner, Emre Kıcıman, Hamid Palangi, Barun Patra, and Robert West. 2024b · 2024
Later among the works it cites.
IRCAN: Mitigating knowledge conflicts in LLM generation via identifying and reweighting context-aware neurons
Shi, Dan, Renren Jin, Tianhao Shen, Weilong Dong, Xinwei Wu, and Deyi Xiong. 2024 · 2024
Later among the works it cites.
Where’s the head? locating knowledge -
Wang, Ziqi, Yiming Deng, Ximing Liu, and Zhiyuan Liu. 2024 · 2024
Later among the works it cites.
Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts
Xie, Jian, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2024 · 2024
Later among the works it cites.
Cub: Benchmarking context utilisation techniques for language models
Hagström, Lovisa, Youna Kim, Haeun Yu, Sang-goo Lee, Richard Johansson, Hyunsoo Cho, and Isabelle Augenstein. 2025 · 2025
Closest in time.
Qwen2.5 technical report
Qwen Team. 2025 · 2025
Closest in time.
Evaluating input feature explanations through a unified diagnostic evaluation framework
Sun, Jingyi, Pepa Atanasova, and Isabelle Augenstein. 2025 · 2025
Closest in time.
On the emergence of position bias in transformers
Wu, Xinyi, Yifei Wang, Stefanie Jegelka, and Ali Jadbabaie. 2025 · 2025
Closest in time.