Fetching the paper…
Reading the bibliography…
Despite the increasing use of large language models (LLMs) for context-grounded tasks like summarization and question-answering, understanding what makes an LLM produce a certain response is challenging.
Bradley-Terry models in R
David Firth · 2005
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
"Why should I trust you?" Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Earlier work this paper cites.
L-Shapley and C-Shapley: Efficient model interpretation for structured data
Jianbo Chen, Le Song, Martin J. Wainwright, and Michael I. Jordan · 2019
Earlier work this paper cites.
Hierarchical interpretations for neural network predictions
Chandan Singh, W. James Murdoch, and Bin Yu · 2019
Cited alongside, same era.
Generating hierarchical explanations on text classification via feature interaction detection
Hanjie Chen, Guangtao Zheng, and Yangfeng Ji · 2020
Cited alongside, same era.
LS-Tree: Model interpretation when the data are linguistic
Jianbo Chen and Michael Jordan · 2020
Cited alongside, same era.
spaCy: Industrial-strength natural language processing in python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd · 2020
Cited alongside, same era.
Towards hierarchical importance attribution: Explaining compositional semantics for neural sequence models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue, and Xiang Ren · 2020
Cited alongside, same era.
Interpretation of NLP models through input marginalization
Double trouble: How to not explain a text classifier’s decisions using counterfactuals synthesized by masked language models?
Thang Pham, Trung Bui, Long Mai, and Anh Nguyen · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Later among the works it cites.
A hierarchical explanation generation method based on feature interaction detection
Yiming Ju, Yuanzhe Zhang, Kang Liu, and Jun Zhao · 2023
Later among the works it cites.
Using Captum to explain generative language models, 2023
Vivek Miglani, Aobo Yang, Aram H. Markosyan, Diego Garcia-Olano, and Narine Kokhlikyan · 2023
Later among the works it cites.
GLIME: General, stable and local LIME explanation
Zeren Tan, Yang Tian, and Jian Li · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Siwon Kim, Jihun Yi, Eunji Kim, and Sungroh Yoon · 2020
Cited alongside, same era.
BERTScore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi · 2020
Cited alongside, same era.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Cited alongside, same era.
An analysis of LIME for text data
Dina Mardaoui and Damien Garreau · 2021
Cited alongside, same era.
BARTScore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu · 2021
Cited alongside, same era.
Scaling instruction-finetuned language models, 2022
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei · 2022
Cited alongside, same era.
SHAP-based explanation methods: A review for NLP interpretability
Edoardo Mosca, Ferenc Szigeti, Stella Tragianni, Daniel Gallagher, and Georg Groh · 2022
Cited alongside, same era.
Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler · 2023
Later among the works it cites.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman · 2023
Later among the works it cites.
IBM Generative AI Python SDK (Tech Preview)
IBM · 2024
Closest in time.
SocialStigmaQA: A benchmark to uncover stigma amplification in generative language models
Manish Nagireddy, Lamogha Chiazor, Moninder Singh, and Ioana Baldini · 2024
Closest in time.
Text to text explanation: Abstractive summarization example
SHAP · 2024
Closest in time.
Text examples
SHAP · 2024
Closest in time.