Fetching the paper…
Reading the bibliography…
How do language models use information provided as context when generating a response? Can we infer whether a particular generated statement is actually grounded in the context, a misinterpretation, or fabricated? To help answer these questions, we introduce the problem of context attribution: pinpointing the parts of the context (if any) that led a model to generate a particular statement.
“The Proof and Measurement of Association between Two Things”
Charles Spearman · 1904
Earlier work this paper cites.
“A value for n-person games”
Lloyd Shapley · 1953
Earlier work this paper cites.
“Design and Analysis of Computer Experiments”
Jerome Sacks, William. Welch, Toby. Mitchell and Henry. Wynn · 1989
Earlier work this paper cites.
“Regression Shrinkage and Selection Via the Lasso”
Robert Tibshirani · 1994
Earlier work this paper cites.
“Natural language processing with Python: analyzing text with the natural language toolkit”
Steven Bird, Ewan Klein and Edward Loper · 2009
Earlier work this paper cites.
“Scikit-learn: Machine Learning in Python”
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot and E. Duchesnay · 2011
Earlier work this paper cites.
“Deep inside convolutional networks: Visualising image classification models and saliency maps”
Karen Simonyan, Andrea Vedaldi and Andrew Zisserman · 2013
Earlier work this paper cites.
“Visualizing and understanding neural models in NLP”
Jiwei Li, Xinlei Chen, Eduard Hovy and Dan Jurafsky · 2015
Earlier work this paper cites.
“Ms marco: A human-generated machine reading comprehension dataset”, 2016
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder and Li Deng · 2016
Earlier work this paper cites.
“Abstractive text summarization using sequence-to-sequence rnns and beyond”
Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre and Bing Xiang · 2016
Earlier work this paper cites.
““Why Should I Trust You?”: Explaining the Predictions of Any Classifier”
Marco Ribeiro, Sameer Singh and Carlos Guestrin · 2016
Earlier work this paper cites.
“Not just a black box: Learning important features through propagating activation differences”
Avanti Shrikumar, Peyton Greenside, Anna Shcherbina and Anshul Kundaje · 2016
Earlier work this paper cites.
“Visualizing and understanding neural machine translation”
Yanzhuo Ding, Yang Liu, Huanbo Luan and Maosong Sun · 2017
Earlier work this paper cites.
“A unified approach to interpreting model predictions”
Scott Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
“Interactive visualization and manipulation of attention-based neural machine translation”
Jaesong Lee, Joong-Hwi Shin and Jun-Seok Kim · 2017
Earlier work this paper cites.
“SmoothGrad: removing noise by adding noise”
D. Smilkov, N. Thorat, B. Kim, F. Viégas and M. Wattenberg · 2017
Earlier work this paper cites.
“HotpotQA: A dataset for diverse, explainable multi-hop question answering”
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov and Christopher Manning · 2018
Earlier work this paper cites.
“DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs”
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh and Matt Gardner · 2019
Earlier work this paper cites.
“Attention is not explanation”
Sarthak Jain and Byron Wallace · 2019
Earlier work this paper cites.
“Natural questions: a benchmark for question answering research”
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin and Kenton Lee · 2019
Earlier work this paper cites.
“Sentence-bert: Sentence embeddings using siamese bert-networks”
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
“bLIMEy: Surrogate Prediction Explanations Beyond LIME”
Kacper Sokol, Alexander Hepburn, Raul Santos-Rodriguez and Peter Flach · 2019
Earlier work this paper cites.
Sofia Serrano and Noah Smith · 2019
Earlier work this paper cites.
“High-dimensional statistics: A non-asymptotic viewpoint”
Martin Wainwright · 2019
Earlier work this paper cites.
“Universal adversarial triggers for attacking and analyzing NLP”
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner and Sameer Singh · 2019
Cited alongside, same era.
“Attention is not not explanation”
Sarah Wiegreffe and Yuval Pinter · 2019
Cited alongside, same era.
“Quantifying attention flow in transformers”
Samira Abnar and Willem Zuidema · 2020
Cited alongside, same era.
“Tydi qa: A benchmark for information-seeking question answering in ty pologically di verse languages”
Jonathan Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev and Jennimaria Palomaki · 2020
Cited alongside, same era.
“Transformers: State-of-the-art natural language processing”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf and Morgan Funtowicz · 2020
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng and Bing Qin · 2023
Later among the works it cites.
Albert Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Chaplot, Diego Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample and Lucile Saulnier · 2023
Later among the works it cites.
“Calibrated language models must hallucinate”
Adam Kalai and Santosh Vempala · 2023
Later among the works it cites.
“Datainf: Efficiently estimating data influence in lora-tuned llms and diffusion models”
Yongchan Kwon, Eric Wu, Kevin Wu and James Zou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Explaining by removing: A unified framework for model explanation”
Ian Covert, Scott Lundberg and Su-In Lee · 2021
Cited alongside, same era.
“BERT meets shapley: Extending SHAP explanations to transformer-based classifiers”
Enja Kokalj, Blaž Škrlj, Nada Lavrač, Senja Pollak and Marko Robnik-Šikonja · 2021
Cited alongside, same era.
“Webgpt: Browser-assisted question-answering with human feedback”
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju and William Saunders · 2021
Cited alongside, same era.
“Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models”
Bernd Bohnet, Vinh Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Soares, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig and Kai Hui · 2022
Cited alongside, same era.
“Data curation alone can stabilize in-context learning”
Ting-Yun Chang and Robin Jia · 2022
Cited alongside, same era.
“Rarr: Researching and revising what language models say, using language models”
Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee and Da-Cheng Juan · 2022
Cited alongside, same era.
“Datamodels: Predicting Predictions from Training Data”
Andrew Ilyas, Sung Park, Logan Engstrom, Guillaume Leclerc and Aleksander Madry · 2022
Cited alongside, same era.
Nelson Liu, Tianyi Zhang and Percy Liang · 2023
Later among the works it cites.
“Factscore: Fine-grained atomic evaluation of factual precision in long form text generation”
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer and Hannaneh Hajishirzi · 2023
Later among the works it cites.
“In-context example selection with influences”
Tai Nguyen and Eric Wong · 2023
Later among the works it cites.
“TRAK: Attributing Model Behavior at Scale”
Sung Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc and Aleksander Madry · 2023
Later among the works it cites.
“Attention sorting combats recency bias in long context language models”
Alexander Peysakhovich and Adam Lerer · 2023
Later among the works it cites.
“Measuring attribution in natural language generation models”
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Tomar, Iulia Turc and David Reitter · 2023
Later among the works it cites.
“Quantifying the plausibility of context reliance in neural machine translation”
Gabriele Sarti, Grzegorz Chrupała, Malvina Nissim and Arianna Bisazza · 2023
Later among the works it cites.
“Unifying corroborative and contributive attributions in large language models”
Theodora Worledge, Judy Shen, Nicole Meister, Caleb Winston and Carlos Guestrin · 2023
Later among the works it cites.
“Universal and transferable adversarial attacks on aligned language models”
Andy Zou, Zifan Wang, J Kolter and Matt Fredrikson · 2023
Later among the works it cites.
“Phi-3 technical report: A highly capable language model locally on your phone”
Marah Abdin, Sam Jacobs, Ammar Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari and Harkirat Behl · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang and Angela Fan · 2024
Closest in time.
“DsDm: Model-Aware Dataset Selection with Datamodels”, 2024
Logan Engstrom, Axel Feldmann and Aleksander Madry · 2024
Closest in time.
“AtP*: An efficient and scalable method for localizing LLM behaviour to components”
János Kramár, Tom Lieberum, Rohin Shah and Neel Nanda · 2024
Closest in time.
“Lost in the middle: How language models use long contexts”
Nelson Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni and Percy Liang · 2024
Closest in time.
“Wait, It’s All Token Noise? Always Has Been: Interpreting LLM Behavior Using Shapley Value”
Behnam Mohammadi · 2024
Closest in time.
“Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks”
Dario Pasquini, Martin Strohmeier and Carmela Troncoso · 2024
Closest in time.
“Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation”, 2024
Jirui Qi, Gabriele Sarti, Raquel Fern’andez and Arianna Bisazza · 2024
Closest in time.
“Decomposing and editing predictions by modeling model computation”
Harshay Shah, Andrew Ilyas and Aleksander Madry · 2024
Closest in time.
“Backtracing: Retrieving the Cause of the Query”
Rose Wang, Pawan Wirawarn, Omar Khattab, Noah Goodman and Dorottya Demszky · 2024
Closest in time.
“Effective Large Language Model Adaptation for Improved Grounding and Citation Generation”
Xi Ye, Ruoxi Sun, Sercan Arik and Tomas Pfister · 2024
Closest in time.