Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are capable of generating plausible explanations of how they arrived at an answer to a question.
Prior distributions for variance parameters in hierarchical models (comment on article by Browne and Draper)
Andrew Gelman · 2006
Earlier work this paper cites.
Bayesian measures of explained variance and pooling in multilevel (hierarchical) models
Andrew Gelman and Iain Pardoe · 2006
Earlier work this paper cites.
Identification of conditional interventional distributions
Ilya Shpitser and Judea Pearl · 2006
Earlier work this paper cites.
Quantifying causal influences
Dominik Janzing, David Balduzzi, Moritz Grosse-Wentrup, and Bernhard Schölkopf · 2013
Earlier work this paper cites.
The no-u-turn sampler: adaptively setting path lengths in hamiltonian monte carlo
Matthew D Hoffman, Andrew Gelman, et al · 2014
Earlier work this paper cites.
Explaining predictions of non-linear classifiers in NLP
Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek · 2016
Earlier work this paper cites.
How social-class stereotypes maintain inequality
Federica Durante and Susan T Fiske · 2017
Earlier work this paper cites.
Learning to explain: An information-theoretic perspective on model interpretation
Jianbo Chen, Le Song, Martin Wainwright, and Michael Jordan · 2018
Earlier work this paper cites.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Earlier work this paper cites.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace · 2019
Earlier work this paper cites.
Is attention interpretable?
Sofia Serrano and Noah A. Smith · 2019
Earlier work this paper cites.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace · 2020
Earlier work this paper cites.
A framework for understanding sources of harm throughout the machine learning life cycle. equity and access in algorithms, mechanisms, and optimization, 1–9, 2021
Harini Suresh and John Guttag · 2021
Earlier work this paper cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld · 2021
Earlier work this paper cites.
Cebab: Estimating the causal effects of real-world concepts on nlp model behavior
Eldar D Abraham, Karel D’Oosterlinck, Amir Feder, Yair Gat, Atticus Geiger, Christopher Potts, Roi Reichart, and Zhengxuan Wu · 2022
Cited alongside, same era.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman · 2022
Cited alongside, same era.
Direct and indirect effects
Judea Pearl · 2022
Cited alongside, same era.
Faithfulness tests for natural language explanations
Pepa Atanasova, Oana-Maria Camburu, Christina Lioma, Thomas Lukasiewicz, Jakob Grue Simonsen, and Isabelle Augenstein · 2023
Cited alongside, same era.
Can large language models explain themselves? a study of llm-generated self-explanations
Shiyuan Huang, Siddarth Mamidanna, Shreedhar Jangam, Yilun Zhou, and Leilani H Gilpin · 2023
Cited alongside, same era.
Do models explain themselves? Counterfactual simulatability of natural language explanations
Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao, He He, Jacob Steinhardt, Zhou Yu, and Kathleen Mckeown · 2024
Later among the works it cites.
Disability myths and stereotypes
Australian Public Service Commission · 2024
Later among the works it cites.
Faithful explanations of black-box NLP models using llm-generated counterfactuals
Yair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin, Amit Sharma, and Roi Reichart · 2024
Later among the works it cites.
The intersection of substance use stigma and anti-black racial stigma: A scoping review
Rashmi Ghonasgi, Maria E Paschke, Rachel P Winograd, Catherine Wright, Eva Selph, and Devin E Banks · 2024
Later among the works it cites.
Towards faithful model explanation in nlp: A survey
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gender bias and stereotypes in large language models
Hadas Kotek, Rikker Dockum, and David Sun · 2023
Cited alongside, same era.
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al · 2023
Cited alongside, same era.
Faithful chain-of-thought reasoning
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch · 2023
Cited alongside, same era.
Question decomposition improves the faithfulness of model-generated reasoning
Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Carson Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamile Lukosiute, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Sam McCandlish, Sheer El Showk, Tamera Lanham, Tim Maxwell, Venkatesa Chandrasekaran, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez · 2023
Cited alongside, same era.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman · 2023
Cited alongside, same era.
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Cited alongside, same era.
Meet claude
Anthropic · 2024
Cited alongside, same era.
Andreas Madsen, Sarath Chandar, and Siva Reddy · 2024
Later among the works it cites.
Introducing llama 3.1: Our most capable models to date
Meta · 2024
Later among the works it cites.
Gender stereotyping
OHCHR · 2024
Later among the works it cites.
Openai model index for researchers
OpenAI · 2024
Later among the works it cites.
On measuring faithfulness or self-consistency of natural language explanations
Letitia Parcalabescu and Anette Frank · 2024
Later among the works it cites.
Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning
Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings · 2024
Later among the works it cites.
The probabilities also matter: A more faithful metric for faithfulness of free-text explanations in large language models
Noah Siegel, Oana-Maria Camburu, Nicolas Heess, and Maria Perez-Ortiz · 2024
Later among the works it cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits · 2076
Closest in time.