Fetching the paper…
Reading the bibliography…
Recent advancements in NLP systems, particularly with the introduction of LLMs, have led to widespread adoption of these systems by a broad spectrum of users across various domains, impacting decision-making, the job market, society, and scientific research.
Medical artificial intelligence and human values
Kun-Hsing Yu, Elizabeth Healey, Tze-Yun Leong, Isaac S. Kohane, and Arjun Kumar Manrai. 2024 · 1904
Earlier work this paper cites.
Large language models in medicine
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023 · 1940
Earlier work this paper cites.
It is not true that transformers are inductive learners: Probing NLI models with external negation
Michael Sullivan. 2024 · 1945
Earlier work this paper cites.
Using test suites in evaluation of machine translation systems
Margaret King and Kirsten Falkedal. 1990 · 1990
Earlier work this paper cites.
Harms of gender exclusivity and challenges in non-binary representation in language technologies
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff M. Phillips, and Kai-Wei Chang. 2021 · 1994
Earlier work this paper cites.
TSNLP - test suites for natural language processing
Sabine Lehmann, Stephan Oepen, Sylvie Regnier-Prost, Klaus Netter, Veronika Lux, Judith Klein, Kirsten Falkedal, Frederik Fouvry, Dominique Estival, Eva Dauphin, Herve Compagnion, Judith Baur, Lorna Balkan, and Doug Arnold. 1996 · 1996
Earlier work this paper cites.
Humans and automation: Use, misuse, disuse, abuse
Raja Parasuraman and Victor Riley. 1997 · 1997
Earlier work this paper cites.
Wt5?! training text-to-text models to explain their predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2004
Earlier work this paper cites.
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang. 2020 · 2007
Earlier work this paper cites.
Using "annotator rationales" to improve machine learning for text categorization
Omar Zaidan, Jason Eisner, and Christine D. Piatko. 2007 · 2007
Earlier work this paper cites.
Topic modeling with contextualized word representation clusters
Laure Thompson and David Mimno. 2020 · 2010
Earlier work this paper cites.
The effect of wording on message propagation: Topic- and author-controlled natural experiments on twitter
Chenhao Tan, Lillian Lee, and Bo Pang. 2014 · 2014
Earlier work this paper cites.
Simlex-999: Evaluating semantic models with (genuine) similarity estimation
Felix Hill, Roi Reichart, and Anna Korhonen. 2015 · 2015
Earlier work this paper cites.
Ira Leviant and Roi Reichart. 2015 · 2015
Earlier work this paper cites.
Predicting judicial decisions of the european court of human rights: a natural language processing perspective
Nikolaos Aletras, Dimitrios Tsarapatsanis, Daniel Preotiuc-Pietro, and Vasileios Lampos. 2016 · 2016
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016 · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi S. Jaakkola. 2016 · 2016
Earlier work this paper cites.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
Behavioral measurement of trust in automation: the trust fall
David Miller, Mishel Johns, Brian Mok, Nikhil Gowda, David Sirkin, Key Lee, and Wendy Ju. 2016 · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
A linguistic evaluation of rule-based, phrase-based, and neural MT engines
Aljoscha Burchardt, Vivien Macketanz, Jon Dehdari, Georg Heigold, Jan-Thorsten Peter, and Philip Williams. 2017 · 2017
Earlier work this paper cites.
Evaluating the morphological competence of machine translation systems
Franck Burlot and François Yvon. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
European union regulations on algorithmic decision-making and a "right to explanation"
Bryce Goodman and Seth R. Flaxman. 2017 · 2017
Earlier work this paper cites.
Learning to reason: End-to-end module networks for visual question answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. 2017 · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. 2017 · 2017
Earlier work this paper cites.
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter W. Battaglia, and Tim Lillicrap. 2017 · 2017
Earlier work this paper cites.
How grammatical is character-level neural machine translation? assessing MT quality with contrastive translation pairs
Rico Sennrich. 2017 · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda B. Viégas, and Martin Wattenberg. 2017 · 2017
Earlier work this paper cites.
Inference is everything: Recasting semantic resources into a unified evaluation framework
Aaron Steven White, Pushpendre Rastogi, Kevin Duh, and Benjamin Van Durme. 2017 · 2017
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
What you can cram into a single \$&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, Germán Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Earlier work this paper cites.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem H. Zuidema. 2018 · 2018
Earlier work this paper cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Do language models understand anything? on the ability of lstms to understand negative polarity items
Jaap Jumelet and Dieuwke Hupkes. 2018 · 2018
Earlier work this paper cites.
The mythos of model interpretability
Zachary C. Lipton. 2018 · 2018
Earlier work this paper cites.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Nicolas Papernot and Patrick D. McDaniel. 2018 · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Earlier work this paper cites.
Please stop explaining black box models for high stakes decisions
Cynthia Rudin. 2018 · 2018
Earlier work this paper cites.
Interpreting neural networks with nearest neighbors
Eric Wallace, Shi Feng, and Jordan L. Boyd-Graber. 2018 · 2018
Earlier work this paper cites.
Mind the GAP: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. 2018 · 2018
Earlier work this paper cites.
What do RNN language models learn about filler-gap dependencies?
Ethan Wilcox, Roger Levy, Takashi Morita, and Richard Futrell. 2018 · 2018
Earlier work this paper cites.
Challenges of using text classifiers for causal inference
Zach Wood-Doughty, Ilya Shpitser, and Mark Dredze. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Earlier work this paper cites.
Interpretable neural predictions with differentiable binary variables
Jasmijn Bastings, Wilker Aziz, and Ivan Titov. 2019 · 2019
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James R. Glass. 2019 · 2019
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James R. Glass. 2019 · 2019
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna M. Wallach, Jennifer T. Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Cem Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019 · 2019
Earlier work this paper cites.
In search of meaning: Lessons, resources and next steps for computational analysis of financial discourse
Mahmoud El-Haj, Paul Rayson, Martin Walker, Steven Young, and Vasiliki Simaki. 2019 · 2019
Earlier work this paper cites.
Saliency learning: Teaching the model where to pay attention
Reza Ghaeini, Xiaoli Z. Fern, Hamed Shahbazi, and Prasad Tadepalli. 2019 · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Earlier work this paper cites.
Gamut: A design probe to understand how data scientists understand machine learning models
Fred Hohman, Andrew Head, Rich Caruana, Robert DeLine, and Steven Mark Drucker. 2019 · 2019
Earlier work this paper cites.
Attention is not explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Earlier work this paper cites.
Against interpretability: a critical examination of the interpretability problem in machine learning
Maya Krishnan. 2019 · 2019
Earlier work this paper cites.
On human predictions with explanations and predictions of machine learning models: A case study on deception detection
Vivian Lai and Chenhao Tan. 2019 · 2019
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2019 · 2019
Earlier work this paper cites.
An examination of the algorithmic accountability act of 2019
Mark MacCarthy. 2019 · 2019
Earlier work this paper cites.
Layer-wise relevance propagation: An overview
Grégoire Montavon, Alexander Binder, Sebastian Lapuschkin, Wojciech Samek, and Klaus-Robert Müller. 2019 · 2019
Earlier work this paper cites.
Definitions, methods, and applications in interpretable machine learning
W James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. 2019 · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model
Jesse Vig. 2019 · 2019
Earlier work this paper cites.
Allennlp interpret: A framework for explaining predictions of NLP models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019 · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019b · 2019
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Earlier work this paper cites.
Unsupervised domain clusters in pretrained language models
Roee Aharoni and Yoav Goldberg. 2020 · 2020
Earlier work this paper cites.
Evaluating saliency map explanations for convolutional neural networks: a user study
Ahmed Alqaraawi, Martin Schuessler, Philipp Weiß, Enrico Costanza, and Nadia Berthouze. 2020 · 2020
Earlier work this paper cites.
Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI
Alejandro Barredo Arrieta, Natalia Díaz Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-Lopez, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera. 2020 · 2020
Earlier work this paper cites.
Language (technology) is power: A critical survey of "bias" in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna M. Wallach. 2020 · 2020
Earlier work this paper cites.
Proxy tasks and subjective measures can be misleading in evaluating explainable AI systems
Zana Buçinca, Phoebe Lin, Krzysztof Z. Gajos, and Elena L. Glassman. 2020 · 2020
Earlier work this paper cites.
Make up your mind! adversarial generation of inconsistent natural language explanations
Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, and Phil Blunsom. 2020 · 2020
Earlier work this paper cites.
A survey of the state of explainable AI for natural language processing
Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen. 2020 · 2020
Earlier work this paper cites.
Analyzing individual neurons in pre-trained language models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov. 2020 · 2020
Earlier work this paper cites.
Explainable AI in industry: practical challenges and lessons learned: implications tutorial
Krishna Gade, Sahin Cem Geyik, Krishnaram Kenthapadi, Varun Mithal, and Ankur Taly. 2020 · 2020
Earlier work this paper cites.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, and et al. 2020 · 2020
Earlier work this paper cites.
Neuron shapley: Discovering the responsible neurons
Amirata Ghorbani and James Y. Zou. 2020 · 2020
Earlier work this paper cites.
Neural module networks for reasoning over text
Nitish Gupta, Kevin Lin, Dan Roth, Sameer Singh, and Matt Gardner. 2020 · 2020
Earlier work this paper cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Earlier work this paper cites.
Learning to faithfully rationalize by construction
Sarthak Jain, Sarah Wiegreffe, Yuval Pinter, and Byron C. Wallace. 2020 · 2020
Earlier work this paper cites.
Roles and utilization of attention heads in transformer-based neural language models
Jae-young Jo and Sung-Hyon Myaeng. 2020 · 2020
Earlier work this paper cites.
Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning
Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna M. Wallach, and Jennifer Wortman Vaughan. 2020 · 2020
Earlier work this paper cites.
Learning the difference that makes A difference with counterfactually-augmented data
Divyansh Kaushik, Eduard H. Hovy, and Zachary Chase Lipton. 2020 · 2020
Earlier work this paper cites.
NILE : Natural language inference with faithful natural language explanations
Sawan Kumar and Partha P. Talukdar. 2020 · 2020
Earlier work this paper cites.
Computational social science: Obstacles and opportunities
David MJ Lazer, Alex Pentland, Duncan J Watts, Sinan Aral, Susan Athey, Noshir Contractor, Deen Freelon, Sandra Gonzalez-Bailon, Gary King, Helen Margetts, et al. 2020 · 2020
Earlier work this paper cites.
Picking bert’s brain: Probing for linguistic dependencies in contextualized embeddings using representational similarity analysis
Michael A. Lepori and R. Thomas McCoy. 2020 · 2020
Earlier work this paper cites.
Asking without telling: Exploring latent ontologies in contextual representations
Julian Michael, Jan A. Botha, and Ian Tenney. 2020 · 2020
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP
John X. Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020 · 2020
Earlier work this paper cites.
Deep neural networks detect suicide risk from textual facebook posts
Yaakov Ophir, Refael Tikochinski, Christa S. C. Asterhan, Itay Sisso, and Roi Reichart. 2020 · 2020
Earlier work this paper cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Earlier work this paper cites.
On the systematicity of probing contextualized word representations: The case of hypernymy in BERT
Abhilasha Ravichander, Eduard H. Hovy, Kaheer Suleman, Adam Trischler, and Jackie Chi Kit Cheung. 2020 · 2020
Earlier work this paper cites.
Adjusting for confounding with text matching
Margaret E. Roberts, Brandon M Stewart, and Richard A. Nielsen. 2020 · 2020
Cited alongside, same era.
Explainable machine learning for scientific insights and discoveries
Ribana Roscher, Bastian Bohn, Marco F. Duarte, and Jochen Garcke. 2020 · 2020
Cited alongside, same era.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber. 2020 · 2020
Cited alongside, same era.
Multi-simlex: A large-scale evaluation of multilingual and crosslingual lexical semantic similarity
Ivan Vulic, Simon Baker, Edoardo Maria Ponti, Ulla Petti, Ira Leviant, Kelly Wing, Olga Majewska, Eden Bar, Matt Malone, Thierry Poibeau, Roi Reichart, and Anna Korhonen. 2020 · 2020
Cited alongside, same era.
Perturbed masking: Parameter-free probing for analyzing and interpreting BERT
Zhiyong Wu, Yun Chen, Ben Kao, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2023 · 2023
Later among the works it cites.
Analyzing transformers in embedding space
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2023 · 2023
Later among the works it cites.
Tools such as chatgpt threaten transparent science; here are our ground rules for their use
Nature Editorials. 2023 · 2023
Later among the works it cites.
Gpts are gpts: An early look at the labor market impact potential of large language models
Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. 2023 · 2023
Later among the works it cites.
Sequential integrated gradients: a simple but effective method for explaining language models
Joseph Enguehard. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Greedy attack and gumbel attack: Generating adversarial examples for discrete data
Puyudi Yang, Jianbo Chen, Cho-Jui Hsieh, Jane-Ling Wang, and Michael I. Jordan. 2020 · 2020
Cited alongside, same era.
On completeness-aware concept-based explanations in deep neural networks
Chih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li, Tomas Pfister, and Pradeep Ravikumar. 2020 · 2020
Cited alongside, same era.
Credit risk evaluation model with textual features from loan descriptions for P2P lending
Weiguo Zhang, Chao Wang, Yue Zhang, and Junbo Wang. 2020 · 2020
Cited alongside, same era.
Masking as an efficient alternative to finetuning for pretrained language models
Mengjie Zhao, Tao Lin, Fei Mi, Martin Jaggi, and Hinrich Schütze. 2020 · 2020
Cited alongside, same era.
Transformer interpretability beyond attention visualization
Hila Chefer, Shir Gur, and Lior Wolf. 2021 · 2021
Cited alongside, same era.
An empirical survey of data augmentation for limited data learning in nlp
Jiaao Chen, Derek Tam, Colin Raffel, Mohit Bansal, and Diyi Yang. 2021 · 2021
Cited alongside, same era.
Are neural nets modular? inspecting functional modularity through differentiable weight masks
Róbert Csordás, Sjoerd van Steenkiste, and Jürgen Schmidhuber. 2021 · 2021
Cited alongside, same era.
It’s not only what you say, it’s also who it’s said to: Counterfactual analysis of interactive behavior in the courtroom
Biaoyan Fang, Trevor Cohn, Timothy Baldwin, and Lea Frermann. 2023 · 2023
Later among the works it cites.
Causal-structure driven augmentations for text OOD generalization
Amir Feder, Yoav Wald, Claudia Shi, Suchi Saria, and David M. Blei. 2023 · 2023
Later among the works it cites.
Counterfactuals of counterfactuals: a back-translation-inspired approach to analyse counterfactual editors
George Filandrianos, Edmund Dervakos, Orfeas Menis-Mastromichalakis, Chrysoula Zerva, and Giorgos Stamou. 2023 · 2023
Later among the works it cites.
Deepdecipher: Accessing and investigating neuron activation in large language models
Albert Garde, Esben Kran, and Fazl Barez. 2023 · 2023
Later among the works it cites.
Faithful explanations of black-box nlp models using llm-generated counterfactuals
Yair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin, Amit Sharma, and Roi Reichart. 2023 · 2023
Later among the works it cites.
A survey of adversarial defenses and robustness in NLP
Shreya Goyal, Sumanth Doddapaneni, Mitesh M. Khapra, and Balaraman Ravindran. 2023 · 2023
Later among the works it cites.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023 · 2023
Later among the works it cites.
A robust information-masking approach for domain counterfactual generation
Pengfei Hong, Rishabh Bhardwaj, Navonil Majumder, Somak Aditya, and Soujanya Poria. 2023 · 2023
Later among the works it cites.
Chatgpt for good? on opportunities and challenges of large language models for education
Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, and et al. 2023 · 2023
Later among the works it cites.
VISIT: visualizing and interpreting the semantic information flow of transformers
Shahar Katz and Yonatan Belinkov. 2023 · 2023
Later among the works it cites.
The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning
Seungone Kim, Se June Joo, Doyoung Kim, Joel Jang, Seonghyeon Ye, Jamin Shin, and Minjoon Seo. 2023 · 2023
Later among the works it cites.
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, and et al. 2023 · 2023
Later among the works it cites.
Towards trustworthy and aligned machine learning: A data-centric survey with causality perspectives
Haoyang Liu, Maheep Chaudhary, and Haohan Wang. 2023 · 2023
Later among the works it cites.
Interpretable-by-design text classification with iteratively generated concept bottleneck
Josh Magnus Ludan, Qing Lyu, Yue Yang, Liam Dugan, Mark Yatskar, and Chris Callison-Burch. 2023 · 2023
Later among the works it cites.
Faithful chain-of-thought reasoning
Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Marianna Apidianaki, and Chris Callison-Burch. 2023 · 2023
Later among the works it cites.
What makes chain-of-thought prompting effective? A counterfactual study
Aman Madaan, Katherine Hermann, and Amir Yazdanbakhsh. 2023 · 2023
Later among the works it cites.
Post-hoc interpretability for neural NLP: A survey
Andreas Madsen, Siva Reddy, and Sarath Chandar. 2023 · 2023
Later among the works it cites.
Inverse scaling: When bigger isn’t better
Ian R. McKenzie, Alexander Lyzhov, Michael Pieler, Alicia Parrish, Aaron Mueller, and et al. 2023 · 2023
Later among the works it cites.
MaNtLE: Model-agnostic natural language explainer
Rakesh Menon, Kerem Zaman, and Shashank Srivastava. 2023 · 2023
Later among the works it cites.
Can llms facilitate interpretation of pre-trained language models?
Basel Mousi, Nadir Durrani, and Fahim Dalvi. 2023 · 2023
Later among the works it cites.
Future lens: Anticipating subsequent tokens from a single hidden state
Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C. Wallace, and David Bau. 2023 · 2023
Later among the works it cites.
On measuring faithfulness of natural language explanations
Letitia Parcalabescu and Anette Frank. 2023 · 2023
Later among the works it cites.
Increasing our ads transparency
Pedro Pavón. 2023 · 2023
Later among the works it cites.
Toward transparent AI: A survey on interpreting the inner structures of deep neural networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. 2023 · 2023
Later among the works it cites.
Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities
Waddah Saeed and Christian W. Omlin. 2023 · 2023
Later among the works it cites.
Mansi Sakarvadia, Arham Khan, Aswathy Ajith, Daniel Grzenda, Nathaniel Hudson, André Bauer, Kyle Chard, and Ian T. Foster. 2023 · 2023
Later among the works it cites.
Response to the march 2023 ’pause giant ai experiments: An open letter’ by yoshua bengio, signed by stuart russell, elon musk, steve wozniak, yuval noah harari and others…
Jim Samuel. 2023 · 2023
Later among the works it cites.
People make better edits: Measuring the efficacy of llm-generated counterfactually augmented data for harmful language detection
Indira Sen, Dennis Assenmacher, Mattia Samory, Isabelle Augenstein, Wil M. P. van der Aalst, and Claudia Wagner. 2023 · 2023
Later among the works it cites.
Eilam Shapira, Reut Apel, Moshe Tennenholtz, and Roi Reichart. 2023 · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023 · 2023
Later among the works it cites.
Interpretable by design: Wrapper boxes combine neural performance with faithful explanations
Yiheng Su, Junyi Jessy Li, and Matthew Lease. 2023 · 2023
Later among the works it cites.
Attribution patching outperforms automated circuit discovery
Aaquib Syed, Can Rager, and Arthur Conmy. 2023 · 2023
Later among the works it cites.
Perspective changes in human listeners are aligned with the contextual transformation of the word embedding space
Refael Tikochinski, Ariel Goldstein, Yaara Yeshurun, Uri Hasson, and Roi Reichart. 2023 · 2023
Later among the works it cites.
CREST: A joint framework for rationalization and counterfactual text generation
Marcos V. Treviso, Alexis Ross, Nuno Miguel Guerreiro, and André F. T. Martins. 2023 · 2023
Later among the works it cites.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023 · 2023
Later among the works it cites.
Navigating cultural chasms: Exploring and unlocking the cultural POV of text-to-image models
Mor Ventura, Eyal Ben-David, Anna Korhonen, and Roi Reichart. 2023 · 2023
Later among the works it cites.
Artificial intelligence in studies—use of chatgpt and ai-based tools among students in germany
Jörg von Garrel and Jana Mayer. 2023 · 2023
Later among the works it cites.
A causal view of entity bias in (large) language models
Fei Wang, Wenjie Mo, Yiwei Wang, Wenxuan Zhou, and Muhao Chen. 2023a · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023b · 2023
Later among the works it cites.
Goal-driven explainable clustering via language descriptions
Zihan Wang, Jingbo Shang, and Ruiqi Zhong. 2023c · 2023
Later among the works it cites.
Causal proxy models for concept-based model explanations
Zhengxuan Wu, Karel D’Oosterlinck, Atticus Geiger, Amir Zur, and Christopher Potts. 2023a · 2023
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in alpaca
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah D. Goodman. 2023b · 2023
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi. 2023 · 2023
Later among the works it cites.
Characterizing mechanisms for factual recall in language models
Qinan Yu, Jack Merullo, and Ellie Pavlick. 2023 · 2023
Later among the works it cites.
Causal matching with text embeddings: A case study in estimating the causal effects of peer review policies
Raymond Zhang, Neha Nayak Kennard, Daniel Scott Smith, Daniel A. McFarland, Andrew McCallum, and Katherine Keith. 2023 · 2023
Later among the works it cites.
On evaluating adversarial robustness of large vision-language models
Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Cheung, and Min Lin. 2023b · 2023
Later among the works it cites.
An invariant learning characterization of controlled text generation
Carolina Zheng, Claudia Shi, Keyon Vafa, Amir Feder, and David M. Blei. 2023 · 2023
Later among the works it cites.
Causal inference from text: Unveiling interactions between variables
Yuxiang Zhou and Yulan He. 2023 · 2023
Later among the works it cites.
Autodan: Automatic and interpretable adversarial attacks on large language models
Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu, Joe Barrow, Zichao Wang, Furong Huang, Ani Nenkova, and Tong Sun. 2023 · 2023
Later among the works it cites.
Llms with chain-of-thought are non-causal reasoners
Guangsheng Bao, Hongbo Zhang, Linyi Yang, Cunxiang Wang, and Yue Zhang. 2024 · 2024
Closest in time.
Mechanistic interpretability for AI safety - A review
Leonard Bereska and Efstratios Gavves. 2024 · 2024
Closest in time.
Mitigating text toxicity with counterfactual generation
Milan Bhan, Jean-Noel Vittaut, Nina Achache, Victor Legrand, Nicolas Chesneau, Annabelle Blangero, Juliette Murris, and Marie-Jeanne Lesot. 2024 · 2024
Closest in time.
Measuring the robustness of nlp models to domain shifts
Nitay Calderon, Naveh Porat, Eyal Ben-David, Alexander Chapanin, Zorik Gekhman, Nadav Oved, Vitaly Shalumov, and Roi Reichart. 2024 · 2024
Closest in time.
Yu Ying Chiu, Liwei Jiang, Maria Antoniak, Chan Young Park, Shuyue Stella Li, Mehar Bhatia, Sahithya Ravi, Yulia Tsvetkov, Vered Shwartz, and Yejin Choi. 2024 · 2024
Closest in time.
Estimating the causal effect of early arxiving on paper acceptance
Yanai Elazar, Jiayao Zhang, David Wadden, Bo Zhang, and Noah A. Smith. 2024 · 2024
Closest in time.
Does fine-tuning llms on new knowledge encourage hallucinations?
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig. 2024 · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva. 2024 · 2024
Closest in time.
Universal neurons in GPT2 language models
Wes Gurnee, Theo Horsley, Zifan Carl Guo, Tara Rezaei Kheirkhah, Qinyi Sun, Will Hathaway, Neel Nanda, and Dimitris Bertsimas. 2024 · 2024
Closest in time.
Scaling up discovery of latent concepts in deep NLP models
Majd Hawasly, Fahim Dalvi, and Nadir Durrani. 2024 · 2024
Closest in time.
STILE: exploring and debugging social biases in pre-trained text representations
Samia Kabir, Lixiang Li, and Tianyi Zhang. 2024 · 2024
Closest in time.
Atp*: An efficient and scalable method for localizing LLM behaviour to components
János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda. 2024 · 2024
Closest in time.
A systematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, and et al. 2024 · 2024
Closest in time.
Prompting large language models for counterfactual generation: An empirical study
Yongqi Li, Mayi Xu, Xin Miao, Shen Zhou, and Tieyun Qian. 2024 · 2024
Closest in time.
Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel A. McFarland, and James Y. Zou. 2024 · 2024
Closest in time.
Local interpretations for explainable natural language processing: A survey
Siwen Luo, Hamish Ivison, Soyeon Caren Han, and Josiah Poon. 2024 · 2024
Closest in time.
From insights to actions: The impact of interpretability and analysis research on NLP
Marius Mosbach, Vagrant Gautam, Tomás Vergara Browne, Dietrich Klakow, and Mor Geva. 2024 · 2024
Closest in time.
Romy Müller. 2024 · 2024
Closest in time.
Llms for generating and evaluating counterfactuals: A comprehensive study
Van Bach Nguyen, Paul Youssef, Jörg Schlötterer, and Christin Seifert. 2024 · 2024
Closest in time.
Advprompter: Fast adaptive adversarial prompting for llms
Anselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos, and Yuandong Tian. 2024 · 2024
Closest in time.
NORMAD: A benchmark for measuring the cultural adaptability of large language models
Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. 2024 · 2024
Closest in time.
Revisiting character-level adversarial attacks
Elias Abad Rocamora, Yongtao Wu, Fanghui Liu, Grigorios G. Chrysos, and Volkan Cevher. 2024 · 2024
Closest in time.
Catfood: Counterfactual augmented training for improving out-of-domain performance and calibration
Rachneet Sachdeva, Martin Tutek, and Iryna Gurevych. 2024 · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob N. Foerster, Tim Rocktäschel, and Roberta Raileanu. 2024 · 2024
Closest in time.
Towards explainability and fairness in swiss judgement prediction: Benchmarking on a multilingual dataset
T. Y. S. S. Santosh, Nina Baumgartner, Matthias Stürmer, Matthias Grabmair, and Joel Niklaus. 2024 · 2024
Closest in time.
Mario Sanz-Guerrero and Javier Arroyo. 2024 · 2024
Closest in time.
Can large language models replace economic choice prediction labs?
Eilam Shapira, Omer Madmon, Roi Reichart, and Moshe Tennenholtz. 2024 · 2024
Closest in time.
Rethinking interpretability in the era of large language models
Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao. 2024 · 2024
Closest in time.
Atomic inference for NLI with generated facts as atoms
Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Oana-Maria Camburu, and Marek Rei. 2024 · 2024
Closest in time.
Interpreting pretrained language models via concept bottlenecks
Zhen Tan, Lu Cheng, Song Wang, Bo Yuan, Jundong Li, and Huan Liu. 2024 · 2024
Closest in time.
Systematic biases in LLM simulations of debates
Amir Taubenfeld, Yaniv Dover, Roi Reichart, and Ariel Goldstein. 2024 · 2024
Closest in time.
Incremental accumulation of linguistic context in artificial and biological neural networks
Refael Tikochinski, Ariel Goldstein, Yoav Meiri, Uri Hasson, and Roi Reichart. 2024 · 2024
Closest in time.
Diffusion lens: Interpreting text encoders in text-to-image pipelines
Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad, and Yonatan Belinkov. 2024 · 2024
Closest in time.
Towards conversational diagnostic AI
Tao Tu, Anil Palepu, Mike Schaekermann, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Nenad Tomasev, Shekoofeh Azizi, Karan Singhal, Yong Cheng, Le Hou, Albert Webson, Kavita Kulkarni, S. Sara Mahdavi, Christopher Semturs, Juraj Gottweis, Joelle K. Barral, Katherine Chou, Gregory S. Corrado, Yossi Matias, Alan Karthikesalingam, and Vivek Natarajan. 2024 · 2024
Closest in time.
Usable XAI: 10 strategies towards exploiting explainability in the LLM era
Xuansheng Wu, Haiyan Zhao, Yaochen Zhu, Yucheng Shi, Fan Yang, Tianming Liu, Xiaoming Zhai, Wenlin Yao, Jundong Li, Mengnan Du, and Ninghao Liu. 2024 · 2024
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Shaochen Zhong, Bing Yin, and Xia Ben Hu. 2024 · 2024
Closest in time.
Multi-round counterfactual generation: Interpreting and improving models of text classification
Huajie Zhang, Yuxin Ying, Fuzhen Zhuang, Haiqin Weng, Sun Ying, Zhao Zhang, Yiqi Tong, and Yan Liu. 2024b · 2024
Closest in time.
Explainability for large language models: A survey
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024 · 2024
Closest in time.
Relying on the unreliable: The impact of language models’ reluctance to express uncertainty
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, and Maarten Sap. 2024 · 2024
Closest in time.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024 · 2024
Closest in time.
Nitay Calderon, Roi Reichart, and Rotem Dror. 2025 · 2025
Closest in time.
Open problems in mechanistic interpretability
Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, Stella Biderman, Adrià Garriga-Alonso, Arthur Conmy, Neel Nanda, Jessica Rumbelow, Martin Wattenberg, Nandi Schoots, Joseph Miller, Eric J. Michaud, Stephen Casper, Max Tegmark, William Saunders, David Bau, Eric Todd, Atticus Geiger, Mor Geva, Jesse Hoogland, Daniel Murfet, and Thomas McGrath. 2025 · 2025
Closest in time.
Probing cross-lingual lexical knowledge from multilingual sentence encoders
Ivan Vulic, Goran Glavas, Fangyu Liu, Nigel Collier, Edoardo Maria Ponti, and Anna Korhonen. 2023 · 2097
Closest in time.