Fetching the paper…
Reading the bibliography…
With the growing popularity of general-purpose Large Language Models (LLMs), comes a need for more global explanations of model behaviors.
An evaluation of the human-interpretability of explanation
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez. 2019 · 1902
Earlier work this paper cites.
Discovery of natural language concepts in individual units of cnns
Seil Na, Yo Joong Choe, Dong-Hyun Lee, and Gunhee Kim. 2019 · 1902
Earlier work this paper cites.
Explaining classifiers with causal concept effect (cace)
Yash Goyal, Amir Feder, Uri Shalit, and Been Kim. 2019 · 1907
Earlier work this paper cites.
Theory and evaluation metrics for learning disentangled representations
Kien Do and Truyen Tran. 2019 · 1908
Earlier work this paper cites.
A new measure of rank correlation
Maurice G Kendall. 1938 · 1938
Earlier work this paper cites.
Coefficient alpha and the internal structure of tests
Lee J Cronbach. 1951 · 1951
Earlier work this paper cites.
Construct validity in psychological tests
Lee J Cronbach and Paul E Meehl. 1955 · 1955
Earlier work this paper cites.
Convergent and discriminant validation by the multitrait-multimethod matrix
Donald T Campbell and Donald W Fiske. 1959 · 1959
Earlier work this paper cites.
The proof and measurement of association between two things
C. Spearman. 1961 · 1961
Earlier work this paper cites.
Psychometric theory new york
Jum C Nunnally and Ira H Bernstein. 1994 · 1994
Earlier work this paper cites.
Faithfulness and reduplicative identity
John J McCarthy and Alan Prince. 1995 · 1995
Earlier work this paper cites.
Introduction to measurement theory
Mary J Allen and Wendy M Yen. 2001 · 2001
Earlier work this paper cites.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2004
Earlier work this paper cites.
Abstracting deep neural networks into concept graphs for concept level interpretability
Avinash Kori, Parth Natekar, Ganapathy Krishnamurthi, and Balaji Srinivasan. 2020 · 2008
Earlier work this paper cites.
External evaluation of topic models
David Newman, Sarvnaz Karimi, and Lawrence Cavedon. 2009 · 2009
Earlier work this paper cites.
Intrinsic probing through dimension selection
Lucas Torroba Hennigen, Adina Williams, and Ryan Cotterell. 2020 · 2010
Earlier work this paper cites.
Automatic evaluation of topic coherence
David Newman, Jey Han Lau, Karl Grieser, and Timothy Baldwin. 2010 · 2010
Earlier work this paper cites.
Optimizing semantic coherence in topic models
David Mimno, Hanna Wallach, Edmund Talley, Miriam Leenders, and Andrew McCallum. 2011 · 2011
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
David Alvarez Melis and Tommi Jaakkola. 2018 · 2018
Cited alongside, same era.
Metrics for explainable ai: Challenges and prospects
Robert R Hoffman, Shane T Mueller, Gary Klein, and Jordan Litman. 2018 · 2018
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. 2018 · 2018
Cited alongside, same era.
Co-attentive multi-task learning for explainable recommendation
Zhongxia Chen, Xiting Wang, Xing Xie, Tong Wu, Guoqing Bu, Yining Wang, and Enhong Chen. 2019b · 2019
Cited alongside, same era.
Explainable recommendation through attentive multi-view learning
Jingyue Gao, Xiting Wang, Yasha Wang, and Xing Xie. 2019 · 2019
Cited alongside, same era.
Self-explaining deep models with logic rule reasoning
Seungeon Lee, Xiting Wang, Sungwon Han, Xiaoyuan Yi, Xing Xie, and Meeyoung Cha. 2022 · 2022
Later among the works it cites.
A unified understanding of deep nlp models for text classification
Zhen Li, Xiting Wang, Weikai Yang, Jing Wu, Zhengyan Zhang, Zhiyuan Liu, Maosong Sun, Hui Zhang, and Shixia Liu. 2022 · 2022
Later among the works it cites.
Analyzing encoded concepts in transformer language models
Hassan Sajjad, Nadir Durrani, Fahim Dalvi, Firoj Alam, Abdul Rafae Khan, and Jia Xu. 2022 · 2022
Later among the works it cites.
A framework for learning ante-hoc explainable models via concepts
Anirban Sarkar, Deepak Vijaykeerthy, Anindya Sarkar, and Vineeth N Balasubramanian. 2022 · 2022
Later among the works it cites.
Reinforcement subgraph reasoning for fake news detection
Ruichao Yang, Xiting Wang, Yiqiao Jin, Chaozhuo Li, Jianxun Lian, and Xing Xie. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Y Zou, and Been Kim. 2019 · 2019
Cited alongside, same era.
Towards a deep and unified understanding of deep neural models in nlp
Chaoyu Guan, Xiting Wang, Quanshi Zhang, Runjin Chen, Di He, and Xing Xie. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
Cxplain: Causal explanations for model interpretation under uncertainty
Patrick Schwab and Walter Karlen. 2019 · 2019
Cited alongside, same era.
Concept whitening for interpretable image recognition
Zhi Chen, Yijie Bei, and Cynthia Rudin. 2020 · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020 · 2020
Cited alongside, same era.
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Cited alongside, same era.
Rishabh Bhardwaj and Soujanya Poria. 2023 · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023 · 2023
Later among the works it cites.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023 · 2023
Later among the works it cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah. 2023 · 2023
Later among the works it cites.
“are your explanations reliable?” investigating the stability of lime in explaining text classifiers by marrying xai and adversarial attack
Christopher Burger, Lingwei Chen, and Thai Le. 2023 · 2023
Later among the works it cites.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2023 · 2023
Later among the works it cites.
From neural activations to concepts: A survey on explaining concepts in neural networks
Jae Hee Lee, Sergio Lanza, and Stefan Wermter. 2023 · 2023
Later among the works it cites.
Loogle: Can long-context language models understand long contexts?
Jiaqi Li, Mengmeng Wang, Zilong Zheng, and Muhan Zhang. 2023 · 2023
Later among the works it cites.
Evaluating the stability of semantic concept representations in cnns for robust explainability
Georgii Mikriukov, Gesina Schwalbe, Christian Hellert, and Korinna Bade. 2023 · 2023
Later among the works it cites.
Explaining black box text modules in natural language with language models
Chandan Singh, Aliyah R Hsu, Richard Antonello, Shailee Jain, Alexander G Huth, Bin Yu, and Jianfeng Gao. 2023 · 2023
Later among the works it cites.
Understanding and enhancing robustness of concept-based models
Sanchit Sinha, Mengdi Huai, Jianhui Sun, and Aidong Zhang. 2023 · 2023
Later among the works it cites.
Explain any concept: Segment anything meets concept-based explanation
Ao Sun, Pingchuan Ma, Yuanyuan Yuan, and Shuai Wang. 2023 · 2023
Later among the works it cites.
Evaluating evaluation metrics: A framework for analyzing nlg evaluation metrics using measurement theory
Ziang Xiao, Susu Zhang, Vivian Lai, and Q Vera Liao. 2023 · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. 2023 · 2023
Later among the works it cites.
Uncovering safety risks in open-source llms through concept activation vector
Zhihao Xu, Ruixuan Huang, Xiting Wang, Fangzhao Wu, Jing Yao, and Xing Xie. 2024 · 2024
Closest in time.
Foundation models meet visualizations: Challenges and opportunities
Weikai Yang, Mengchen Liu, Zheng Wang, and Shixia Liu. 2024 · 2024
Closest in time.
Distillation with explanations from large language models
Hanyu Zhang, Xiting Wang, Xiang Ao, and Qing He. 2024 · 2024
Closest in time.
Multi-level recommendation reasoning over knowledge graphs with reinforcement learning
Xiting Wang, Kunpeng Liu, Dongjie Wang, Le Wu, Yanjie Fu, and Xing Xie. 2022 · 2098
Closest in time.