Fetching the paper…
Reading the bibliography…
In this survey, we address the key challenges in Large Language Models (LLM) research, focusing on the importance of interpretability.
Conducting systematic literature reviews and systematic mapping studies
Balbir Barn, Souvik Barat, and Tony Clark · 2017
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model
Jesse Vig · 2019
Earlier work this paper cites.
exdil: A tool for classifying and explaining hospital discharge letters
Fabio Mercorio, Mario Mezzanzanica, and Andrea Seveso · 2020
Earlier work this paper cites.
Financial sentiment analysis: an investigation into common mistakes and silver bullets
Frank Xing, Lorenzo Malandri, Yue Zhang, and Erik Cambria · 2020
Earlier work this paper cites.
Changes in evidence for studies assessing interventions for covid-19 reported in preprints: meta-research study
Theodora Oikonomidi, Isabelle Boutron, Olivier Pierre, Guillaume Cabanac, Philippe Ravaud, and Covid-19 Nma Consortium · 2020
Earlier work this paper cites.
Explainable ai: a narrative review at the crossroad of knowledge discovery, knowledge representation and representation learning
Ikram Chraibi Kaadoud, Lina Fahed, and Philippe Lenca · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Earlier work this paper cites.
A survey on the explainability of supervised machine learning
Nadia Burkart and Marco F Huber · 2021
Earlier work this paper cites.
ContrXT: Generating contrastive explanations from any text classifier
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani, and Andrea Seveso · 2021
Earlier work this paper cites.
Interpreting language models through knowledge graph extraction
Vinitra Swamy, Angelika Romanou, and Martin Jaggi · 2021
Earlier work this paper cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
T Wu, M Tulio Ribeiro, J Heer, and D Weld · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Beware the rationalization trap! when language model explainability diverges from our mental models of language, 2022
Rita Sevastjanova and Mennatallah El-Assady · 2022
Earlier work this paper cites.
Xai for myo-controlled prosthesis: Explaining emg data for hand gesture classification
Noemi Gozzi, Lorenzo Malandri, Fabio Mercorio, and Alessandra Pedrocchi · 2022
Earlier work this paper cites.
A survey on xai for cyber physical systems in medicine
Nicola Alimonda, Luca Guidotto, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, and Giovanni Tosi · 2022
Earlier work this paper cites.
Contrastive explanations of text classifiers as a service
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani, and Andrea Seveso · 2022
Earlier work this paper cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Visual classification via description from large language models
Sachit Menon and Carl Vondrick · 2022
Earlier work this paper cites.
Rethinking with retrieval: Faithful large language model inference
Hangfeng He, Hongming Zhang, and Dan Roth · 2022
Earlier work this paper cites.
Improved logical reasoning of language models via differentiable symbolic programming
Hanlin Zhang, Ziyang Li, Jiani Huang, Mayur Naik, and Eric Xing · 2022
Cited alongside, same era.
Explanations from large language models make small reasoners better
Shiyang Li, Jianshu Chen, Yelong Shen, Zhiyu Chen, Xinlu Zhang, Zekun Li, Hong Wang, Jing Qian, Baolin Peng, Yi Mao, et al · 2022
Cited alongside, same era.
The unreliability of explanations in few-shot prompting for textual reasoning
Xi Ye and Greg Durrett · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan · 2022
Cited alongside, same era.
Roscoe: A suite of metrics for scoring step-by-step reasoning
Olga Golovneva, Moya Peng Chen, Spencer Poff, Martin Corredor, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz · 2022
Cited alongside, same era.
Answering questions by meta-reasoning over multiple chains of thought
Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, and Jonathan Berant · 2023
Later among the works it cites.
Inseq: An interpretability toolkit for sequence generation models
Gabriele Sarti, Nils Feldhus, Ludwig Sickert, and Oskar van der Wal · 2023
Later among the works it cites.
Towards understanding in-context learning with contrastive demonstrations and saliency maps
Zongxia Li, Paiheng Xu, Fuxiao Liu, and Hyemi Song · 2023
Later among the works it cites.
Lmexplainer: a knowledge-enhanced explainer for language models
Zichen Chen, Ambuj K Singh, and Misha Sra · 2023
Later among the works it cites.
Explaining black box text modules in natural language with language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can ChatGPT’s responses boost traditional natural language processing?
Mostafa Amin, Erik Cambria, and Björn Schuller · 2023
Cited alongside, same era.
Large language models in medicine
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting · 2023
Cited alongside, same era.
The importance of human-labeled data in the era of llms
Yang Liu · 2023
Cited alongside, same era.
Leveraging group contrastive explanations for handling fairness
Alessandro Castelnovo, Nicole Inverardi, Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, and Andrea Seveso · 2023
Cited alongside, same era.
A comprehensive review on financial explainable ai
Wei Jie Yeo, Wihan van der Heever, Rui Mao, Erik Cambria, Ranjan Satapathy, and Gianmarco Mengaldo · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Cited alongside, same era.
Large language models in medical education: Opportunities, challenges, and future directions
Alaa Abd-Alrazaq, Rawan AlSaad, Dari Alhuwail, Arfan Ahmed, Padraig Mark Healy, Syed Latifi, Sarah Aziz, Rafat Damseh, Sadam Alabed Alrazak, Javaid Sheikh, et al · 2023
Cited alongside, same era.
Chandan Singh, Aliyah R Hsu, Richard Antonello, Shailee Jain, Alexander G Huth, Bin Yu, and Jianfeng Gao · 2023
Later among the works it cites.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman · 2023
Later among the works it cites.
Explainable automated debugging via large language model-driven scientific debugging
Sungmin Kang, Bei Chen, Shin Yoo, and Jian-Guang Lou · 2023
Later among the works it cites.
Post hoc explanations of language models can improve language models
Satyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun, Sameer Singh, and Himabindu Lakkaraju · 2023
Later among the works it cites.
Language in a bottle: Language model guided concept bottlenecks for interpretable image classification
Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, and Mark Yatskar · 2023
Later among the works it cites.
Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional images
Nitzan Bitton-Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt, Yuval Elovici, Gabriel Stanovsky, and Roy Schwartz · 2023
Later among the works it cites.
Chatgraph: Interpretable text classification by converting chatgpt knowledge to graphs
Yucheng Shi, Hehuan Ma, Wenliang Zhong, Gengchen Mai, Xiang Li, Tianming Liu, and Junzhou Huang · 2023
Later among the works it cites.
Techs: Temporal logical graph networks for explainable extrapolation reasoning
Qika Lin, Jun Liu, Rui Mao, Fangzhi Xu, and Erik Cambria · 2023
Later among the works it cites.
Amirhossein Aminimehr, Pouya Khani, Amirali Molaei, Amirmohammad Kazemeini, and Erik Cambria · 2023
Later among the works it cites.
Eight things to know about large language models
Samuel R Bowman · 2023
Later among the works it cites.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2023
Later among the works it cites.
Ai transparency in the age of llms: A human-centered research roadmap
Q Vera Liao and Jennifer Wortman Vaughan · 2023
Later among the works it cites.
Qianqian Xie, Weiguang Han, Yanzhao Lai, Min Peng, and Jimin Huang · 2023
Later among the works it cites.
Generative ai in eu law: Liability, privacy, intellectual property, and cybersecurity
Claudio Novelli, Federico Casolari, Philipp Hacker, Giorgio Spedicato, and Luciano Floridi · 2024
Closest in time.
Model-contrastive explanations through symbolic reasoning
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, and Andrea Seveso · 2024
Closest in time.