Fetching the paper…
Reading the bibliography…
In order to oversee advanced AI systems, it is important to understand their underlying decision-making process.
Snowball: A language for stemming algorithms
Martin F Porter. 2001 · 2001
Earlier work this paper cites.
Interpretability and explainability: A machine learning zoo mini-tour
Ricards Marcinkevics and Julia E. Vogt. 2020 · 2012
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian J. Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Earlier work this paper cites.
e-SNLI: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2018 · 2018
Earlier work this paper cites.
Measuring association between labels and free-text rationales
Sarah Wiegreffe, Ana Marasović, and Noah A. Smith. 2020 · 2018
Earlier work this paper cites.
Eraser: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen Rajani, Eric P. Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2019 · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
On interpretability of artificial neural networks: A survey
Fenglei Fan, Jinjun Xiong, Mengzhou Li, and Ge Wang. 2020 · 2020
Earlier work this paper cites.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Earlier work this paper cites.
SemEval-2020 task 4: Commonsense validation and explanation
Cunxiang Wang, Shuailong Liang, Yili Jin, Yilong Wang, Xiaodan Zhu, and Yue Zhang. 2020 · 2020
Earlier work this paper cites.
Explanations for CommonsenseQA: New Dataset and Models
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021 · 2021
Cited alongside, same era.
The struggles of feature-based explanations: Shapley values vs. minimal sufficient subsets
Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom. 2021 · 2021
Cited alongside, same era.
Teach me to explain: A review of datasets for explainable natural language processing
Sarah Wiegreffe and Ana Marasović. 2021 · 2021
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. 2022 · 2022
Cited alongside, same era.
Faithful reasoning using large language models
Antonia Creswell and Murray Shanahan. 2022 · 2022
Cited alongside, same era.
Interpretable by design: Learning predictors by composing interpretable queries
Aditya Chattopadhyay, Stewart Slocum, Benjamin D. Haeffele, René Vidal, and Donald Geman. 2023 · 2023
Later among the works it cites.
Challenges with unsupervised llm knowledge discovery
Sebastian Farquhar, Vikrant Varma, Zachary Kenton, Johannes Gasteiger, Vladimir Mikulik, and Rohin Shah. 2023 · 2023
Later among the works it cites.
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Thomas Henighan, Timothy Maxwell, Timothy Telleen-Lawton, Tristan Hume, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez. 2023 · 2023
Later among the works it cites.
The alignment problem from a deep learning perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Selection-inference: Exploiting large language models for interpretable logical reasoning
Antonia Creswell, Murray Shanahan, and Irina Higgins. 2022 · 2022
Cited alongside, same era.
Explaining chest x-ray pathologies in natural language
Maxime Kayser, Cornelius Emde, Oana-Maria Camburu, Guy Parsons, Bartlomiej Papiez, and Thomas Lukasiewicz. 2022 · 2022
Cited alongside, same era.
Externalized reasoning oversight: a research direction for language model alignment
Tamera Lanham. 2022 · 2022
Cited alongside, same era.
Huspacy: an industrial-strength hungarian natural language processing toolkit
György Orosz, Zsolt Szántó, Péter Berkecz, Gergő Szabó, and Richárd Farkas. 2022 · 2022
Cited alongside, same era.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. 2022 · 2022
Cited alongside, same era.
Faithfulness tests for natural language explanations
Pepa Atanasova, Oana-Maria Camburu, Christina Lioma, Thomas Lukasiewicz, Jakob Grue Simonsen, and Isabelle Augenstein. 2023 · 2023
Cited alongside, same era.
On measuring faithfulness or self-consistency of natural language explanations
Letitia Parcalabescu and Anette Frank. 2023 · 2023
Later among the works it cites.
Question decomposition improves the faithfulness of model-generated reasoning
Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Carson Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Sam McCandlish, Sheer El Showk, Tamera Lanham, Tim Maxwell, Venkatesa Chandrasekaran, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez. 2023 · 2023
Later among the works it cites.
Preventing language models from hiding their reasoning
Fabien Roger and Ryan Greenblatt. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Miles Turpin, Julian Michael, Ethan Perez, and Sam Bowman. 2023 · 2023
Later among the works it cites.
Honesty is the best policy: Defining and mitigating ai deception
Francis Rhys Ward, Francesco Belardinelli, Francesca Toni, and Tom Everitt. 2023 · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023 · 2023
Later among the works it cites.