Fetching the paper…
Reading the bibliography…
As transparency becomes key for robotics and AI, it will be necessary to evaluate the methods through which transparency is provided, including automatically generated natural language (NL) explanations.
Measurable counterfactual local explanations for any classifier
Adam White and Artur S. d’Avila Garcez. 2019 · 1908
Earlier work this paper cites.
One explanation does not fit all: A toolkit and taxonomy of AI explainability techniques
Vijay Arya, Rachel K. E. Bellamy, Pin-Yu Chen, Amit Dhurandhar, Michael Hind, Samuel C. Hoffman, Stephanie Houde, Q. Vera Liao, Ronny Luss, Aleksandra Mojsilovic, Sami Mourad, Pablo Pedemonte, Ramya Raghavendra, John T. Richards, Prasanna Sattigeri, Karthikeyan Shanmugam, Moninder Singh, Kush R. Varshney, Dennis Wei, and Yunfeng Zhang. 2019 · 1909
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries Chin-Yew
Chin-Yew Lin. 1971 · 1971
Earlier work this paper cites.
The role of dialogue in providing scaffolded instruction
Annemarie Sullivan Palincsar. 1986 · 1986
Earlier work this paper cites.
Evaluation in the context of natural language generation
C. Mellish and R. Dale. 1998 · 1998
Earlier work this paper cites.
Learning summarization by using similarities
Nicole Tourigny and Laurence Capus. 1998 · 1998
Earlier work this paper cites.
Explanations from intelligent systems: Theoretical foundations and implications for practice
Shirley Gregor and Izak Benbasat. 1999 · 1999
Earlier work this paper cites.
IBM Research Report Bleu : a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, Wei-jing Zhu, and Yorktown Heights. 2001 · 2001
Earlier work this paper cites.
Metodología de análisis de contenido. Teoría y práctica
Klaus Krippendorff. 1980 · 2004
Earlier work this paper cites.
Comparing automatic and human evaluation of NLG systems
Anja Belz and Ehud Reiter. 2006 · 2006
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Earlier work this paper cites.
Simplicity and probability in causal explanation
Tania Lombrozo. 2007 · 2007
Earlier work this paper cites.
Can we evaluate the quality of generated text?
David Hardcastle and Donia Scott. 2008 · 2008
Earlier work this paper cites.
System building cost vs. output quality in data-to-text generation
Anja Belz and Eric Kow. 2009 · 2009
Earlier work this paper cites.
Question answering based on semantic graphs
Lorand Dali, Delia Rusu, Blaž Fortuna, Dunja Mladenić, and Marko Grobelnik. 2009 · 2009
Earlier work this paper cites.
A study into preferred explanations of virtual agent behavior
Maaike Harbers, Karel Van Den Bosch, and John Jules Ch Meyer. 2009 · 2009
Earlier work this paper cites.
Most relevant explanation in bayesian networks
Changhe Yuan, Heejin Lim, and Tsai-Ching Lu. 2011 · 2011
Earlier work this paper cites.
Computing Inter-Rater Reliability for Observational Data: An Overview and Tutorial
Kevin A. Hallgren. 2012 · 2012
Earlier work this paper cites.
Using integer linear programming for content selection, lexicalization, and aggregation to produce compact texts from OWL ontologies
Gerasimos Lampouras and Ion Androutsopoulos. 2013 · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Automatic Summarization of Tweets in Providing Indonesian Trending Topic Explanation
Yosef Ardhito Winatmoko and Masayu Leylia Khodra. 2013 · 2013
Earlier work this paper cites.
Cluster-based prediction of user ratings for stylistic surface realisation
Nina Dethlefs, Heriberto Cuayáhuitl, Helen Hastie, Verena Rieser, and Oliver Lemon. 2014 · 2014
Earlier work this paper cites.
A snapshot of NLG evaluation practices 2005 - 2014
Dimitra Gkatzia and Saad Mahamood. 2015 · 2014
Earlier work this paper cites.
A comparative evaluation methodology for NLG in interactive systems
Helen Hastie and Anja Belz. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Evaluating Explanations
David B. Leake. 2014 · 2014
Earlier work this paper cites.
Anomaly detection in vessel tracks using Bayesian networks
Steven Mascaro, Ann Nicholson, and Kevin Korb. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Nlprov: Natural language provenance
Daniel Deutch, Nave Frost, and Amir Gilad. 2016 · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
On evaluation of natural language processing tasks: Is gold standard evaluation methodology a good solution?
Vojtěch Kovář, Miloš Jakubíček, and Aleš Horák. 2016 · 2016
Cited alongside, same era.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Cited alongside, same era.
Statistical natural language generation from tabular non-textual data
Joy Mahapatra, Sudip Kumar Naskar, and Sivaji Bandyopadhyay. 2016 · 2016
Cited alongside, same era.
”why should i trust you?” explaining the predictions of any classifier
Some insights towards a unified semantic representation of explanation for eXplainable artificial intelligence
Ismaïl Baaj, Jean-Philippe Poli, and Wassila Ouerdane. 2019 · 2019
Later among the works it cites.
The effects of example-based explanations in a machine learning interface
Carrie J. Cai, Jonas Jongejan, and Jess Holbrook. 2019 · 2019
Later among the works it cites.
A Survey of Explainable AI Terminology
Miruna-Adriana Clinciu and Helen Hastie. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
ELI5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 2019
Later among the works it cites.
Assessing the factual accuracy of generated text
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
How people explain action (and autonomous intelligent systems should too)
Maartje M.A. De Graaf and Bertram F. Malle. 2017 · 2017
Cited alongside, same era.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and B. Kim. 2017 · 2017
Cited alongside, same era.
Learning From Explanations Using Sentiment and Advice in RL
Samantha Krening, Brent Harrison, Karen M. Feigh, Charles Lee Isbell, Mark Riedl, and Andrea Thomaz. 2017 · 2017
Cited alongside, same era.
PASS: A Dutch data-to-text system for soccer, targeted towards specific audiences
Chris van der Lee, Emiel Krahmer, and Sander Wubben. 2017 · 2017
Cited alongside, same era.
A study of snippet length and informativeness behaviour, performance and user experience
David Maxwell, Leif Azzopardi, and Yashar Moshfeghi. 2017 · 2017
Cited alongside, same era.
Explainable AI: beware of inmates running the asylum
Tim Miller, Piers Hower, and Liz Sonenberg. 2017 · 2017
Cited alongside, same era.
Ben Goodrich, Vinay Rao, Peter J. Liu, and Mohammad Saleh. 2019 · 2019
Later among the works it cites.
Peter A. Jansen, Elizabeth Wainwright, Steven Marmorstein, and Clayton T. Morrison. 2019 · 2019
Later among the works it cites.
Text Generation from Knowledge Graphs with Graph Transformers
Rik Koncel-Kedziorski, Dhanush Bekal, Yi Luan, Mirella Lapata, and Hannaneh Hajishirzi. 2019 · 2019
Later among the works it cites.
A grounded interaction protocol for explainable artificial intelligence
Prashan Madumal, Liz Sonenberg, Tim Miller, and Frank Vetere. 2019 · 2019
Later among the works it cites.
On bayesian new edge prediction and anomaly detection in computer networks
Silvia Metelli and Nicholas Heard. 2019 · 2019
Later among the works it cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Later among the works it cites.
Sentence-BERT: Sentence embeddings using siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2020 · 2019
Later among the works it cites.
Human Compatible: Artificial Intelligence and the Problem of Control
S. Russell. 2019 · 2019
Later among the works it cites.
Counterfactual Explanations of Machine Learning Predictions: Opportunities and Challenges for AI Safety
Kacper Sokol and Peter A. Flach. 2019 · 2019
Later among the works it cites.
Disentangling the properties of human evaluation methods: A classification system to support comparability, meta-evaluation and reproducibility testing
Anya Belz, Simon Mille, and David M. Howcroft. 2020 · 2020
Later among the works it cites.
Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser. 2020 · 2020
Later among the works it cites.
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Later among the works it cites.
NILE : Natural language inference with faithful natural language explanations
Sawan Kumar and Partha Talukdar. 2020 · 2020
Later among the works it cites.
Qed: A framework and dataset for explanations in question answering
Matthew Lamm, Jennimaria Palomaki, Chris Alberti, Daniel Andor, Eunsol Choi, Livio Baldini Soares, and Michael Collins. 2020 · 2020
Later among the works it cites.
Anomaly detection in smart homes using Bayesian networks
Sasan Saqaeeyan, Hamid Haj Seyyed Javadi, and Hossein Amirkhani. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Anomaly detection system for water networks in northern ethiopia using bayesian inference
Zaid Tashman, Christoph Gorder, Sonali Parthasarathy, Mohamad M. Nasr-Azadani, and Rachel Webre. 2020 · 2020
Later among the works it cites.
Fact-based content weighting for evaluating abstractive summarisation
Xinnuo Xu, Ondřej Dušek, Jingyi Li, Verena Rieser, and Ioannis Konstas. 2020 · 2020
Later among the works it cites.
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Article 22 EU GDPR ”Automated individual decision-making, including profiling”
European Commission. 2018 · 2021
Closest in time.
Bert word embeddings tutorial
Chris McCormick and Nick Ryan. 2019 · 2021
Closest in time.
Does Gender Influence Online Survey Participation? A Record-Linkage Analysis of University Faculty Online Survey Response Behavior
William Smith. 2008 · 2021
Closest in time.