Fetching the paper…
Reading the bibliography…
Although AI holds promise for improving human decision making in societally critical domains, it remains an open question how human-AI teams can reliably outperform AI alone and human alone in challenging prediction tasks (also known as complementary performance).
Trust between humans and machines, and the design of decision aids
Bonnie M Muir. 1987 · 1987
Earlier work this paper cites.
Daubert v. Merrell Dow Pharmaceuticals, Inc
Supreme Court of the United States. 1993 · 1993
Earlier work this paper cites.
Humans and automation: Use, misuse, disuse, abuse
Raja Parasuraman and Victor Riley. 1997 · 1997
Earlier work this paper cites.
Getting access to what goes on in people’s heads?: reflections on the think-aloud technique. In Proceedings of the second Nordic conference on Human-computer interaction . ACM, 101–110
Janni Nielsen, Torkil Clemmensen, and Carsten Yssing. 2002 · 2002
Earlier work this paper cites.
The structure and function of explanations
Tania Lombrozo. 2006 · 2006
Earlier work this paper cites.
Supporting trust calibration and the effective use of decision aids by presenting dynamic system confidence information
John M McGuirl and Nadine B Sarter. 2006 · 2006
Earlier work this paper cites.
Dataset shift in machine learning
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence. 2009 · 2009
Earlier work this paper cites.
State Court Processing Statistics, 1990-2009: Felony Defendants in Large Urban Counties
United States Department of Justice. Office of Justice Programs. Bureau of Justice Statistics. 2014 · 2009
Earlier work this paper cites.
Trust calibration for automated decision aids
Maranda McBride and Shona Morgan. 2010 · 2010
Earlier work this paper cites.
Transfer learning
Lisa Torrey and Jude Shavlik. 2010 · 2010
Earlier work this paper cites.
Machine learning in non-stationary environments: Introduction to covariate shift adaptation
Masashi Sugiyama and Motoaki Kawanabe. 2012 · 2012
Earlier work this paper cites.
A design methodology for trust cue calibration in cognitive agents. In International conference on virtual, augmented and mixed reality . Springer, 251–262
Ewart J de Visser, Marvin Cohen, Amos Freedy, and Raja Parasuraman. 2014 · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of ICCV
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Prediction policy problems
Jon Kleinberg, Jens Ludwig, Sendhil Mullainathan, and Ziad Obermeyer. 2015 · 2015
Earlier work this paper cites.
Are well-calibrated users effective users? Associations between calibration of trust and performance on an automation-aided task
Stephanie M Merritt, Deborah Lee, Jennifer L Unnerstall, and Kelli Huber. 2015 · 2015
Earlier work this paper cites.
Machine Bias
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. 2016 · 2016
Earlier work this paper cites.
The effects of automatic speech recognition quality on human transcription latency. In Proceedings of the 13th Web for All Conference . 1–8
Yashesh Gaur, Walter S Lasecki, Florian Metze, and Jeffrey P Bigham. 2016 · 2016
Earlier work this paper cites.
Machine learning basics
Ian Goodfellow, Y Bengio, and A Courville. 2016 · 2016
Earlier work this paper cites.
Interacting with predictions: Visual inspection of black-box machine learning models. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems . ACM, 5686–5697
Josua Krause, Adam Perer, and Kenney Ng. 2016 · 2016
Earlier work this paper cites.
Interpretable decision sets: A joint framework for description and prediction. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1675–1684
Himabindu Lakkaraju, Stephen H Bach, and Jure Leskovec. 2016 · 2016
Earlier work this paper cites.
The mythos of model interpretability
Zachary C Lipton. 2016 · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier. In Proceedings of KDD
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
State of Wisconsin, Plaintiff-Respondent, v. Eric L. Loomis, Defendant-Appellant
Supreme Court of Wisconsin. 2016 · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions. In Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 1885–1894
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Scribe: deep integration of human and machine intelligence to caption speech in real time
Walter S Lasecki, Christopher D Miller, Iftekhar Naim, Raja Kushalnagar, Adam Sadilek, Daniel Gildea, and Jeffrey P Bigham. 2017 · 2017
Cited alongside, same era.
Sent to Prison by a Software Program’s Secret Algorithms
Adam Liptak. 2017 · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions. In Proceedings of the 31st international conference on neural information processing systems . 4768–4777
Scott M Lundberg and Su-In Lee. 2017 · 2017
Cited alongside, same era.
Counterfactual explanations without opening the black box: Automated decisions and the GDPR
Sandra Wachter, Brent Mittelstadt, and Chris Russell. 2017 · 2017
Cited alongside, same era.
Automatic alt-text: Computer-generated image descriptions for blind users on a social network service. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing . 1180–1192
Local Decision Pitfalls in Interactive Machine Learning: An Investigation into Feature Selection in Sentiment Analysis
Tongshuang Wu, Daniel S Weld, and Jeffrey Heer. 2019 · 2019
Later among the works it cites.
Unremarkable ai: Fitting intelligent decision support into critical, clinical decision-making processes. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . 1–11
Qian Yang, Aaron Steinfeld, and John Zimmerman. 2019 · 2019
Later among the works it cites.
Understanding the Effect of Accuracy on Trust in Machine Learning Models. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . ACM, 279
Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. 2019 · 2019
Later among the works it cites.
A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic Retinopathy. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–12
Emma Beede, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M Vardoulakis. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shaomei Wu, Jeffrey Wieland, Omid Farivar, and Julie Schiller. 2017 · 2017
Cited alongside, same era.
Explaining explanations: An overview of interpretability of machine learning. In 2018 IEEE 5th International Conference on data science and advanced analytics (DSAA) . IEEE, 80–89
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal. 2018 · 2018
Cited alongside, same era.
Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 3608–3617
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. 2018 · 2018
Cited alongside, same era.
Human decisions and machine predictions
Jon Kleinberg, Himabindu Lakkaraju, Jure Leskovec, Jens Ludwig, and Sendhil Mullainathan. 2018 · 2018
Cited alongside, same era.
Explainable machine-learning predictions for the prevention of hypoxaemia during surgery
Scott M Lundberg, Bala Nair, Monica S Vavilala, Mayumi Horibe, Michael J Eisses, Trevor Adams, David E Liston, Daniel King-Wai Low, Shu-Fang Newman, Jerry Kim, et al · 2018
Cited alongside, same era.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. 2018 · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm. 2019 · 2019
Cited alongside, same era.
Proxy tasks and subjective measures can be misleading in evaluating explainable AI systems. In Proceedings of the 25th International Conference on Intelligent User Interfaces . 454–464
Zana Buçinca, Phoebe Lin, Krzysztof Z Gajos, and Elena L Glassman. 2020 · 2020
Later among the works it cites.
Feature-Based Explanations Don’t Help People Detect Misclassifications of Online Toxicity. In Proceedings of the International AAAI Conference on Web and Social Media , Vol. 14. 95–106
Samuel Carton, Qiaozhu Mei, and Paul Resnick. 2020 · 2020
Later among the works it cites.
Music Creation by Example. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems
Emma Frid, Ceslo Gomes, and Zeyu Jin. 2020 · 2020
Later among the works it cites.
Bhavya Ghai, Q Vera Liao, Yunfeng Zhang, Rachel Bellamy, and Klaus Mueller. 2020 · 2020
Later among the works it cites.
Interpreting Interpretability: Understanding Data Scientists’ Use of Interpretability Tools for Machine Learning. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–14
Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan. 2020 · 2020
Later among the works it cites.
“Why is ‘Chicago’ deceptive?” Towards Building Model-Driven Tutorials for Humans. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13
Vivian Lai, Han Liu, and Chenhao Tan. 2020 · 2020
Later among the works it cites.
A transfer learning method with deep residual network for pediatric pneumonia diagnosis
Gaobo Liang and Lixin Zheng. 2020 · 2020
Later among the works it cites.
The limits of human predictions of recidivism
Zhiyuan “Jerry” Lin, Jongbin Jung, Sharad Goel, Jennifer Skeem, et al · 2020
Later among the works it cites.
Novice-AI Music Co-Creation via AI-Steering Tools for Deep Generative Models. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13
Ryan Louie, Andy Coenen, Cheng Zhi Huang, Michael Terry, and Carrie J Cai. 2020 · 2020
Later among the works it cites.
International evaluation of an AI system for breast cancer screening
Scott Mayer McKinney, Marcin Sieniek, Varun Godbole, Jonathan Godwin, Natasha Antropova, Hutan Ashrafian, Trevor Back, Mary Chesus, Greg C Corrado, Ara Darzi, et al · 2020
Later among the works it cites.
The Language Interpretability Tool: Extensible, Interactive Visualizations and Analysis for NLP Models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . 107–118
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, et al · 2020
Later among the works it cites.
CheXplain: Enabling Physicians to Explore and Understand Data-Driven, AI-Enabled Medical Imaging Analysis. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13
Yao Xie, Melody Chen, David Kao, Ge Gao, and Xiang ‘Anthony’ Chen. 2020 · 2020
Later among the works it cites.
Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency . 295–305
Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. 2020 · 2020
Later among the works it cites.
A comprehensive survey on transfer learning
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He. 2020 · 2020
Later among the works it cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–16
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021 · 2021
Closest in time.
You’d Better Stop! Understanding Human Reliance on Machine Learning Models under Covariate Shift. In 13th ACM Web Science Conference 2021 . 120–129
Chun-Wei Chiang and Ming Yin. 2021 · 2021
Closest in time.
Wilds: A benchmark of in-the-wild distribution shifts. In International Conference on Machine Learning . PMLR, 5637–5664
Pang Wei Koh, Shiori Sagawa, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, et al · 2021
Closest in time.
Manipulating and measuring model interpretability. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–52
Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Wortman Vaughan, and Hanna Wallach. 2021 · 2021
Closest in time.
Are Explanations Helpful? A Comparative Study of the Effects of Explanations in AI-Assisted Decision-Making. In 26th International Conference on Intelligent User Interfaces . 318–328
Xinru Wang and Ming Yin. 2021 · 2021
Closest in time.
Adversarial Examples for Evaluating Reading Comprehension Systems. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . 2021–2031
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.