Fetching the paper…
Reading the bibliography…
People are relying on AI agents to assist them with various tasks.
A method for comparing two hierarchical clusterings
Edward B Fowlkes and Colin L Mallows · 1983
Earlier work this paper cites.
Controlling the false discovery rate: a practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg · 1995
Earlier work this paper cites.
On the use of the adjusted rand index as a metric for evaluating supervised classification
Jorge M Santos and Mark Embrechts · 2009
Earlier work this paper cites.
Limits in decision making arise from limits in memory retrieval
Gyslain Giguère and Bradley C Love · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Overrides of medication-related clinical decision support alerts in outpatients
Karen C Nanji, Sarah P Slight, Diane L Seger, Insook Cho, Julie M Fiskio, Lisa M Redden, Lynn A Volk, and David W Bates · 2014
Earlier work this paper cites.
Visual category learning
Jennifer J Richler and Thomas J Palmeri · 2014
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Reminders of past choices bias decisions for reward in humans
Aaron M Bornstein, Mel W Khaw, Daphna Shohamy, and Nathaniel D Daw · 2017
Earlier work this paper cites.
End-to-end learning of driving models from large-scale video datasets
Huazhe Xu, Yang Gao, Fisher Yu, and Trevor Darrell · 2017
Earlier work this paper cites.
Do explanations make vqa models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Predict responsibly: Improving fairness and accuracy by learning to defer
David Madras, Toni Pitassi, and Richard Zemel · 2018
Earlier work this paper cites.
Factsheets: Increasing trust in ai services through supplier’s declarations of conformity
Matthew Arnold, Rachel KE Bellamy, Michael Hind, Stephanie Houde, Sameep Mehta, Aleksandra Mojsilović, Ravi Nair, K Natesan Ramamurthy, Alexandra Olteanu, David Piorkowski, et al · 2019
Earlier work this paper cites.
Guidelines for human-ai interaction
Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al · 2019
Earlier work this paper cites.
Beyond accuracy: The role of mental models in human-ai team performance
Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S Lasecki, Daniel S Weld, and Eric Horvitz · 2019
Earlier work this paper cites.
Updates in human-ai teams: Understanding and addressing the performance/compatibility tradeoff
Gagan Bansal, Besmira Nushi, Ece Kamar, Daniel S Weld, Walter S Lasecki, and Eric Horvitz · 2019
Earlier work this paper cites.
" hello ai": Uncovering the onboarding needs of medical practitioners for human-ai collaborative decision-making
Carrie J Cai, Samantha Winter, David Steiner, Lauren Wilcox, and Michael Terry · 2019
Earlier work this paper cites.
What can ai do for me? evaluating machine learning interpretations in cooperative play
Shi Feng and Jordan Boyd-Graber · 2019
Earlier work this paper cites.
Will you accept an imperfect ai? exploring designs for adjusting end-user expectations of ai systems
Rafal Kocielnik, Saleema Amershi, and Paul N Bennett · 2019
Earlier work this paper cites.
An evaluation of the human-interpretability of explanation
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, and Finale Doshi-Velez · 2019
Earlier work this paper cites.
On human predictions with explanations and predictions of machine learning models: A case study on deception detection
Vivian Lai and Chenhao Tan · 2019
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
The algorithmic automation problem: Prediction, triage, and human effort
Maithra Raghu, Katy Blumer, Greg Corrado, Jon Kleinberg, Ziad Obermeyer, and Sendhil Mullainathan · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Errudite: Scalable, reproducible, and testable error analysis
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld · 2019
Earlier work this paper cites.
Understanding the effect of accuracy on trust in machine learning models
Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach · 2019
Earlier work this paper cites.
A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy
Emma Beede, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M Vardoulakis · 2020
Earlier work this paper cites.
Tweeteval: Unified benchmark and comparative evaluation for tweet classification
Francesco Barbieri, Jose Camacho-Collados, Leonardo Neves, and Luis Espinosa-Anke · 2020
Earlier work this paper cites.
Proxy tasks and subjective measures can be misleading in evaluating explainable ai systems
Zana Bucinca, Phoebe Lin, Krzysztof Z Gajos, and Elena L Glassman · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhu, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel S Weld · 2020
Cited alongside, same era.
Human evaluation of spoken vs. visual explanations for open-domain qa
Ana Valeria Gonzalez, Gagan Bansal, Angela Fan, Robin Jia, Yashar Mehdad, and Srinivasan Iyer · 2020
Cited alongside, same era.
Content moderation, ai, and the question of scale
Tarleton Gillespie · 2020
Cited alongside, same era.
Evaluating explainable ai: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal · 2020
Cited alongside, same era.
Adaptive testing of computer vision models
Irena Gao, Gabriel Ilharco, Scott Lundberg, and Marco Tulio Ribeiro · 2022
Later among the works it cites.
Describing sets of images with textual-pca
Oded Hupert, Idan Schwartz, and Lior Wolf · 2022
Later among the works it cites.
Imagenet-x: Understanding model mistakes with factor of variation annotations
Badr Youbi Idrissi, Diane Bouchacourt, Randall Balestriero, Ivan Evtimov, Caner Hazirbas, Nicolas Ballas, Pascal Vincent, Michal Drozdzal, David Lopez-Paz, and Mark Ibrahim · 2022
Later among the works it cites.
Distilling model failures as directions in latent space
Saachi Jain, Hannah Lawrence, Ankur Moitra, and Aleksander Madry · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning
Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan · 2020
Cited alongside, same era.
" why is’ chicago’deceptive?" towards building model-driven tutorials for humans
Vivian Lai, Han Liu, and Chenhao Tan · 2020
Cited alongside, same era.
Consistent estimators for learning to defer to an expert
Hussein Mozannar and David Sontag · 2020
Cited alongside, same era.
Dynasent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela · 2020
Cited alongside, same era.
Misplaced trust: Measuring the interference of machine learning in human decision-making
Harini Suresh, Natalie Lao, and Ilaria Liccardi · 2020
Cited alongside, same era.
No explainability without accountability: An empirical study of explanations and feedback in interactive ml
Alison Smith-Renner, Ron Fan, Melissa Birchfield, Tongshuang Wu, Jordan Boyd-Graber, Daniel S Weld, and Leah Findlater · 2020
Cited alongside, same era.
Neighbours matter: Image captioning with similar images
Qingzhong Wang, Jiuniu Wang, Antoni B Chan, Siyu Huang, Haoyi Xiong, Xingjian Li, and Dejing Dou · 2020
Cited alongside, same era.
Improving human-ai partnerships in child welfare: Understanding worker practices, challenges, and desires for algorithmic decision support
Anna Kawakami, Venkatesh Sivaraman, Hao-Fei Cheng, Logan Stapleton, Yanghuidi Cheng, Diana Qing, Adam Perer, Zhiwei Steven Wu, Haiyi Zhu, and Kenneth Holstein · 2022
Later among the works it cites.
Human-ai collaboration via conditional delegation: A case study of content moderation
Vivian Lai, Samuel Carton, Rajat Bhatnagar, Q Vera Liao, Yunfeng Zhang, and Chenhao Tan · 2022
Later among the works it cites.
A thorough review of models, evaluation metrics, and datasets on image captioning
Gaifang Luo, Lijun Cheng, Chao Jing, Can Zhao, and Guozhu Song · 2022
Later among the works it cites.
Vision transformer based model for describing a set of images as a story
Zainy M Malakan, Ghulam Mubashar Hassan, and Ajmal Mian · 2022
Later among the works it cites.
Teaching humans when to defer to a classifier via exemplars
Hussein Mozannar, Arvind Satyanarayan, and David Sontag · 2022
Later among the works it cites.
Chatgpt: Introducing chatgpt
OpenAI · 2022
Later among the works it cites.
Adaptive testing and debugging of nlp models
Marco Tulio Ribeiro and Scott Lundberg · 2022
Later among the works it cites.
Seal: Interactive tool for systematic error analysis and labeling
Nazneen Rajani, Weixin Liang, Lingjiao Chen, Meg Mitchell, and James Zou · 2022
Later among the works it cites.
Selective regression under fairness criteria
Abhin Shah, Yuheng Bu, Joshua K Lee, Subhro Das, Rameswar Panda, Prasanna Sattigeri, and Gregory W Wornell · 2022
Later among the works it cites.
When does dough become a bagel? analyzing the remaining mistakes on imagenet
Vijay Vasudevan, Benjamin Caine, Raphael Gontijo-Lopes, Sara Fridovich-Keil, and Rebecca Roelofs · 2022
Later among the works it cites.
Uncalibrated models can improve human-ai collaboration
Kailas Vodrahalli, Tobias Gerstenberg, and James Zou · 2022
Later among the works it cites.
Discovering bugs in vision models using off-the-shelf image generation and captioning
Olivia Wiles, Isabela Albuquerque, and Sven Gowal · 2022
Later among the works it cites.
On distinctive image captioning via comparing and reweighting
Jiuniu Wang, Wenjia Xu, Qingzhong Wang, and Antoni B Chan · 2022
Later among the works it cites.
Drml: Diagnosing and rectifying vision models using language
Yuhui Zhang, Jeff Z HaoChen, Shih-Cheng Huang, Kuan-Chieh Wang, James Zou, and Serena Yeung · 2022
Later among the works it cites.
Describing differences between text distributions with natural language
Ruiqi Zhong, Charlie Snell, Dan Klein, and Jacob Steinhardt · 2022
Later among the works it cites.
Escape: Countering systematic errors from machine’s blind spots via interactive visual analysis
Yongsu Ahn, Yu-Ru Lin, Panpan Xu, and Zeng Dai · 2023
Closest in time.
Learning personalized decision support policies
Umang Bhatt, Valerie Chen, Katherine M Collins, Parameswaran Kamalaruban, Emma Kallina, Adrian Weller, and Ameet Talwalkar · 2023
Closest in time.
Improving human-ai collaboration with descriptions of ai behavior
Ángel Alexander Cabrera, Adam Perer, and Jason I Hong · 2023
Closest in time.
Raymond Fok and Daniel S Weld · 2023
Closest in time.
Brihi Joshi, Ziyi Liu, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang, Yejin Choi, and Xiang Ren · 2023
Closest in time.
Training towards critical use: Learning to situate ai predictions relative to human knowledge
Anna Kawakami, Luke Guerdan, Yanghuidi Cheng, Kate Glazko, Matthew Lee, Scott Carter, Nikos Arechiga, Haiyi Zhu, and Kenneth Holstein · 2023
Closest in time.
Shuai Ma, Ying Lei, Xinru Wang, Chengbo Zheng, Chuhan Shi, Ming Yin, and Xiaojuan Ma · 2023
Closest in time.
Who should predict? exact algorithms for learning to defer to humans
Hussein Mozannar, Hunter Lang, Dennis Wei, Prasanna Sattigeri, Subhro Das, and David Sontag · 2023
Closest in time.
https://www.prolific.co/
Prolific · 2023
Closest in time.
Mass-producing failures of multimodal systems with language models
Shengbang Tong, Erik Jones, and Jacob Steinhardt · 2023
Closest in time.
Explanations can reduce overreliance on ai systems during decision-making
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S Bernstein, and Ranjay Krishna · 2023
Closest in time.
Peter Zhang · 2023
Closest in time.