Fetching the paper…
Reading the bibliography…
Forming a reliable judgement of a machine learning (ML) model's appropriateness for an application ecosystem is critical for its responsible use, and requires considering a broad range of factors including harms, benefits, and responsibilities.
The future of data analysis
John W Tukey. 1962 · 1962
Earlier work this paper cites.
The Interpretation of Cultures
Clifford Geertz. 1973 · 1973
Earlier work this paper cites.
Foundation of evaluation
Cornelis Joost Van Rijsbergen. 1974 · 1974
Earlier work this paper cites.
Image-Music-Text
Roland Barthes. 1977 · 1977
Earlier work this paper cites.
Institutional ecology, ‘translations’ and boundary objects: Amateurs and professionals in Berkeley’s Museum of Vertebrate Zoology, 1907-39
Susan Leigh Star and James R Griesemer. 1989 · 1989
Earlier work this paper cites.
Cost-sensitive classification: Empirical evaluation of a hybrid genetic decision tree induction algorithm
Peter D Turney. 1994 · 1994
Earlier work this paper cites.
Automatically Identifying Gender Issues in Machine Translation using Perturbations. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 1991–1995
Hila Gonen and Kellie Webster. 2020 · 1995
Earlier work this paper cites.
Testing heuristics: We have it all wrong
John N Hooker. 1995 · 1995
Earlier work this paper cites.
Evaluating natural language processing systems: An analysis and review . Vol. 1083
Karen Sparck Jones and Julia R Galliers. 1995 · 1995
Earlier work this paper cites.
Analysis and visualization of classifier performance with nonuniform class and cost distributions. In Proceedings of AAAI-97 Workshop on AI Approaches to Fraud Detection & Risk Management . 57–63
Foster Provost and Tom Fawcett. 1997 · 1997
Earlier work this paper cites.
Nothing about us without us
James I Charlton. 1998 · 1998
Earlier work this paper cites.
Is statistics too difficult?
Frank Hampel and Eth Zurich. 1998 · 1998
Earlier work this paper cites.
The generative lexicon
James Pustejovsky. 1998 · 1998
Earlier work this paper cites.
Testing: a roadmap. In Proceedings of the Conference on the Future of Software Engineering . 61–72
Mary Jean Harrold. 2000 · 2000
Earlier work this paper cites.
Statistical modeling: The two cultures (with comments and a rejoinder by the author)
Leo Breiman. 2001 · 2001
Earlier work this paper cites.
Test driven development: A practical guide
Dave Astels. 2003 · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003 . 142–147
Erik Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
A structured experiment of test-driven development
Boby George and Laurie Williams. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Why question machine learning evaluation methods. In AAAI workshop on evaluation methods for machine learning . 6–11
Nathalie Japkowicz. 2006 · 2006
Earlier work this paper cites.
Donald Martin, Jr., Vinodkumar Prabhakaran, Jill Kuhlberg, Andrew Smart, and William S. Isaac. 2020 · 2006
Earlier work this paper cites.
Integrating technology readiness into technology acceptance: The TRAM model
Chien-Hsin Lin, Hsin-Yu Shih, and Peter J Sher. 2007 · 2007
Earlier work this paper cites.
Direct Importance Estimation with Model Selection and Its Application to Covariate Shift Adaptation
Masashi Sugiyama, Shinichi Nakajima, Hisashi Kashima, Paul Buenau, and Motoaki Kawanabe. 2007 · 2007
Earlier work this paper cites.
Contrastive Training for Improved Out-of-Distribution Detection
Jim Winkens, Rudy Bunel, Abhijit Guha Roy, Robert Stanforth, Vivek Natarajan, Joseph R Ledsam, Patricia MacWilliams, Pushmeet Kohli, Alan Karthikesalingam, Simon Kohl, et al · 2007
Earlier work this paper cites.
Ways of seeing
John Berger. 2008 · 2008
Earlier work this paper cites.
Metaphors we live by
George Lakoff and Mark Johnson. 2008 · 2008
Earlier work this paper cites.
A utility-driven approach to question ranking in social QA. In Proceedings of The 23rd International Conference on Computational Linguistics (COLING 2010) . 125–133
Razvan Bunescu and Yunfeng Huang. 2010 · 2010
Earlier work this paper cites.
Evaluating Machine Translation Utility via Semantic Role Labels.. In LREC . Citeseer
Chi-kiu Lo and Dekai Wu. 2010 · 2010
Earlier work this paper cites.
Subjective natural language problems: Motivations, applications, characterizations, and implications. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies . 107–112
Cecilia Ovesdotter Alm. 2011 · 2011
Earlier work this paper cites.
The art of software testing
Glenford J Myers, Corey Sandler, and Tom Badgett. 2011 · 2011
Earlier work this paper cites.
Evaluation: From Precision, Recall and F-Factor to ROC, Informedness, Markedness & Correlation
David Martin Ward Powers. 2011 · 2011
Earlier work this paper cites.
WILDS: A Benchmark of in-the-Wild Distribution Shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang. 2020 · 2012
Earlier work this paper cites.
The problem of area under the curve. In 2012 IEEE International conference on information science and technology . IEEE, 567–573
David Martin Ward Powers. 2012a · 2012
Earlier work this paper cites.
Introduction to Peircean visual semiotics
Tony Jappy. 2013 · 2013
Earlier work this paper cites.
Soft skills in software engineering: A study of its demand by software companies in Uruguay. In 2013 6th international workshop on cooperative and human aspects of software engineering (CHASE) . IEEE, 133–136
Gerardo Matturro. 2013 · 2013
Earlier work this paper cites.
What the F-measure doesn’t measure: Features, Flaws, Fallacies and Fixes
David Martin Ward Powers. 2014 · 2014
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty. 2015 · 2015
Earlier work this paper cites.
Philosophy of engineering: What it is and why it matters
William Bulleit, Jon Schmidt, Irfan Alvi, Erik Nelson, and Tonatiuh Rodriguez-Nikl. 2015 · 2015
Earlier work this paper cites.
Could Big Data be the end of theory in science? A few remarks on the epistemology of data-driven science
Fulvio Mazzocchi. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016 · 2016
Earlier work this paper cites.
Complementarity, F-score, and NLP Evaluation. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16) . 261–266
Leon Derczynski. 2016 · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
The social impact of natural language processing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . 591–598
Dirk Hovy and Shannon L Spruit. 2016 · 2016
Earlier work this paper cites.
Indigenous data sovereignty: Toward an agenda
Tahu Kukutai and John Taylor. 2016 · 2016
Earlier work this paper cites.
TGIF: A new dataset and benchmark on animated GIF description. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 4641–4650
Yuncheng Li, Yale Song, Liangliang Cao, Joel Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo. 2016 · 2016
Earlier work this paper cites.
Technology and the virtues: A philosophical guide to a future worth wanting
Shannon Vallor. 2016 · 2016
Earlier work this paper cites.
Fairness in machine learning
Solon Barocas, Moritz Hardt, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
The ML test score: A rubric for ML production readiness and technical debt reduction. In 2017 IEEE International Conference on Big Data (Big Data) . IEEE, 1123–1132
Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D Sculley. 2017 · 2017
Earlier work this paper cites.
NeuralPower: Predict and deploy energy-efficient convolutional neural networks. In Asian Conference on Machine Learning . PMLR, 622–637
Ermao Cai, Da-Cheng Juan, Dimitrios Stamoulis, and Diana Marculescu. 2017 · 2017
Earlier work this paper cites.
Fair prediction with disparate impact: A study of bias in recidivism prediction instruments
Alexandra Chouldechova. 2017 · 2017
Earlier work this paper cites.
Algorithmic decision making and the cost of fairness. In Proceedings of the 23rd acm sigkdd international conference on knowledge discovery and data mining . 797–806
Sam Corbett-Davies, Emma Pierson, Avi Feller, Sharad Goel, and Aziz Huq. 2017 · 2017
Earlier work this paper cites.
Towards linguistically generalizable NLP systems: A workshop and shared task
Allyson Ettinger, Sudha Rao, Hal Daumé III, and Emily M Bender. 2017 · 2017
Earlier work this paper cites.
Alexandre Lacoste, Thomas Boquet, Negar Rostamzadeh, Boris Oreshkin, Wonchang Chung, and David Krueger. 2017 · 2017
Earlier work this paper cites.
On Chomsky and the two cultures of statistical learning
Peter Norvig. 2017 · 2017
Earlier work this paper cites.
Watset: Automatic induction of synsets from a graph of synonyms. In 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017 . Association for Computational Linguistics, 1579–1590
Dmitry Ustalov, Alexander Panchenko, and Chris Biemann. 2017 · 2017
Earlier work this paper cites.
Domain Adaptation with Adversarial Training and Graph Embeddings. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 1077–1087
Firoj Alam, Shafiq Joty, and Muhammad Imran. 2018 · 2018
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Fairness in machine learning: Lessons from political philosophy. In Conference on Fairness, Accountability and Transparency . PMLR, 149–159
Reuben Binns. 2018 · 2018
Cited alongside, same era.
The measure and mismeasure of fairness: A critical review of fair machine learning
Sam Corbett-Davies and Sharad Goel. 2018 · 2018
Cited alongside, same era.
Lecture notes on fair division
Ulle Endriss. 2018 · 2018
Cited alongside, same era.
Learn-to-score: Efficient 3D scene exploration by predicting view utility. In Proceedings of the European conference on computer vision (ECCV) . 437–452
Benjamin Hepp, Debadeepta Dey, Sudipta N Sinha, Ashish Kapoor, Neel Joshi, and Otmar Hilliges. 2018 · 2018
Cited alongside, same era.
Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 conference on fairness, accountability, and transparency . 33–44
Inioluwa Deborah Raji, Andrew Smart, Rebecca N White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Beyond Technical Skills in Software Testing: Automated versus Manual Testing. In Proceedings of the IEEE/ACM 42nd International Conference on Software Engineering Workshops . 161–164
Mary Sánchez-Gordón, Laxmi Rijal, and Ricardo Colomo-Palacios. 2020 · 2020
Later among the works it cites.
Green AI
Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. 2020 · 2020
Later among the works it cites.
Basic rights: Subsistence, affluence, and US foreign policy
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexandre Lacoste, Boris Oreshkin, Wonchang Chung, Thomas Boquet, Negar Rostamzadeh, and David Krueger. 2018 · 2018
Cited alongside, same era.
Delayed impact of fair machine learning. In International Conference on Machine Learning . PMLR, 3150–3158
Lydia T Liu, Sarah Dean, Esther Rolf, Max Simchowitz, and Moritz Hardt. 2018 · 2018
Cited alongside, same era.
AI and Big Data: A blueprint for a human rights, social and ethical impact assessment
Alessandro Mantelero. 2018 · 2018
Cited alongside, same era.
Preventing disparities: Bayesian and frequentist methods for assessing fairness in machine learning decision-support models
Douglas S McNair. 2018 · 2018
Cited alongside, same era.
Image to image translation for domain adaptation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 4500–4509
Zak Murez, Soheil Kolouri, David Kriegman, Ravi Ramamoorthi, and Kyungnam Kim. 2018 · 2018
Cited alongside, same era.
Connecting pixels to privacy and utility: Automatic redaction of private information in images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 8466–8475
Tribhuvanesh Orekondy, Mario Fritz, and Bernt Schiele. 2018 · 2018
Cited alongside, same era.
Artificial intelligence & human rights: Opportunities & risks
Filippo A Raso, Hannah Hilligoss, Vivek Krishnamurthy, Christopher Bavitz, and Levin Kim. 2018 · 2018
Cited alongside, same era.
Fashion-gen: The generative fashion dataset and challenge
Negar Rostamzadeh, Seyedarian Hosseini, Thomas Boquet, Wojciech Stokowiec, Ying Zhang, Christian Jauvin, and Chris Pal. 2018 · 2018
Cited alongside, same era.
Henry Shue. 2020 · 2020
Later among the works it cites.
Reliance on metrics is a fundamental challenge for AI. In Proceedings of the Ethics of Data Science Conference
RL Thomas and D Uminsky. 2020 · 2020
Later among the works it cites.
Machine learning testing: Survey, landscapes and horizons
Jie M Zhang, Mark Harman, Lei Ma, and Yang Liu. 2020a · 2020
Later among the works it cites.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li. 2020b · 2020
Later among the works it cites.
Not one but many tradeoffs: Privacy vs. utility in differentially private machine learning. In Proceedings of the 2020 ACM SIGSAC Conference on Cloud Computing Security Workshop . 15–26
Benjamin Zi Hao Zhao, Mohamed Ali Kaafar, and Nicolas Kourtellis. 2020 · 2020
Later among the works it cites.
AI, big data, and the future of consent
Adam J Andreotta, Nin Kirkham, and Marco Rizzi. 2021 · 2021
Later among the works it cites.
What We Can’t Measure, We Can’t Understand: Challenges to Demographic Data Procurement in the Pursuit of Fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . 249–260
McKane Andrus, Elena Spitzer, Jeffrey Brown, and Alice Xiang. 2021 · 2021
Later among the works it cites.
Designing Disaggregated Evaluations of AI Systems: Choices, Considerations, and Tradeoffs
Solon Barocas, Anhong Guo, Ece Kamar, Jacquelyn Krones, Meredith Ringel Morris, Jennifer Wortman Vaughan, Duncan Wadsworth, and Hanna Wallach. 2021 · 2021
Later among the works it cites.
Toward a Perspectivist Turn in Ground Truthing for Predictive Computing
Valerio Basile, Federico Cabitza, Andrea Campagner, and Michael Fell. 2021 · 2021
Later among the works it cites.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . 610–623
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Later among the works it cites.
Deep learning for AI
Yoshua Bengio, Yann Lecun, and Geoffrey Hinton. 2021 · 2021
Later among the works it cites.
The values encoded in machine learning research
Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. 2021 · 2021
Later among the works it cites.
Systematic Inequalities in Language Technology Performance across the World’s Languages
Damián Blasi, Antonios Anastasopoulos, and Graham Neubig. 2021 · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Later among the works it cites.
What Will it Take to Fix Benchmarking in Natural Language Understanding?. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 4843–4855
Samuel Bowman and George Dahl. 2021 · 2021
Later among the works it cites.
Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21) . 2633–2650
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Later among the works it cites.
Overinterpretation reveals image classification model pathologies
Brandon Carter, Siddhartha Jain, Jonas W Mueller, and David Gifford. 2021 · 2021
Later among the works it cites.
Mandoline: Model Evaluation under Distribution Shift. In International Conference on Machine Learning . PMLR, 1617–1629
Mayee Chen, Karan Goel, Nimit S Sohoni, Fait Poms, Kayvon Fatahalian, and Christopher Ré. 2021 · 2021
Later among the works it cites.
Learning to be Fair: A Consequentialist Approach to Equitable Decision-Making
Alex Chohlas-Wood, Madison Coots, Emma Brunskill, and Sharad Goel. 2021 · 2021
Later among the works it cites.
Excavating AI: The politics of images in machine learning training sets
Kate Crawford and Trevor Paglen. 2021 · 2021
Later among the works it cites.
Mind the gap! On the future of AI research
Emma Dahlin. 2021 · 2021
Later among the works it cites.
Head2Toe: Utilizing Intermediate Representations for Better OOD Generalization
Utku Evci, Vincent Dumoulin, Hugo Larochelle, and Michael Curtis Mozer. 2021 · 2021
Later among the works it cites.
Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task Learning. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 2591–2597
Tommaso Fornaciari, Alexandra Uma, Silviu Paun, Barbara Plank, Dirk Hovy, and Massimo Poesio. 2021 · 2021
Later among the works it cites.
The (im) possibility of fairness: Different value systems require different mechanisms for fair decision making
Sorelle A Friedler, Carlos Scheidegger, and Suresh Venkatasubramanian. 2021 · 2021
Later among the works it cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021 · 2021
Later among the works it cites.
ALT-MAS: A Data-Efficient Framework for Active Testing of Machine Learning Algorithms
Huong Ha, Sunil Gupta, Santu Rana, and Svetha Venkatesh. 2021 · 2021
Later among the works it cites.
“I don’t think these devices are very culturally sensitive.”—The impact of errors on African Americans in Automated Speech Recognition
Courtney Heldreth, Michal Lahav, Zion Mengesha, Juliana Sublewski, and Elyse Tuennerman. 2021 · 2021
Later among the works it cites.
Towards accountability for machine learning datasets: Practices from software engineering and infrastructure. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . 560–575
Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. 2021 · 2021
Later among the works it cites.
Towards Benchmarking the Utility of Explanations for Model Debugging. In Proceedings of the First Workshop on Trustworthy Natural Language Processing . 68–73
Maximilian Idahl, Lijun Lyu, Ujwal Gadiraju, and Avishek Anand. 2021 · 2021
Later among the works it cites.
Measurement and fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . 375–385
Abigail Z Jacobs and Hanna Wallach. 2021 · 2021
Later among the works it cites.
Learning high fidelity depths of dressed humans by watching social media dance videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12753–12762
Yasamin Jafarian and Hyun Soo Park. 2021 · 2021
Later among the works it cites.
Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research
Bernard Koch, Emily Denton, Alex Hanna, and Jacob G Foster. 2021 · 2021
Later among the works it cites.
Active testing: Sample-efficient model evaluation. In International Conference on Machine Learning . PMLR, 5753–5763
Jannik Kossen, Sebastian Farquhar, Yarin Gal, and Tom Rainforth. 2021 · 2021
Later among the works it cites.
Are We Learning Yet? A Meta Review of Evaluation Failures Across Machine Learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt. 2021 · 2021
Later among the works it cites.
Documenting Computer Vision Datasets: An Invitation to Reflexive Data Practices. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency . 161–172
Milagros Miceli, Tianling Yang, Laurens Naudts, Martin Schuessler, Diana Serbanescu, and Alex Hanna. 2021 · 2021
Later among the works it cites.
Evaluating the Robustness of Neural Language Models to Input Perturbations. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 1558–1570
Milad Moradi and Matthias Samwald. 2021 · 2021
Later among the works it cites.
On Releasing Annotator-Level Labels and Information in Datasets. In Proceedings of The Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop . 133–138
Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021 · 2021
Later among the works it cites.
AI and the Everything in the Whole Wide World Benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
Inioluwa Deborah Raji, Emily Denton, Emily M Bender, Alex Hanna, and Amandalynne Paullada. 2021 · 2021
Later among the works it cites.
How do AI systems fail socially?: an engineering risk analysis approach. In 2021 IEEE International Symposium on Ethics in Engineering, Science and Technology (ETHICS) . 1–8
Shalaleh Rismani and Ajung Moon. 2021 · 2021
Later among the works it cites.
Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Association for Computational Linguistics, Online, 4486–4503
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. 2021 · 2021
Later among the works it cites.
Thinking Beyond Distributions in Testing Machine Learned Models. In NeurIPS 2021 Workshop on Distribution Shifts: Connecting Methods and Applications
Negar Rostamzadeh, Ben Hutchinson, Christina Greer, and Vinodkumar Prabhakaran. 2021 · 2021
Later among the works it cites.
Re-Imagining Algorithmic Fairness in India and Beyond. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (Virtual Event, Canada) (FAccT ’21) . Association for Computing Machinery, New York, NY, USA, 315–328
Nithya Sambasivan, Erin Arnesen, Ben Hutchinson, Tulsee Doshi, and Vinodkumar Prabhakaran. 2021a · 2021
Later among the works it cites.
“Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. In proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–15
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021b · 2021
Later among the works it cites.
Do datasets have politics? Disciplinary values in computer vision dataset development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. 2021 · 2021
Later among the works it cites.
Targeting the Benchmark: On Methodology in Current Natural Language Processing Research. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers) . 670–674
David Schlangen. 2021 · 2021
Later among the works it cites.
Consequentialism
Walter Sinnott-Armstrong. 2021 · 2021
Later among the works it cites.
Generalizing to Unseen Domains: A Survey on Domain Generalization. In Proceedings of IJCAI 2021
Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Wenjun Zeng, and Tao Qin. 2021 · 2021
Later among the works it cites.
Fashion iq: A new dataset towards retrieving images by natural language feedback. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11307–11317
Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris. 2021 · 2021
Later among the works it cites.
OpenAttack: An Open-source Textual Adversarial Attack Toolkit. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations . 363–371
Guoyang Zeng, Fanchao Qi, Qianrui Zhou, Tingji Zhang, Zixian Ma, Bairu Hou, Yuan Zang, Zhiyuan Liu, and Maosong Sun. 2021 · 2021
Later among the works it cites.
Anatomy of an AI System
Kate Crawford and Vladan Joler. 2018 · 2022
Closest in time.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022 · 2022
Closest in time.
A deep insight into measuring face image utility with general and face-specific image quality metrics. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 905–914
Biying Fu, Cong Chen, Olaf Henniger, and Naser Damer. 2022 · 2022
Closest in time.
Assessing the Fairness of AI Systems: AI Practitioners’ Processes, Challenges, and Needs for Support
Michael Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan, and Hanna Wallach. 2022 · 2022
Closest in time.
Simulated Adversarial Testing of Face Recognition Models
Nataniel Ruiz, Adam Kortylewski, Weichao Qiu, Cihang Xie, Sarah Adel Bargal, Alan Yuille, and Stan Sclaroff. 2022 · 2022
Closest in time.