Nuanced metrics for measuring unintended bias with real data for text classification
Original
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019 · 1903
Earlier work this paper cites.
Intrinsic versus extrinsic evaluations of parsing systems
Diego Mollá and Ben Hutchinson. 2003 · 2003
Earlier work this paper cites.
The relationship between precision-recall and roc curves
Jesse Davis and Mark Goadrich. 2006 · 2006
Earlier work this paper cites.
Classification with a reject option using a hinge loss
Peter L. Bartlett and Marten H. Wegkamp. 2008 · 2008
Earlier work this paper cites.
Detection of harassment on web 2.0
Dawei Yin, Zhenzhen Xue, Liangjie Hong, Brian D Davison, April Kontostathis, and Lynne Edwards. 2009 · 2009
Earlier work this paper cites.
Fairness and robustness in invariant learning: A case study in toxicity classification
Original
Robert Adragna, Elliot Creager, David Madras, and Richard Zemel. 2020 · 2011
Earlier work this paper cites.
Modeling the detection of textual cyberbullying
Karthik Dinakar, Roi Reichart, and Henry Lieberman. 2011 · 2011
Earlier work this paper cites.
Wilds: A benchmark of in-the-wild distribution shifts
Original
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Sara Beery, et al. 2020 · 2012
Earlier work this paper cites.
Antisocial behavior in online discussion communities
Justin Cheng, Cristian Danescu-Niculescu-Mizil, and Jure Leskovec. 2015 · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Mahdi Pakdaman Naeini, Gregory F. Cooper, and Milos Hauskrecht. 2015 · 2015
Earlier work this paper cites.
Introduction to uncertainty quantification
T. J. Sullivan. 2015 · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Original
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. 2016 · 2016
Earlier work this paper cites.
Learning with rejection
Corinna Cortes, Giulia DeSalvo, and Mehryar Mohri. 2016 · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016 · 2016
Earlier work this paper cites.
Anyone can become a troll: Causes of trolling behavior in online discussions
Justin Cheng, Michael Bernstein, Cristian Danescu-Niculescu-Mizil, and Jure Leskovec. 2017 · 2017
Earlier work this paper cites.