Fetching the paper…
Reading the bibliography…
The labels used to train machine learning (ML) models are of paramount importance.
Calibration of probabilities: The state of the art
Lichtenstein, S.; Fischhoff, B.; and Phillips, L. D. 1977 · 1977
Earlier work this paper cites.
Maximum likelihood estimation of observer error-rates using the EM algorithm
Dawid, A. P.; and Skene, A. M. 1979 · 1979
Earlier work this paper cites.
Knowledge Discovery in Large Image Databases: Dealing with Uncertainties in Ground Truth
Smyth, P.; Burl, M. C.; Fayyad, U. M.; and Perona, P. 1994 · 1994
Earlier work this paper cites.
An Introduction to Vote-Counting Schemes
Levin, J.; and Nalebuff, B. 1995 · 1995
Earlier work this paper cites.
On the reality of cognitive illusions
Tversky, A.; and Kahneman, D. 1996 · 1996
Earlier work this paper cites.
Reviewing intuitive decision-making and uncertainty: the implications for medical education
Hall, K. H. 2002 · 2002
Earlier work this paper cites.
A Bayesian truth serum for subjective data
Prelec, D. 2004 · 2004
Earlier work this paper cites.
Beyer, L.; Hénaff, O. J.; Kolesnikov, A.; Zhai, X.; and van den Oord, A. 2020 · 2006
Earlier work this paper cites.
Uncertain Judgements: Eliciting Expert Probabilities
O’Hagan, A.; Buck, C. E.; Daneshkhah, A.; Eiser, J. R.; Garthwaite, P. H.; Jenkinson, D. J.; Oakley, J. E.; and Rakow, T. 2006 · 2006
Earlier work this paper cites.
ImageNet: Constructing a large-scale image database
Fei-Fei, L.; Deng, J.; and Li, K. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. 2009 · 2009
Earlier work this paper cites.
Whose vote should count more: Optimal integration of labels from labelers of unknown expertise
Whitehill, J.; Wu, T.-f.; Bergsma, J.; Movellan, J.; and Ruvolo, P. 2009 · 2009
Earlier work this paper cites.
Visual recognition with humans in the loop
Branson, S.; Wah, C.; Schroff, F.; Babenko, B.; Welinder, P.; Perona, P.; and Belongie, S. 2010 · 2010
Earlier work this paper cites.
SHELF: the Sheffield elicitation framework (version 2.0)
Oakley, J. E.; and O’Hagan, A. 2010 · 2010
Earlier work this paper cites.
The optimism bias
Sharot, T. 2011 · 2011
Earlier work this paper cites.
Lay understanding of probability distributions
Goldstein, D. G.; and Rothschild, D. 2014 · 2014
Earlier work this paper cites.
Learning classification models with soft-label information
Nguyen, Q.; Valizadegan, H.; and Hauskrecht, M. 2014 · 2014
Earlier work this paper cites.
The benefits of a model of annotation
Passonneau, R. J.; and Carpenter, B. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Earlier work this paper cites.
Explaining and Harnessing Adversarial Examples
Goodfellow, I.; Shlens, J.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; Dean, J.; et al. 2015 · 2015
Earlier work this paper cites.
Posterior probability matching and human perceptual decision making
Murray, R. F.; Patel, K.; and Yee, A. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Eliciting categorical data for optimal aggregation
Ho, C.-J.; Frongillo, R.; and Chen, Y. 2016 · 2016
Earlier work this paper cites.
Adversarial machine learning at scale
Kurakin, A.; Goodfellow, I.; and Bengio, S. 2016 · 2016
Cited alongside, same era.
Ambiguity Helps: Classification With Disagreements in Crowdsourced Annotations
Sharmanska, V.; Hernandez-Lobato, D.; Miguel Hernandez-Lobato, J.; and Quadrianto, N. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; and Wojna, Z. 2016 · 2016
Cited alongside, same era.
Bayesian Aggregation of Categorical Distributions with Applications in Crowdsourcing
Augustin, A.; Venanzi, M.; Rogers, A.; and Jennings, N. R. 2017 · 2017
Cited alongside, same era.
Building machines that learn and think like people
Lake, B. M.; Ullman, T. D.; Tenenbaum, J. B.; and Gershman, S. J. 2017 · 2017
Cited alongside, same era.
Knowledge distillation: A survey
Gou, J.; Yu, B.; Maybank, S. J.; and Tao, D. 2021 · 2021
Later among the works it cites.
Uncertain Decisions Facilitate Better Preference Learning
Laidlaw, C.; and Russell, S. 2021 · 2021
Later among the works it cites.
Orbit: A real-world few-shot dataset for teachable object recognition
Massiceti, D.; Zintgraf, L.; Bronskill, J.; Theodorou, L.; Harris, M. T.; Cutrell, E.; Morrison, C.; Hofmann, K.; and Stumpf, S. 2021 · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks
Northcutt, C. G.; Athalye, A.; and Mueller, J. 2021 · 2021
Later among the works it cites.
On Releasing Annotator-Level Labels and Information in Datasets
Prabhakaran, V.; Davani, A. M.; and Diaz, M. 2021 · 2021
Later among the works it cites.
Forecast aggregation via peer prediction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, W.; Dai, B.; Humayun, A.; Tay, C.; Yu, C.; Smith, L. B.; Rehg, J. M.; and Song, L. 2017 · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
Pereyra, G.; Tucker, G.; Chorowski, J.; Kaiser, Ł.; and Hinton, G. 2017 · 2017
Cited alongside, same era.
Majority voting and pairing with multiple noisy labeling
Sheng, V. S.; Zhang, J.; Gu, B.; and Wu, X. 2017 · 2017
Cited alongside, same era.
DiverseNet: When One Right Answer Is Not Enough
Firman, M.; Campbell, N. D. F.; Agapito, L.; and Brostow, G. J. 2018 · 2018
Cited alongside, same era.
Prolific. ac—A subject pool for online experiments
Palan, S.; and Schitter, C. 2018 · 2018
Cited alongside, same era.
Active learning with confidence-based answers for crowdsourcing labeling tasks
Song, J.; Wang, H.; Gao, Y.; and An, B. 2018 · 2018
Cited alongside, same era.
Multi-label inference for crowdsourcing
Zhang, J.; and Wu, X. 2018 · 2018
Cited alongside, same era.
Wang, J.; Liu, Y.; and Chen, Y. 2021 · 2021
Later among the works it cites.
Delving deep into label smoothing
Zhang, C.-B.; Jiang, P.-T.; Hou, Q.; Wei, Y.; Han, Q.; Li, Z.; and Cheng, M.-M. 2021 · 2021
Later among the works it cites.
On the Utility of Prediction Sets in Human-AI Teams
Babbar, V.; Bhatt, U.; and Weller, A. 2022 · 2022
Closest in time.
Perspectives on Incorporating Expert Feedback into Model Updates
Chen, V.; Bhatt, U.; Heidari, H.; Weller, A.; and Talwalkar, A. 2022 · 2022
Closest in time.
Transfer and Marginalize: Explaining Away Label Noise with Privileged Information
Collier, M.; Jenatton, R.; Kokiopoulou, E.; and Berent, J. 2022 · 2022
Closest in time.
CrowdWorkSheets: Accounting for Individual and Collective Identities Underlying Crowdsourced Dataset Annotation
Díaz, M.; Kivlichan, I.; Rosen, R.; Baker, D.; Amironesei, R.; Prabhakaran, V.; and Denton, E. 2022 · 2022
Closest in time.
Priors in bayesian deep learning: A review
Fortuin, V. 2022 · 2022
Closest in time.
Jury learning: Integrating dissenting voices into machine learning models
Gordon, M. L.; Lam, M. S.; Park, J. S.; Patel, K.; Hancock, J.; Hashimoto, T.; and Bernstein, M. S. 2022 · 2022
Closest in time.
Pixmix: Dreamlike pictures comprehensively improve safety measures
Hendrycks, D.; Zou, A.; Mazeika, M.; Tang, L.; Li, B.; Song, D.; and Steinhardt, J. 2022 · 2022
Closest in time.
Going Beyond One-Hot Encoding in Classification: Can Human Uncertainty Improve Model Performance?
Koller, C.; Kauermann, G.; and Zhu, X. X. 2022 · 2022
Closest in time.
Eliciting Confidence for Improving Crowdsourced Audio Annotations
Méndez, A. E.; Cartwright, M.; Bello, J. P.; and Nov, O. 2022 · 2022
Closest in time.
Schmarje, L.; Grossmann, V.; Zelenka, C.; Dippel, S.; Kiko, R.; Oszust, M.; Pastell, M.; Stracke, J.; Valros, A.; Volkmann, N.; et al. 2022 · 2022
Closest in time.
Bayesian modeling of human–AI complementarity
Steyvers, M.; Tejeda, H.; Kerrigan, G.; and Smyth, P. 2022 · 2022
Closest in time.
Provably Improving Expert Predictions with Conformal Prediction
Straitouri, E.; Wang, L.; Okati, N.; and Rodriguez, M. G. 2022 · 2022
Closest in time.
Reliance on metrics is a fundamental challenge for AI
Thomas, R. L.; and Uminsky, D. 2022 · 2022
Closest in time.
Plex: Towards Reliability using Pretrained Large Model Extensions
Tran, D.; Liu, J.; Dusenberry, M. W.; Phan, D.; Collier, M.; Ren, J.; Han, K.; Wang, Z.; Mariet, Z.; Hu, H.; Band, N.; Rudner, T. G. J.; Singhal, K.; Nado, Z.; van Amersfoort, J.; Kirsch, A.; Jenatton, R.; Thain, N.; Yuan, H.; Buchanan, K.; Murphy, K.; Sculley, D.; Gal, Y.; Ghahramani, Z.; Snoek, J.; and Lakshminarayanan, B. 2022 · 2022
Closest in time.
Scaling and Disagreements: Bias, Noise, and Ambiguity
Uma, A.; Almanea, D.; and Poesio, M. 2022 · 2022
Closest in time.
Uncalibrated Models Can Improve Human-AI Collaboration
Vodrahalli, K.; Gerstenberg, T.; and Zou, J. 2022 · 2022
Closest in time.
To Aggregate or Not? Learning with Separate Noisy Labels
Wei, J.; Zhu, Z.; Luo, T.; Amid, E.; Kumar, A.; and Liu, Y. 2022 · 2022
Closest in time.