Fetching the paper…
Reading the bibliography…
The construction of most supervised learning datasets revolves around collecting multiple labels for each instance, then aggregating the labels to form a type of "gold-standard".
Maximum likelihood estimation of observer error-rates using the EM algorithm
A. Dawid and A. Skene · 1979
Earlier work this paper cites.
Matrix Analysis
R. A. Horn and C. R. Johnson · 1985
Earlier work this paper cites.
Empirical likelihood ratio confidence regions
A. Owen · 1990
Earlier work this paper cites.
Probability in Banach Spaces
M. Ledoux and M. Talagrand · 1991
Earlier work this paper cites.
Building a large annotated corpus of English: the Penn Treebank
M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz · 1994
Earlier work this paper cites.
The Nature of Statistical Learning Theory
V. Vapnik · 1995
Earlier work this paper cites.
Weak Convergence and Empirical Processes: With Applications to Statistics
A. W. van der Vaart and J. A. Wellner · 1996
Earlier work this paper cites.
Efficient and Adaptive Estimation for Semiparametric Models
P. Bickel, C. A. J. Klaassen, Y. Ritov, and J. Wellner · 1998
Earlier work this paper cites.
Asymptotic Statistics
A. W. van der Vaart · 1998
Earlier work this paper cites.
Smooth discrimination analysis
E. Mammen and A. B. Tsybakov · 1999
Earlier work this paper cites.
Asymptotics in Statistics: Some Basic Concepts
L. Le Cam and G. L. Yang · 2000
Earlier work this paper cites.
RCV1: A new benchmark collection for text categorization research
D. Lewis, Y. Yang, T. Rose, and F. Li · 2004
Earlier work this paper cites.
Boosting as a regularized path to a maximum margin classifier
S. Rosset, J. Zhu, and T. Hastie · 2004
Earlier work this paper cites.
Theory of classification: a survey of some recent advances
S. Boucheron, O. Bousquet, and G. Lugosi · 2005
Earlier work this paper cites.
Convexity, classification, and risk bounds
P. L. Bartlett, M. I. Jordan, and J. McAuliffe · 2006
Earlier work this paper cites.
The rise of crowdsourcing
J. Howe · 2006
Earlier work this paper cites.
UCI machine learning repository, 2007
A. Asuncion and D. J. Newman · 2007
Earlier work this paper cites.
ImageNet: a large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei · 2009
Cited alongside, same era.
The Elements of Statistical Learning
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Cited alongside, same era.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Cited alongside, same era.
Lectures on Stochastic Programming: Modeling and Theory
A. Shapiro, D. Dentcheva, and A. Ruszczyński · 2009
Cited alongside, same era.
Whose vote should count more: Optimal integration of labels from labelers of unknown expertise
J. Whitehill, T. Wu, J. Bergsma, J. Movellan, and P. Ruvolo · 2009
Cited alongside, same era.
Normal Approximation by Stein’s method
L. H. Chen, L. Goldstein, and Q.-M. Shao · 2010
Cited alongside, same era.
Learning from untrusted data
M. Charikar, J. Steinhardt, and G. Valiant · 2017
Later among the works it cites.
50 years of data science
D. L. Donoho · 2017
Later among the works it cites.
Snorkel: rapid training data creation with weak supervision
A. Ratner, S. H. Bach, H. Ehrenberg, J. Fries, S. Wu, and C. Ré · 2017
Later among the works it cites.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Later among the works it cites.
A modern maximum-likelihood theory for high-dimensional logistic regression
E. Candès and P. Sur · 2019
Later among the works it cites.
Human uncertainty makes classification more robust
J. C. Peterson, R. M. Battleday, T. L. Griffiths, and O. Russakovsky · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning from crowds
V. C. Raykar, S. Yu, L. H. Zhao, G. H. Valadez, C. Florin, L. Bogoni, and L. Moy · 2010
Cited alongside, same era.
The multidimensional wisdom of crowds
P. Welinder, S. Branson, P. Perona, and S. Belongie · 2010
Cited alongside, same era.
Robust 1-bit compressed sensing and sparse logistic regression: A convex programming approach
Y. Plan and R. Vershynin · 2013
Cited alongside, same era.
Budget-optimal task allocation for reliable crowdsourcing systems
D. Karger, S. Oh, and D. Shah · 2014
Cited alongside, same era.
On the absolute constants in the Berry-Esseen-type inequalities
I. Shevtsova · 2014
Cited alongside, same era.
Qualitatively characterizing neural network optimization problems
I. Goodfellow, O. Vinyals, and A. Saxe · 2015
Cited alongside, same era.
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
M. J. Wainwright · 2019
Later among the works it cites.
The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression
E. Candès and P. Sur · 2020
Later among the works it cites.
Learning from imperfect annotations
E. A. Platanios, M. Al-Shedivat, E. Xing, and T. Mitchell · 2020
Later among the works it cites.
NeurIPS 2021 datasets and benchmarks track
A. Beygelzimer, P. Liang, J. Wortman Vaughan, and Y. Dauphin · 2021
Later among the works it cites.
Mathematical Foundations of Infinite-Dimensional Statistical Models
E. Giné and R. Nickl · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Later among the works it cites.
Patterns, Predictions, and Actions: A story about machine learning
M. Hardt and B. Recht · 2022
Closest in time.
LAION-5B: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev · 2022
Closest in time.
DataComp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, E. Orgad, R. Entezari, G. Daras, S. Pratt, V. Ramanujan, Y. Bitton, K. Marathe, S. Mussmann, R. Vencu, M. Cherti, R. Krishna, P. W. Koh, O. Saukh, A. Ratner, S. Song, H. Hajishirzi, A. Farhadi, R. Beaumont, S. Oh, A. Dimakis, J. Jitsev, Y. Carmon, V. Shankar, and L. Schmidt · 2023
Closest in time.