Fetching the paper…
Reading the bibliography…
Labeling training data is increasingly the largest bottleneck in deploying machine learning systems.
Probability of error of some adaptive pattern-recognition machines
H. J. Scudder · 1965
Earlier work this paper cites.
Learning with a probabilistic teacher
A. K. Agrawala · 1970
Earlier work this paper cites.
Maximum likelihood estimation of observer error-rates using the EM algorithm
A. P. Dawid and A. M. Skene · 1979
Earlier work this paper cites.
Automatic acquisition of hyponyms from large text corpora
M. A. Hearst · 1992
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
A. Blum and T. Mitchell · 1998
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G. E. Hinton · 2002
Earlier work this paper cites.
Framewise phoneme classification with bidirectional LSTM and other neural network architectures
A. Graves and J. Schmidhuber · 2005
Earlier work this paper cites.
Learning to extract relations from the Web using minimal supervision
R. C. Bunescu and R. J. Mooney · 2007
Earlier work this paper cites.
Modeling annotators: A generative approach to learning from annotator rationales
O. F. Zaidan and J. Eisner · 2008
Earlier work this paper cites.
Semi-Supervised Learning
O. Chapelle, B. Schölkopf, and A. Zien, editors · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning from measurements in exponential families
P. Liang, M. I. Jordan, and D. Klein · 2009
Earlier work this paper cites.
Distant supervision for relation extraction without labeled data
M. Mintz, S. Bills, R. Snow, and D. Jurafsky · 2009
Earlier work this paper cites.
Generalized expectation criteria for semi-supervised learning with weakly labeled data
G. S. Mann and A. McCallum · 2010
Earlier work this paper cites.
A survey on transfer learning
S. J. Pan and Q. Yang · 2010
Earlier work this paper cites.
Modeling relations and their mentions without labeled text
S. Riedel, L. Yao, and A. McCallum · 2010
Earlier work this paper cites.
Knowledge-based weak supervision for information extraction of overlapping relations
R. Hoffmann, C. Zhang, X. Ling, L. Zettlemoyer, and D. S. Weld · 2011
Earlier work this paper cites.
Human computation: A survey and taxonomy of a growing field
A. J. Quinn and B. B. Bederson · 2011
Earlier work this paper cites.
Finding a “kneedle” in a haystack: Detecting knee points in system behavior
V. Satopaa, J. Albrecht, D. Irwin, and B. Raghavan · 2011
Cited alongside, same era.
A survey of crowdsourcing systems
M.-C. Yuen, I. King, and K.-S. Leung · 2011
Cited alongside, same era.
Pattern learning for relation extraction with a hierarchical topic model
E. Alfonseca, K. Filippova, J.-Y. Delort, and G. Garrido · 2012
Cited alongside, same era.
Active Learning
B. Settles · 2012
Cited alongside, same era.
Reducing wrong labels in distant supervision for relation extraction
S. Takamatsu, I. Sato, and H. Nakagawa · 2012
Cited alongside, same era.
A Bayesian approach to discovering truth from conflicting sources for data integration
B. Zhao, B. I. Rubinstein, J. Gemmell, and J. Han · 2012
Cited alongside, same era.
The Mobilize center: an NIH big data to knowledge center to advance human movement research and improve mobility
J. P. Ku, J. L. Hicks, T. Hastie, J. Leskovec, C. Ré, and S. L. Delp · 2015
Later among the works it cites.
A survey on truth discovery
Y. Li, J. Gao, C. Meng, Q. Li, L. Su, B. Zhao, W. Fan, and J. Han · 2015
Later among the works it cites.
Overview of the BioCreative V chemical disease relation (CDR) task
C.-H. Wei, Y. Peng, R. Leaman, D. A. P., C. J. Mattingly, J. Li, T. Wiegers, and Z. Lu · 2015
Later among the works it cites.
TensorFlow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Later among the works it cites.
The MetaCyc database of metabolic pathways and enzymes and the BioCyc collection of pathway/genome databases
R. Caspi, R. Billington, L. Ferrer, H. Foerster, C. A. Fulcher, I. M. Keseler, A. Kothari, M. Krummenacker, M. Latendresse, L. A. Mueller, Q. Ong, S. Paley, P. Subhraveti, D. S. Weaver, and P. D. Karp · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aggregating crowdsourced binary ratings
N. Dalvi, A. Dasgupta, R. Kumar, and V. Rastogi · 2013
Cited alongside, same era.
A CTD–Pfizer collaboration: Manual curation of 88,000 scientific articles text mined for drug–disease and drug–phenotype interactions
A. P. Davis et al · 2013
Cited alongside, same era.
Error rate analysis of labeling by crowdsourcing
H. Li, B. Yu, and D. Zhou · 2013
Cited alongside, same era.
Combining generative and discriminative model scores for distant supervision
B. Roth and D. Klakow · 2013
Cited alongside, same era.
Improved pattern learning for bootstrapped entity extraction
S. Gupta and C. D. Manning · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
What do a million news articles look like?
D. Corney, D. Albakour, M. Martinez, and S. Moussa · 2016
Later among the works it cites.
Google’s hand-fed AI now gives answers, not just search results, 2016
C. Metz · 2016
Later among the works it cites.
The comparative toxicogenomics database: update 2017
D. A. P., C. J. Grondin, R. J. Johnson, D. Sciaky, B. L. King, R. McMorran, J. Wiegers, T. Wiegers, and C. J. Mattingly · 2016
Later among the works it cites.
Data programming: Creating large training sets, quickly
A. Ratner, C. De Sa, S. Wu, D. Selsam, and C. Ré · 2016
Later among the works it cites.
Spectral methods meet EM: A provably optimal algorithm for crowdsourcing
Y. Zhang, X. Chen, D. Zhou, and M. I. Jordan · 2016
Later among the works it cites.
Technical report, International Data Corporation, 2017
Worldwide semiannual cognitive/artificial intelligence systems spending guide · 2017
Closest in time.
Learning the structure of generative models without labeled data
S. H. Bach, B. He, A. Ratner, and C. Ré · 2017
Closest in time.
Baidu’s Andrew Ng on the future of artificial intelligence, 2017
L. Eadicicco · 2017
Closest in time.
HoloClean: Holistic data repairs with probabilistic inference
T. Rekatsinas, X. Chu, I. F. Ilyas, and C. Ré · 2017
Closest in time.
SLiMFast: Guaranteed results for data fusion and source reliability
T. Rekatsinas, M. Joglekar, H. Garcia-Molina, A. Parameswaran, and C. Ré · 2017
Closest in time.
Label-free supervision of neural networks with physics and other domain knowledge
R. Stewart and S. Ermon · 2017
Closest in time.
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Closest in time.
DeepDive: Declarative knowledge base construction
C. Zhang, C. Ré, M. Cafarella, C. De Sa, A. Ratner, J. Shin, F. Wang, and S. Wu · 2017
Closest in time.