Fetching the paper…
Reading the bibliography…
Many machine learning projects for new application areas involve teams of humans who label data for a particular purpose, from hiring crowdworkers to the paper's authors labeling the data themselves.
Towards Traceability in Data Ecosystems using a Bill of Materials Model
Iain Barclay, Alun Preece, Ian Taylor, and Dinesh Verma. 2019 · 1904
Earlier work this paper cites.
ORES: Lowering Barriers with Participatory Machine Learning in Wikipedia
Aaron Halfaker and R Stuart Geiger. 2019 · 1909
Earlier work this paper cites.
Abigail Z. Jacobs and Hanna Wallach. 2019 · 1912
Earlier work this paper cites.
Inioluwa Deborah Raji and Jingying Yang. 2019 · 1912
Earlier work this paper cites.
Work with new electronic ‘brains’ opens field for army math experts
WD Mellin. 1957 · 1957
Earlier work this paper cites.
The discovery of grounded theory; strategies for qualitative research
Barney G Glaser, Anselm L Strauss, and Elizabeth Strutzel. 1968 · 1968
Earlier work this paper cites.
Laboratory Life: The Social Construction of Scientific Facts
Bruno Latour and Steve Woolgar. 1979 · 1979
Earlier work this paper cites.
Professional Vision
Charles Goodwin. 1994 · 1994
Earlier work this paper cites.
Python Library Reference
Guido van Rossum. 1995 · 1995
Earlier work this paper cites.
Seeing like a state: How certain schemes to improve the human condition have failed
James C. Scott. 1998 · 1998
Earlier work this paper cites.
Sorting Things Out: Classification and its Consequences
Geoffrey C Bowker and Susan Leigh Star. 1999 · 1999
Earlier work this paper cites.
Circulating Reference: Sampling the Soil in the Amazon Forest
Bruno Latour. 1999 · 1999
Earlier work this paper cites.
SciPy: Open source scientific tools for Python
Eric Jones, Travis Oliphant, Pearu Peterson, et al · 2001
Earlier work this paper cites.
Databases, Felons, and Voting: Bias and Partisanship of the Florida Felons List in the 2000 Elections
Guy Stuart. 2004 · 2004
Earlier work this paper cites.
Academic research record-keeping: Best practices for individuals, group leaders, and institutions
Alan A Schreier, Kenneth Wilson, and David Resnik. 2006 · 2006
Earlier work this paper cites.
Matplotlib: A 2D Graphics Environment
John D. Hunter. 2007 · 2007
Earlier work this paper cites.
IPython: a System for Interactive Scientific Computing
Fernando Pérez and Brian E. Granger. 2007 · 2007
Earlier work this paper cites.
Consolidated criteria for reporting qualitative research (COREQ): a 32-item checklist for interviews and focus groups
A. Tong, P. Sainsbury, and J. Craig. 2007 · 2007
Earlier work this paper cites.
ACE (Automatic Content Extraction) English annotation guidelines for entities version 6.6
Linguistic Data Consortium. 2008 · 2008
Earlier work this paper cites.
Annotation Tool Development for Large-Scale Corpus Creation Projects at the Linguistic Data Consortium.. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08) , Vol. 8
Kazuaki Maeda, Haejoong Lee, Shawn Medero, Julie Medero, Robert Parker, and Stephanie M. Strassel. 2008 · 2008
Earlier work this paper cites.
recaptcha: Human-based character recognition via web security measures
Luis Von Ahn, Benjamin Maurer, Colin McMillen, David Abraham, and Manuel Blum. 2008 · 2008
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction (2nd ed.)
Jerome Friedman, Trevor Hastie, and Robert Tibshirani. 2009 · 2009
Earlier work this paper cites.
Towards a ‘science’ of corpus annotation: a new methodological challenge for corpus linguistics
Eduard Hovy and Julia Lavid. 2010 · 2010
Earlier work this paper cites.
Data Structures for Statistical Computing in Python. In Proceedings of the 9th Python in Science Conference , Stéfan van der Walt and Jarrod Millman (Eds.). 51–56
Wes McKinney. 2010 · 2010
Cited alongside, same era.
The NumPy Array: A Structure for Efficient Numerical Computation
S. van der Walt, S. C. Colbert, and G. Varoquaux. 2011 · 2011
Cited alongside, same era.
The conundrum of sharing research data
Christine L Borgman. 2012 · 2012
Cited alongside, same era.
Eliminating spammers and ranking annotators for crowdsourced labeling tasks
Vikas C Raykar and Shipeng Yu. 2012 · 2012
Cited alongside, same era.
DMP Online and DMPTool: Different Strategies Towards a Shared Goal
Andrew Sallans and Martin Donnelly. 2012 · 2012
Cited alongside, same era.
GATE Teamware: a web-based, collaborative text annotation framework
How Robust Are Multirater Interrater Reliability Indices to Changes in Frequency Distribution?
David Quarfoot and Richard A. Levine. 2016 · 2016
Later among the works it cites.
Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (CHI ’17) . ACM, New York, NY, USA, 2334–2346
Joseph Chee Chang, Saleema Amershi, and Ece Kamar. 2017 · 2017
Later among the works it cites.
Teaching Integrity in Empirical Economics: The Pedagogy of Reproducible Science in Undergraduate Education
N. Medeiros and R.J. Ball. 2017 · 2017
Later among the works it cites.
Computational grounded theory: A methodological framework
Laura K Nelson. 2017 · 2017
Later among the works it cites.
Automatically tracking metadata and provenance of machine learning experiments. In Machine Learning Systems workshop at NIPS
Sebastian Schelter, Joos-Hendrik Böse, Johannes Kirschnick, Thoralf Klein, and Stephan Seufert. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kalina Bontcheva, Hamish Cunningham, Ian Roberts, Angus Roberts, Valentin Tablan, Niraj Aswani, and Genevieve Gorrell. 2013 · 2013
Cited alongside, same era.
An introduction to statistical learning
Gareth James, Daniela Witten, Trevor Hastie, and Robert Tibshirani. 2013 · 2013
Cited alongside, same era.
Analyzing media messages: Using quantitative content analysis in research
Daniel Riff, Stephen Lacy, and Frederick Fico. 2013 · 2013
Cited alongside, same era.
Measuring crowd truth: Disagreement metrics combined with worker behavior filters. In CrowdSem 2013 Workshop
Guillermo Soberón, Lora Aroyo, Chris Welty, Oana Inel, Hui Lin, and Manfred Overmeen. 2013 · 2013
Cited alongside, same era.
Open Science: One Term, Five Schools of Thought
Benedikt Fecher and Sascha Friesike. 2014 · 2014
Cited alongside, same era.
Ten Simple Rules for the Care and Feeding of Scientific Data
Alyssa Goodman, Alberto Pepe, Alexander W. Blocker, Christine L. Borgman, Kyle Cranmer, Merce Crosas, Rosanne Di Stefano, Yolanda Gil, Paul Groth, Margaret Hedstrom, David W. Hogg, Vinay Kashyap, Ashish Mahabal, Aneta Siemiginowska, and Aleksandra Slavkovic. 2014 · 2014
Cited alongside, same era.
On the choice of measures of reliability and validity in the content-analysis of texts
Anton Oleinik, Irina Popova, Svetlana Kirdina, and Tatyana Shatalova. 2014 · 2014
Cited alongside, same era.
Later among the works it cites.
Good enough practices in scientific computing
Greg Wilson, Jennifer Bryan, Karen Cranston, Justin Kitzes, Lex Nederbragt, and Tracy K. Teal. 2017 · 2017
Later among the works it cites.
Data statements for NLP: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Later among the works it cites.
Ethical and Socially-Aware Data Labels. In Annual International Symposium on Information Management and Big Data . Springer, 320–327
Elena Beretta, Antonio Vetrò, Bruno Lepri, and Juan Carlos De Martin. 2018 · 2018
Later among the works it cites.
Automating inequality: How high-tech tools profile, police, and punish the poor
Virginia Eubanks. 2018 · 2018
Later among the works it cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumeé III, and Kate Crawford. 2018 · 2018
Later among the works it cites.
Increasing Trust in AI Services through Supplier’s Declarations of Conformity
Michael Hind, Sameep Mehta, Aleksandra Mojsilovic, Ravi Nair, Karthikeyan Natesan Ramamurthy, Alexandra Olteanu, and Kush R Varshney. 2018 · 2018
Later among the works it cites.
The dataset nutrition label: A framework to drive higher data quality standards
Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski. 2018 · 2018
Later among the works it cites.
The Practice of Reproducible Research : Case Studies and Lessons from the Data-Intensive Sciences
Justin Kitzes, Daniel Turek, and Fatma Deniz. 2018 · 2018
Later among the works it cites.
doccano: Text Annotation Tool for Human
Hiroki Nakayama, Takahiro Kubo, Junya Kamura, Yasufumi Taniguchi, and Xu Liang. 2018 · 2018
Later among the works it cites.
Binder 2.0 - Reproducible, Interactive, Sharable Environments for Science at Scale. In Proceedings of the 17th Python in Science Conference , Fatih Akici, David Lippa, Dillon Niederhut, and M Pacer (Eds.). 113 – 120
Project Jupyter, Matthias Bussonnier, Jessica Forde, Jeremy Freeman, Brian Granger, Tim Head, Chris Holdgraf, Kyle Kelley, Gladys Nalvarte, Andrew Osheroff, M Pacer, Yuvi Panda, Fernando Perez, Benjamin Ragan Kelley, and Carol Willing. 2018 · 2018
Later among the works it cites.
Automating Large-scale Data Quality Verification
Sebastian Schelter, Dustin Lange, Philipp Schmidt, Meltem Celikel, Felix Biessmann, and Andreas Grafberger. 2018 · 2018
Later among the works it cites.
Responsible research with crowds: pay crowdworkers at least minimum wage
M Six Silberman, Bill Tomlinson, Rochelle LaPlante, Joel Ross, Lilly Irani, and Andrew Zaldivar. 2018 · 2018
Later among the works it cites.
Decision Provenance: Harnessing Data Flow for Accountable Systems
Jatinder Singh, Jennifer Cobbe, and Chris Norval. 2019 · 2018
Later among the works it cites.
Seaborn: Statistical Data Visualization Using Matplotlib
Michael Waskom, Olga Botvinnik, Drew O’Kane, Paul Hobson, Joel Ostblom, Saulius Lukauskas, David C Gemperline, Tom Augspurger, Yaroslav Halchenko, John B. Cole, Jordi Warmenhoven, Julian de Ruiter, Cameron Pye, Stephan Hoyer, Jake Vanderplas, Santi Villalba, Gero Kunter, Eric Quintero, Pete Bachant, Marcel Martin, Kyle Meyer, Alistair Miles, Yoav Ram, Thomas Brunner, Tal Yarkoni, Mike Lee Williams, Constantine Evans, Clark Fitzgerald, Brian, and Adel Qalieh. 2018 · 2018
Later among the works it cites.
Automated Management of Deep Learning Experiments. In Proceedings of the 3rd International Workshop on Data Management for End-to-End Machine Learning (DEEM’19) . ACM, New York, NY, USA, 8:1–8:4
Gharib Gharibi, Vijay Walunj, Rakan Alanazi, Sirisha Rella, and Yugyung Lee. 2019 · 2019
Closest in time.
Reliability and Inter-rater Reliability in Qualitative Research: Norms and Guidelines for CSCW and HCI Practice
Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019 · 2019
Closest in time.
Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency . ACM, 220–229
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019 · 2019
Closest in time.