Fetching the paper…
Reading the bibliography…
Advances in machine learning are closely tied to the creation of datasets.
A guideline of selecting and reporting intraclass correlation coefficients for reliability research. j chiropr med. 2016; 15 (2): 155–63, 2000
TK Koo and MY Li · 2000
Earlier work this paper cites.
The unreasonable effectiveness of data
Alon Halevy, Peter Norvig, and Fernando Pereira · 2009
Earlier work this paper cites.
Optimal experimental design
Valerii Fedorov · 2010
Earlier work this paper cites.
Methodological guidelines for publishing government linked data
Boris Villazón-Terrazas, Luis M Vilches-Blázquez, Oscar Corcho, and Asunción Gómez-Pérez · 2011
Earlier work this paper cites.
Foundations of data quality management
Wenfei Fan and Floris Geerts · 2012
Earlier work this paper cites.
Methodological guidelines for publishing library data as linked data
Yusniel Hidalgo-Delgado, Reina Estrada-Nelson, Bin Xu, Boris Villazon-Terrazas, Amed Leiva-Mederos, and Andrés Tello · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
The dataset nutrition label: A framework to drive higher data quality standards
Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski · 2018
Earlier work this paper cites.
Artificial intelligence faces reproducibility crisis, 2018
Matthew Hutson · 2018
Earlier work this paper cites.
An empirical analysis of journal policy effectiveness for computational reproducibility
Victoria Stodden, Jennifer Seiler, and Zhaokun Ma · 2018
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Cited alongside, same era.
Balanced datasets are not enough: Estimating and mitigating gender bias in deep image representations
Tianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang, and Vicente Ordonez · 2019
Cited alongside, same era.
An end-to-end framework for integrating and publishing linked open government data
Rabeb Abida, Emna Hachicha Belghith, and Anthony Cleve · 2020
Cited alongside, same era.
Data readiness report, 2020
Shazia Afzal, Rajmohan C, Manish Kesarwani, Sameep Mehta, and Hima Patel · 2020
Cited alongside, same era.
Mt-adapted datasheets for datasets: Template and repository, 2020
Marta R. Costa-jussà, Roger Creus, Oriol Domingo, Albert Domínguez, Miquel Escobar, Cayetana López, Marina Garcia, and Margarita Geleta · 2020
Cited alongside, same era.
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna · 2021
Later among the works it cites.
Clustering word embeddings with self-organizing maps. application on laroseda – a large romanian sentiment data set
Anca Maria Tache, Mihaela Gaman, and Radu Tudor Ionescu · 2021
Later among the works it cites.
Jampatoisnli: A jamaican patois natural language inference dataset
Ruth-Ann Armstrong, John Hewitt, and Christopher Manning · 2022
Later among the works it cites.
Kasia S Chmielinski, Sarah Newman, Matt Taylor, Josh Joseph, Kemi Thomas, Jessica Yurkofsky, and Yue Chelsea Qiu · 2022
Later among the works it cites.
Tackling documentation debt: A survey on algorithmic fairness datasets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benjamin Haibe-Kains, George Alexandru Adam, Ahmed Hosny, Farnoosh Khodakarami, Massive Analysis Quality Control (MAQC) Society Board of Directors Shraddha Thakkar 35 Kusko Rebecca 36 Sansone Susanna-Assunta 37 Tong Weida 35 Wolfinger Russ D. 38 Mason Christopher E. 39 Jones Wendell 40 Dopazo Joaquin 41 Furlanello Cesare 42, Levi Waldron, Bo Wang, Chris McIntosh, Anna Goldenberg, Anshul Kundaje, et al · 2020
Cited alongside, same era.
An ethical highlighter for people-centric dataset creation, 2020
Margot Hanley, Apoorv Khandelwal, Hadar Averbuch-Elor, Noah Snavely, and Helen Nissenbaum · 2020
Cited alongside, same era.
Overcome obstacles to get to ai at scale
IBM · 2020
Cited alongside, same era.
Identifying and correcting label bias in machine learning
Heinrich Jiang and Ofir Nachum · 2020
Cited alongside, same era.
A guide for writing data statements for natural language processing, 2021
Emily M Bender, Batya Friedman, and Angelina McMillan-Major · 2021
Cited alongside, same era.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2021
Cited alongside, same era.
Reduced, reused and recycled: The life of a dataset in machine learning research
Bernard Koch, Emily Denton, Alex Hanna, and Jacob Gates Foster · 2021
Cited alongside, same era.
Alessandro Fabris, Stefano Messina, Gianmaria Silvello, and Gian Antonio Susto · 2022
Later among the works it cites.
Model card user studies
HuggingFace · 2022
Later among the works it cites.
Design guidelines for inclusive speaker verification evaluation datasets, 2022
Wiebke Toussaint Hutiri, Lauriane Gorce, and Aaron Yi Ding · 2022
Later among the works it cites.
Advances, challenges and opportunities in creating data for trustworthy ai
Weixin Liang, Girmaw Abebe Tadesse, Daniel Ho, L Fei-Fei, Matei Zaharia, Ce Zhang, and James Zou · 2022
Later among the works it cites.
Towards better Data Science to address racial bias and health equity
Elaine O Nsoesie and Sandro Galea · 2022
Later among the works it cites.
Data cards: Purposeful and transparent dataset documentation for responsible ai
Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson · 2022
Later among the works it cites.
Datasheet for subjective and objective quality assessment datasets, 2023
Nabajeet Barman, Yuriy Reznik, and Maria Martini · 2023
Later among the works it cites.
Huggingface dataset card guidebook, 2021
HuggingFace · 2023
Later among the works it cites.
Augmented datasheets for speech datasets and ethical decision-making
Orestis Papakyriakopoulos, Anna Seo Gyeong Choi, William Thong, Dora Zhao, Jerone Andrews, Rebecca Bourke, Alice Xiang, and Allison Koenecke · 2023
Later among the works it cites.