Fetching the paper…
Reading the bibliography…
Recent literature has underscored the importance of dataset documentation work for machine learning, and part of this work involves addressing "documentation debt" for datasets that have been used widely but documented sparsely.
Prodigy
Edward Mullen. 2013 · 2013
Earlier work this paper cites.
The Cop And The Girl From The Coffee Shop
Terry Towers. 2013 · 2013
Earlier work this paper cites.
Smashwords Year in Review 2014 and Plans for 2015
Mark Coker. 2014 · 2014
Earlier work this paper cites.
The Library of Congress by the Numbers in 2013
Audrey Fischer. 2014 · 2014
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016 · 2016
Earlier work this paper cites.
Google swallows 11,000 novels to improve AI’s conversation
Richard Lea. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
Building Consentful Tech
Una Lee and Dann Toliver. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
List of dirty naughty obscene and otherwise bad words
Jacob Emerick. 2018 · 2018
Earlier work this paper cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2018 · 2018
Earlier work this paper cites.
The dataset nutrition label: A framework to drive higher data quality standards
Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski. 2018 · 2018
Cited alongside, same era.
How a cabal of romance writers cashed in on Amazon Kindle Unlimited
Sarah Jeong. 2018 · 2018
Cited alongside, same era.
Homemade BookCorpus
Sosuke Kobayashi. 2018 · 2018
Cited alongside, same era.
The Authors Who Love Amazon
Alana Semuels. 2018 · 2018
Cited alongside, same era.
Mind the gap: A balanced corpus of gendered ambiguous pronouns
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge. 2018 · 2018
Cited alongside, same era.
Plagiarism, “book-stuffing”, clickfarms … the rotten side of self-publishing
Alison Flood. 2019 · 2019
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Smashwords 2020 Year in Review and 2021 Preview
Mark Coker. 2020 · 2020
Later among the works it cites.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Later among the works it cites.
The Dataset Nutrition Label
Sarah Holland, Ahmed Hosny, and Sarah Newman. 2020 · 2020
Later among the works it cites.
Google: BERT now used on almost every English query
Barry Schwartz. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 2019
Cited alongside, same era.
Mithralabel: Flexible dataset nutritional labels for responsible data science. In
Chenkai Sun, Abolfazl Asudeh, HV Jagadish, Bill Howe, and Julia Stoyanovich. 2019 · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 2019
Cited alongside, same era.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Closest in time.
Introducing the NeurIPS 2021 Paper Checklist
Alina Beygelzimer, Yann Dauphin, Percy Liang, and Jennifer Wortman Vaughan. 2021 · 2021
Closest in time.
BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation. In
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021 · 2021
Closest in time.
Curtis G. Northcutt, Anish Athalye, and Jonas Mueller. 2021 · 2021
Closest in time.
An AI is training counselors to deal with teens in crisis
Abby Ohlheiser. 2021 · 2021
Closest in time.
Lawyer and Judicial Competency in the Era of Artificial Intelligence: Ethical Requirements for Documenting Datasets and Machine Learning Models
Mark L. Shope. 2021 · 2021
Closest in time.
Announcing the NeurIPS 2021 Datasets and Benchmarks Track
Joaquin Vanschoren and Serena Yeung. 2021 · 2021
Closest in time.