Fetching the paper…
Reading the bibliography…
NLP community is currently investing a lot more research and resources into development of deep learning models than training data.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a · 1905
Earlier work this paper cites.
The meaning-frequency relationship of words
George Kingsley Zipf. 1945 · 1945
Earlier work this paper cites.
Vom Weltbild der deutschen Sprache
Leo Weisgerber. 1953 · 1953
Earlier work this paper cites.
Acquiring a Single New Word
Susan Carey and Elsa Bartlett. 1978 · 1978
Earlier work this paper cites.
Explanation in Linguistics. The Logical Problem of Language Acquisition
Norbert Hornstein and David Lightfoot. 1985 · 1985
Earlier work this paper cites.
100 million words of English: The British National Corpus (BNC)
Geoffrey Neil Leech. 1992 · 1992
Earlier work this paper cites.
Izbrannyje Trudy , volume 2
Yu.D. Apresyan. 1995 · 1995
Earlier work this paper cites.
Acquiring vocabulary through reading: Effects of frequency and contextual richness
Rick Zahar, Tom Cobb, and Nina Spada. 2001 · 2001
Earlier work this paper cites.
Proceedings of the ACL-02 Workshop on Effective Tools and Methodologies for Teaching Natural Language Processing and Computational Linguistics
Dragomir Radev and Chris Brew. 2002 · 2002
Earlier work this paper cites.
Proceedings of the Second ACL Workshop on Effective Tools and Methodologies for Teaching NLP and CL
Chris Brew and Dragomir Radev, editors. 2005 · 2005
Earlier work this paper cites.
How Can We Accelerate Progress Towards Human-like Linguistic Generalization?
Tal Linzen. 2020 · 2005
Earlier work this paper cites.
From Usage to Grammar: The Mind’s Response to Repetition
Joan Bybee. 2006 · 2006
Earlier work this paper cites.
Universal linguistic inductive biases via meta-learning
R. Thomas McCoy, Erin Grant, Paul Smolensky, Thomas L. Griffiths, and Tal Linzen. 2020 · 2006
Earlier work this paper cites.
Children’s first language acquistion from a usage-based perspective
Elena Lieven and Michael Tomasello. 2008 · 2008
Earlier work this paper cites.
Proceedings of the Third Workshop on Issues in Teaching Computational Linguistics
Martha Palmer, Chris Brew, and Fei Xia, editors. 2008 · 2008
Earlier work this paper cites.
Adversarial Watermarking Transformer: Towards Tracing Text Provenance with Data Hiding
Sahar Abdelnabi and Mario Fritz. 2021 · 2009
Earlier work this paper cites.
The Unreasonable Effectiveness of Data
A. Halevy, P. Norvig, and F. Pereira. 2009 · 2009
Earlier work this paper cites.
Learning to use words: Event-related potentials index single-shot contextual word learning
Arielle Borovsky, Marta Kutas, and Jeff Elman. 2010 · 2010
Earlier work this paper cites.
Watermarking the Outputs of Structured Prediction with an application in Statistical Machine Translation
Ashish Venugopal, Jakob Uszkoreit, David Talbot, Franz Och, and Juri Ganitkevitch. 2011 · 2011
Earlier work this paper cites.
When Do You Need Billions of Words of Pretraining Data?
Yian Zhang, Alex Warstadt, Haau-Sing Li, and Samuel R. Bowman. 2020 · 2011
Earlier work this paper cites.
Extracting Training Data from Large Language Models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and Colin Raffel. 2020 · 2012
Earlier work this paper cites.
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna. 2020 · 2012
Earlier work this paper cites.
Proceedings of the Fourth Workshop on Teaching NLP and CL
Ivan Derzhanski and Dragomir Radev, editors. 2013 · 2013
Earlier work this paper cites.
Aspects of the Theory of Syntax
Noam Chomsky. 2014 · 2014
Earlier work this paper cites.
The ubiquity of frequency effects in first language acquisition*
Ben Ambridge, Evan Kidd, Caroline F. Rowland, and Anna L. Theakston. 2015 · 2015
Cited alongside, same era.
Frequency effects in grammar
Holger Diessel and Martin Hilpert. 2016, May 09 · 2016
Cited alongside, same era.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. 2017 · 2017
Cited alongside, same era.
Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science
Emily M. Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
Annotation Artifacts in Natural Language Inference Data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Later among the works it cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2020 · 2020
Later among the works it cites.
Social Biases in NLP Models as Barriers for Persons with Disabilities
Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl. 2020 · 2020
Later among the works it cites.
Lessons from archives: Strategies for collecting sociocultural data in machine learning
Eun Seo Jo and Timnit Gebru. 2020 · 2020
Later among the works it cites.
GPT-3 Bot Posed as a Human on AskReddit for a Week
Philip. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
According to Chomsky, words such as ’book’ and ’carburetor’ are genetically determined
Chris Knight. 2018 · 2018
Cited alongside, same era.
One paper, nine reviews
Rachel Bawden. 2019 · 2019
Cited alongside, same era.
Racial Bias in Hate Speech and Abusive Language Detection Datasets
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Posing Fair Generalization Tasks for Natural Language Inference
Atticus Geiger, Ignacio Cases, Lauri Karttunen, and Christopher Potts. 2019 · 2019
Cited alongside, same era.
Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
What Can We Do to Improve Peer Review in NLP?
Anna Rogers and Isabelle Augenstein. 2020 · 2020
Later among the works it cites.
Jack Bandy and Nicholas Vincent. 2021 · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Closest in time.
Fair ML Tools Require Problematic ML Models
Jacob Buckman. 2021 · 2021
Closest in time.
Google pledges changes to research oversight after internal revolt
Jeffrey Dastin Dave, Paresh. 2021 · 2021
Closest in time.
Documenting the English Colossal Clean Crawled Corpus
Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner. 2021 · 2021
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Closest in time.
Competency Problems: On Finding and Removing Artifacts in Language Data
Matt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters, Alexis Ross, Sameer Singh, and Noah Smith. 2021 · 2021
Closest in time.
A Criticism of Stochastic Parrots
Yoav Goldberg. 2021 · 2021
Closest in time.
Proceedings of the Fifth Workshop on Teaching NLP
David Jurgens, Varada Kolhatkar, Lucy Li, Margot Mieskes, and Ted Pedersen, editors. 2021 · 2021
Closest in time.
Question and Answer Test-Train Overlap in Open-Domain Question Answering Datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021 · 2021
Closest in time.
The Slodderwetenschap (Sloppy Science) of Stochastic Parrots
Michael Lissack. 2021 · 2021
Closest in time.
GPT-3 Powers the Next Generation of Apps
Ashley Pilipiszyn. 2021 · 2021
Closest in time.
"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021b · 2021
Closest in time.
Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021 · 2021
Closest in time.
On stochastic parrots
Suresh Venkatasubramanian. 2021 · 2021
Closest in time.
Adversarial Examples for Evaluating Reading Comprehension Systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.