Fetching the paper…
Reading the bibliography…
Large language models' (LLMs) abilities are drawn from their pretraining data, and model development begins with data curation.
Logic and conversation
H.P. Grice. 1975 · 1975
Earlier work this paper cites.
Harms of gender exclusivity and challenges in non-binary representation in language technologies
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff Phillips, and Kai-Wei Chang. 2021 · 1994
Earlier work this paper cites.
Understanding work using the occupational information network (o* net): Implications for practice and research
Norman G Peterson, Michael D Mumford, Walter C Borman, P Richard Jeanneret, Edwin A Fleishman, Kerry Y Levin, Michael A Campion, Melinda S Mayfield, Frederick P Morgeson, Kenneth Pearlman, et al. 2001 · 2001
Earlier work this paper cites.
Studying the amateur artist: A perspective on disguising data collected in human subjects research on the internet
Amy Bruckman. 2002 · 2002
Earlier work this paper cites.
Steps to creating a content strategy for your organization
Ellen Wagner. 2002 · 2002
Earlier work this paper cites.
Registers of Language , chapter 2. John Wiley & Sons, Ltd
Asif Agha. 2005 · 2005
Earlier work this paper cites.
English as a lingua franca
Barbara Seidlhofer. 2005 · 2005
Earlier work this paper cites.
Introduction to Information Retrieval
Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze. 2008 · 2008
Earlier work this paper cites.
English on the internet and a ‘post-varieties’ approach to language
Philip Seargeant and Caroline Tagg. 2011 · 2011
Earlier work this paper cites.
Who creates content?
Grant Blank. 2013 · 2013
Earlier work this paper cites.
What to do about bad language on the internet
Jacob Eisenstein. 2013 · 2013
Earlier work this paper cites.
Why are more African countries adopting english as an official language
Patrick Plonski, Asratie Teferra, and Rachel Brady. 2013 · 2013
Earlier work this paper cites.
Compact Language Detector 2
Dick Sites. 2013 · 2013
Earlier work this paper cites.
langdetect
Nakatani Shuyo. 2014 · 2014
Earlier work this paper cites.
“The sum of all human knowledge”: A systematic review of scholarly research on the content of Wikipedia
Mostafa Mesgari, Chitu Okoli, Mohamad Mehdi, Finn Årup Nielsen, and Arto Lanamäki. 2015 · 2015
Earlier work this paper cites.
Computational sociolinguistics: A survey
Dong Nguyen, A Seza Doğruöz, Carolyn P Rosé, and Franciska De Jong. 2016 · 2016
Earlier work this paper cites.
Building and evaluating web corpora representing national varieties of English
Paul Cook and Laurel J Brinton. 2017 · 2017
Earlier work this paper cites.
Community identity and user engagement in a multi-community landscape
Justine Zhang, William Hamilton, Cristian Danescu-Niculescu-Mizil, Dan Jurafsky, and Jure Leskovec. 2017 · 2017
Earlier work this paper cites.
India and the Anglosphere: Race, identity and hierarchy in international relations
Alexander Davis. 2018 · 2018
Earlier work this paper cites.
A fast, compact, accurate model for language identification of codemixed text
Yuan Zhang, Jason Riesa, Daniel Gillick, Anton Bakalov, Jason Baldridge, and David Weiss. 2018 · 2018
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019 · 2019
Earlier work this paper cites.
Improving fairness in machine learning systems: What do industry practitioners need?
Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé, Miro Dudik, and Hanna Wallach. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Toward gender-inclusive coreference resolution
Yang Trista Cao and Hal Daumé III. 2020 · 2020
Earlier work this paper cites.
Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. 2020 · 2020
Cited alongside, same era.
The Pile: An 800GB dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020 · 2020
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Cited alongside, same era.
Compact Language Detector v3 (CLD3)
Alex Salcianu, Andy Golding, Anton Bakalov, Chris Alberti, Daniel Andor, David Weiss, Emily Pitler, Greg Coppola, Jason Riesa, Kuzman Ganchev, Michael Ringgaard, Nan Hua, Ryan McDonald, Slav Petrov, Stefan Istrate, and Terry Koo. 2020 · 2020
Which humans?
Mohammad Atari, Mona J Xue, Peter S Park, Damián E Blasi, and Joseph Henrich. 2023 · 2023
Later among the works it cites.
Megawika: Millions of reports and their sources across 50 diverse languages
Samuel Barham, Orion Weller, Michelle Yuan, Kenton Murray, Mahsa Yarmohammadi, Zhengping Jiang, Siddharth Vashishtha, Alexander Martin, Anqi Liu, Aaron Steven White, Jordan Boyd-Graber, and Benjamin Van Durme. 2023 · 2023
Later among the works it cites.
Speak, memory: An archaeology of books known to ChatGPT/GPT-4
Kent Chang, Mackenzie Cramer, Sandeep Soni, and David Bamman. 2023 · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023 · 2023
Later among the works it cites.
RedPajama: an open dataset for training large language models
Together Computer. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The Anglosphere: A genealogy of a racialized identity in international relations
Srdjan Vucetic. 2020 · 2020
Cited alongside, same era.
CCNet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020 · 2020
Cited alongside, same era.
Phans, stans and cishets: Self-presentation effects on content propagation in tumblr
Michael Miller Yoder, Qinlan Shen, Yansen Wang, Alex Coda, Yunseok Jang, Yale Song, Kapil Thadani, and Carolyn P. Rosé. 2020 · 2020
Cited alongside, same era.
What we can’t measure, we can’t understand: Challenges to demographic data procurement in the pursuit of fairness
McKane Andrus, Elena Spitzer, Jeffrey Brown, and Alice Xiang. 2021 · 2021
Cited alongside, same era.
Documenting large webtext corpora: A case study on the Colossal Clean Crawled Corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021 · 2021
Cited alongside, same era.
An empirical exploration in quality filtering of text data
Leo Gao. 2021 · 2021
Cited alongside, same era.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021 · 2021
Cited alongside, same era.
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023 · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy Lillicrap, Angeliki Lazaridou, …, Demis Hassabis, Koray Kavukcuoglu, Jeffrey Dean, and Oriol Vinyals. 2023 · 2023
Later among the works it cites.
Demystifying prompts in language models via perplexity estimation
Hila Gonen, Srini Iyer, Terra Blevins, Noah Smith, and Luke Zettlemoyer. 2023 · 2023
Later among the works it cites.
Scaling expert language models with unsupervised domain discovery
Suchin Gururangan, Margaret Li, Mike Lewis, Weijia Shi, Tim Althoff, Noah A Smith, and Luke Zettlemoyer. 2023 · 2023
Later among the works it cites.
Mordecai 3: A neural geoparser and event geocoder
Andrew Halterman. 2023 · 2023
Later among the works it cites.
Large language models struggle to learn long-tail knowledge
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023 · 2023
Later among the works it cites.
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, et al. 2023 · 2023
Later among the works it cites.
Words as gatekeepers: Measuring discipline-specific terms and meanings in scholarly publications
Li Lucy, Jesse Dodge, David Bamman, and Katherine Keith. 2023 · 2023
Later among the works it cites.
Navid Madani, Rabiraj Bandyopadhyay, Briony Swire-Thompson, Michael Miller Yoder, and Kenneth Joseph. 2023 · 2023
Later among the works it cites.
When less is more: Investigating data pruning for pretraining LLMs at scale
Max Marion, Ahmet Üstün, Luiza Pozzobon, Alex Wang, Marzieh Fadaee, and Sara Hooker. 2023 · 2023
Later among the works it cites.
Scaling data-constrained language models
Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
“I’m fully who I am”: Towards centering transgender and non-binary voices to measure biases in open language generation
Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023 · 2023
Later among the works it cites.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023 · 2023
Later among the works it cites.
Characterizing and identifying socially shared self-descriptions in product reviews
Lu Sun, F Maxwell Harper, Chia-Jung Lee, Vanessa Murdock, and Barbara Poblete. 2023 · 2023
Later among the works it cites.
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang. 2023 · 2023
Later among the works it cites.
Who’s in and who’s out? a case study of multimodal clip-filtering in datacomp
Rachel Hong, William Agnew, Tadayoshi Kohno, and Jamie Morgenstern. 2024 · 2024
Closest in time.
Dolma: an open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo. 2024 · 2024
Closest in time.
Richer countries and richer representations
Kaitlyn Zhou, Kawin Ethayarajh, and Dan Jurafsky. 2022 · 2085
Closest in time.