Fetching the paper…
Reading the bibliography…
General-purpose artificial intelligence (AI) systems are built on massive swathes of public web data, assembled into corpora such as C4, RefinedWeb, and Dolma.
A large-scale study of robots. txt
Yang Sun, Ziming Zhuang, and C Lee Giles · 2007
Earlier work this paper cites.
Classification of web robots: an empirical study based on over one billion requests
Junsup Lee, Sungdeok Cha, Dongkun Lee, and Hyungkyu Lee · 2009
Earlier work this paper cites.
Analysis of web logs: challenges and findings
Maria Carla Calzarossa and Luisa Massari · 2010
Earlier work this paper cites.
statsmodels: Econometric and statistical modeling with python
Skipper Seabold and Josef Perktold · 2010
Earlier work this paper cites.
A Billion Wicked Thoughts: What the World’s Largest Experiment Reveals about Human Desire
Ogi Ogas and Sai Gaddam · 2011
Earlier work this paper cites.
Temporal analysis of crawling activities of commercial web robots
Maria Carla Calzarossa and Luisa Massari · 2012
Earlier work this paper cites.
Web robot detection based on pattern-matching technique
Shinil Kwon, Young-Gab Kim, and Sungdeok Cha · 2012
Earlier work this paper cites.
Accountable algorithms
Joshua Alexander Kroll · 2015
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith · 2017
Earlier work this paper cites.
Auditing algorithms for discrimination
Pauline T Kim · 2017
Earlier work this paper cites.
Shreya Shankar, Yoni Halpern, Eric Breck, James Atwood, Jimbo Wilson, and D Sculley · 2017
Earlier work this paper cites.
pmdarima: Arima estimators for Python, 2017–
Taylor G. Smith et al · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman · 2018
Earlier work this paper cites.
Twenty years of web scraping and the computer fraud and abuse act
Andrew Sellars · 2018
Earlier work this paper cites.
Does object recognition work for everyone?
Terrance De Vries, Ishan Misra, Changhan Wang, and Laurens Van der Maaten · 2019
Earlier work this paper cites.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
Moments in time dataset: one million videos for event understanding
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
Fair learning
Mark A Lemley and Bryan Casey · 2020
Earlier work this paper cites.
Identifying sensitive urls at web-scale
Srdjan Matic, Costas Iordanou, Georgios Smaragdakis, and Nikolaos Laoutaris · 2020
Earlier work this paper cites.
The text file that runs the internet
David Pierce · 2020
Earlier work this paper cites.
Jack Bandy and Nicholas Vincent · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Earlier work this paper cites.
Quality at a glance: An audit of web-crawled multilingual datasets
Isaac Caswell, Julia Kreutzer, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, et al · 2021
Earlier work this paper cites.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, et al · 2021
Earlier work this paper cites.
The problem of zombie datasets: A framework for deprecating datasets
Frances Corry, Hamsini Sridharan, Alexandra Sasha Luccioni, Mike Ananny, Jason Schultz, and Kate Crawford · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Earlier work this paper cites.
Memorization vs. generalization : Quantifying data leakage in NLP performance evaluation
Aparna Elangovan, Jiayuan He, and Karin Verspoor · 2021
Earlier work this paper cites.
Datahunter: A system for finding datasets based on scientific problem descriptions
Michael Färber and Ann-Kathrin Leisinger · 2021
Earlier work this paper cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford · 2021
Earlier work this paper cites.
Towards accountability for machine learning datasets: Practices from software engineering and infrastructure
Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell · 2021
Earlier work this paper cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, et al · 2021
Earlier work this paper cites.
What’s in the box? an analysis of undesirable content in the common crawl corpus
Alexandra Sasha Luccioni and Joseph D Viviano · 2021
Earlier work this paper cites.
Understanding gender and racial disparities in image recognition models
Rohan Mahadev and Anindya Chakravarti · 2021
Earlier work this paper cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna · 2021
Earlier work this paper cites.
Changing the world by changing the data
Anna Rogers · 2021
Earlier work this paper cites.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models, 2021
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel · 2021
Cited alongside, same era.
Challenges in detoxifying language models
Silo language models: Isolating legal risk in a nonparametric datastore
Sewon Min, Suchin Gururangan, Eric Wallace, Weijia Shi, Hannaneh Hajishirzi, Noah A Smith, and Luke Zettlemoyer · 2023
Later among the works it cites.
Scaling data-constrained language models
Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
AI Bot Blocking
Originality.ai · 2023
Later among the works it cites.
Gaia search: Hugging face and pyserini interoperability for nlp training data exploration
Aleksandra Piktus, Odunayo Ogundepo, Christopher Akiki, Akintunde Oladipo, Xinyu Zhang, Hailey Schoelkopf, Stella Biderman, Martin Potthast, and Jimmy Lin · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Cited alongside, same era.
Detoxifying language models risks marginalizing minority voices
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein · 2021
Cited alongside, same era.
mt5: A massively multilingual pre-trained text-to-text transformer, 2021
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2021
Cited alongside, same era.
Constitutional AI: Harmlessness from AI Feedback, December 2022
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan · 2022
Cited alongside, same era.
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, May 2022
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2022
Cited alongside, same era.
Stella Biderman, Kieran Bicheno, and Leo Gao · 2022
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al · 2022
Cited alongside, same era.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Cited alongside, same era.
On the challenges of using black-box apis for toxicity evaluation in research, 2023
Luiza Pozzobon, Beyza Ermis, Patrick Lewis, and Sara Hooker · 2023
Later among the works it cites.
Paul tremblay, mona awad vs. openai, inc., et al., 2023
Joseph R. Saveri, Cadio Zirpoli, Christopher K.L. Young, and Kathleen J. McMahon · 2023
Later among the works it cites.
Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Pete Walsh, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Datafinder: Scientific dataset recommendation from natural language descriptions
Vijay Viswanathan, Luyu Gao, Tongshuang Wu, Pengfei Liu, and Graham Neubig · 2023
Later among the works it cites.
WGA negotiations—status as of may 1, 2023, May 2023
Writers Guild of America · 2023
Later among the works it cites.
Terms & conditions, 2024
Aleph Alpha · 2024
Closest in time.
Usage policy, 2024
Anthropic · 2024
Closest in time.
Mint-1t: Scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens, 2024
Anas Awadalla, Le Xue, Oscar Lo, Manli Shu, Hannah Lee, Etash Kumar Guha, Matt Jordan, Sheng Shen, Mohamed Awadalla, Silvio Savarese, Caiming Xiong, Ran Xu, Yejin Choi, and Ludwig Schmidt · 2024
Closest in time.
A critical analysis of the largest source for generative ai training data: Common crawl
Stefan Baack · 2024
Closest in time.
When Online Content Disappears
Athena Chapekis, Samuel Bestvater, Emma Remy, and Gonzalo Rivero · 2024
Closest in time.
A survey of web content control for generative ai, 2024
Michael Dinzinger, Florian Heß, and Michael Granitzer · 2024
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al · 2024
Closest in time.
Generative ai prohibited use policy, 2024
Google · 2024
Closest in time.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al · 2024
Closest in time.
Securing the ai software supply chain
Isaac Hepworth, Kara Olive, Kingshuk Dasgupta, Michael Le, Mark Lodato, Mihai Maruseac, Sarah Meiklejohn, Shamik Chaudhuri, and Tehila Minkus · 2024
Closest in time.
Testimony before the senate judiciary subcommittee on privacy, technology, and the law: Oversight of a.i.: The future of journalism
Jeff Jarvis · 2024
Closest in time.
Acceptable use policies for foundation models: Considerations for policymakers and developers
Kevin Klyman · 2024
Closest in time.
Data authenticity, consent, & provenance for ai are all broken: what will it take to fix them?
Shayne Longpre, Robert Mahari, Naana Obeng-Marnu, William Brannon, Tobin South, Katy Gero, Sandy Pentland, and Jad Kabbara · 2024
Closest in time.
Perplexity is a bullshit machine
Dhruv Mehrotra and Tim Marchman · 2024
Closest in time.
“i searched for a religious song in amharic and got sexual content instead”: Investigating online harm in low-resourced languages on youtube
Hellina Hailu Nigatu and Inioluwa Deborah Raji · 2024
Closest in time.
Hello gpt-4o: We’re announcing gpt-4o, our new flagship model that can reason across audio, vision, and text in real time., 2024
OpenAI · 2024
Closest in time.
Exclusive: Multiple ai companies bypassing web standard to scrape publisher sites, licensing firm says
Katie Paul · 2024
Closest in time.
Nyt v. openai: The times’s about-face, April 2024
Audrey Pope · 2024
Closest in time.
Nightshade: Prompt-specific poisoning attacks on text-to-image generative models, 2024
Shawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu, Haitao Zheng, and Ben Y. Zhao · 2024
Closest in time.
URL https://haveibeentrained.com/
SpawningAI, 2024 · 2024
Closest in time.
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset, February 2024
Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, and Igor Gitman · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al · 2024
Closest in time.
Who blocks openAI, google AI and Common Crawl?
Ben Welsh · 2024
Closest in time.
Wildchat: 1m chatgpt interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng · 2024
Closest in time.
Multimodal c4: An open, billion-scale corpus of images interleaved with text
Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, and Yejin Choi · 2024
Closest in time.