Fetching the paper…
Reading the bibliography…
Many generative foundation models (or GFMs) are trained on publicly available data and use public infrastructure, but 1) may degrade the "digital commons" that they depend on, and 2) do not have processes in place to return value captured to data producers and stakeholders.
Governing the commons: The evolution of institutions for collective action
Elinor Ostrom · 1990
Earlier work this paper cites.
The MIT Press, 2007
Understanding Knowledge as a Commons: From Theory to Practice · 2007
Earlier work this paper cites.
Economic impact of open source software on innovation and the competitiveness of the information and communication technologies (ict) sector in the eu
Rishab Aiyer Ghosh · 2007
Earlier work this paper cites.
The master switch: The rise and fall of information empires
Tim Wu · 2012
Earlier work this paper cites.
The spreading of misinformation online
Michela Del Vicario, Alessandro Bessi, Fabiana Zollo, Fabio Petroni, Antonio Scala, Guido Caldarelli, H Eugene Stanley, and Walter Quattrociocchi · 2016
Earlier work this paper cites.
The substantial interdependence of wikipedia and google: A case study on the relationship between peer production communities and information technologies
Connor McMahon, Isaac Johnson, and Brent Hecht · 2017
Earlier work this paper cites.
Amplifying the impact of open access: Wikipedia and the diffusion of science
Misha Teplitskiy, Grace Lu, and Eamon Duede · 2017
Earlier work this paper cites.
Caveat emptor, computational social science: Large-scale missing data in a widely-published reddit corpus
Devin Gaffney and J Nathan Matias · 2018
Earlier work this paper cites.
Generating sentiment-preserving fake online reviews using neural language models and their human- and machine-based detection, 2019
David Ifeoluwa Adelani, Haotian Mai, Fuming Fang, Huy H. Nguyen, Junichi Yamagishi, and Isao Echizen · 2019
Earlier work this paper cites.
Use and fair use: Statement on shared images in facial recognition ai
Ryan Merkley · 2019
Earlier work this paper cites.
Deepfake bot submissions to federal public comment websites cannot be distinguished from human submissions
Max Weiss · 2019
Earlier work this paper cites.
Concerns about democracy in the digital age
Janna Anderson and Lee Rainie · 2020
Earlier work this paper cites.
Digital commons
Melanie Dulong de Rosnay and Felix Stalder · 2020
Earlier work this paper cites.
What bert is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Closing the ai accountability gap: Defining an end-to-end framework for internal algorithmic auditing, 2020
Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes · 2020
Earlier work this paper cites.
Do neural language models overcome reporting bias?
Vered Shwartz and Yejin Choi · 2020
Earlier work this paper cites.
Why getting paid for your data is a bad deal
Hayley Tsukayama · 2020
Earlier work this paper cites.
Defending against neural fake news, 2020
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi · 2020
Earlier work this paper cites.
DIALOGPT : Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan · 2020
Earlier work this paper cites.
Problematic machine behavior: A systematic literature review of algorithm audits, 2021
Jack Bandy · 2021
Earlier work this paper cites.
Fueling the fire: How social media intensifies us political polarization–and what can be done about it
PM Barrett, J Hendrix, and JG Sims · 2021
Cited alongside, same era.
The impact of open source software and hardware on technological independence, competitiveness and innovation in the eu economy
Knut Blind, Mirko Böhm, Paula Grzegorzewska, Andrew Katz, Sachiko Muto, Sivan Pätsch, and Torben Schubert · 2021
Cited alongside, same era.
Truth, lies, and automation: How language models could change disinformation
Ben Buchanan, Andrew Lohn, and Micah Musser · 2021
Cited alongside, same era.
A data dividend that works: steps toward building an equitable data economy
Yakov Feygin, Hanlin Li, Chirag Lala, Brent Hecht, Nicholas Vincent, Luisa Scarcella, and Matthew Prewitt · 2021
Cited alongside, same era.
A theoretical analysis of the repetition problem in text generation
Zihao Fu, Wai Lam, Anthony Man-Cho So, and Bei Shi · 2021
Cited alongside, same era.
Auto-generated content
Ahrefs · 2023
Closest in time.
Invasive diffusion: How one unwilling illustrator found herself turned into an ai model, November 2022
Andy Baio · 2023
Closest in time.
’ai should exclude living artists from its database,’ says one painter whose works were used to fuel image generators, 2022
Vittoria Benzine · 2023
Closest in time.
Deep reinforcement learning from human preferences, 2023
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2023
Closest in time.
Democracy and the epistemic commons, February 2021
The Consilience Project · 2023
Closest in time.
A tracker of generative ai-related lawsuits, January 2023
Hayden Field · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fair learning
Mark A Lemley and Bryan Casey · 2021
Cited alongside, same era.
The internet tax freedom act and federal preemption
Congressional Research Service · 2021
Cited alongside, same era.
Open source software and global entrepreneurship
Nataliya Wright, Frank Nagle, and Shane M Greenstein · 2021
Cited alongside, same era.
Easily accessible text-to-image generation amplifies demographic stereotypes at large scale, 2022
Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan · 2022
Cited alongside, same era.
On the opportunities and risks of foundation models, 2022
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Ben Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, Julian Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Rob Reich, Hongyu Ren, Frieda Rong, Yusuf Roohani, Camilo Ruiz, Jack Ryan, Christopher Ré, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishnan Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tramèr, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang · 2022
Cited alongside, same era.
Why meta’s latest large language model survived only three days online, 2022
Will Douglas Heaven · 2022
Cited alongside, same era.
Capturing failures of large language models via human cognitive biases
Erik Jones and Jacob Steinhardt · 2022
Cited alongside, same era.
The most visited website in every country, 2022
Domantas G · 2023
Closest in time.
Generative language models and automated influence operations: Emerging threats and potential mitigations, 2023
Josh A. Goldstein, Girish Sastry, Micah Musser, Renee DiResta, Matthew Gentzel, and Katerina Sedova · 2023
Closest in time.
Spammy automatically generated content
Google · 2023
Closest in time.
Researchers find stable diffusion amplifies stereotypes
Justin Hendrix · 2023
Closest in time.
Radiocarbon dating and bomb carbon
Beta Analytic Testing Laboratory · 2023
Closest in time.
The internet’s favorite website, 2016
Adrienne LaFrance · 2023
Closest in time.
Why are there so many wikipedia articles in swedish and cebuano?, 2021
Ivan Lokhov · 2023
Closest in time.
Social media is polluting society. moderation alone won’t fix the problem, 2022
Nathaniel Lubin and Thomas Krendl Gilbert · 2023
Closest in time.
Is chat gpt biased against conservatives? an empirical study
Robert W McGee · 2023
Closest in time.
How a scots wikipedia scandal highlighted ai’s data problem, 2020
Nicolas Rivero · 2023
Closest in time.
The case for the digital commons, Jun 2021
Divya Siddarth and Glen Weyl · 2023
Closest in time.
Will we run out of ml data? evidence from projecting dataset size trends
Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, and Anson Ho · 2023
Closest in time.
Why is tabnine better than github copilot?, 2022
Anirudh VK · 2023
Closest in time.
tudents are using ai to write their papers, because of course they are, 2022
Claire Woodcock · 2023
Closest in time.