Fetching the paper…
Reading the bibliography…
The race to train language models on vast, diverse, and inconsistently documented datasets has raised pressing concerns about the legal and ethical risks for practitioners.
Masking copyright decisionmaking: The meaninglessness of substantial similarity
Amy B Cohen · 1986
Earlier work this paper cites.
The berne convention: Its history and its key role in the future
Peter Burger · 1988
Earlier work this paper cites.
No sweat copyright and other protection of works of information after feist v. rural telephone
Jane C Ginsburg · 1992
Earlier work this paper cites.
Toward a better understanding of substantial similarity in copyright infringement cases
Jarrod M Mohler · 1999
Earlier work this paper cites.
Copyright in 1791: An essay concerning the founers’ view of the copyright power granted to congress in article i, section 8, clause 8 of the us constitution
L Patterson · 2003
Earlier work this paper cites.
Special issue on open source software development, 2003
Georg Von Krogh and Eric Von Hippel · 2003
Earlier work this paper cites.
From sony to grokster, the failure of the copyright doctrines of contributory infringement and vicarious liability to resolve the war between content and destructive technologies
Craig A Grossman · 2005
Earlier work this paper cites.
On achieving and evaluating language-independence in nlp
Emily M Bender · 2011
Earlier work this paper cites.
Judging similarity
Shyamkrishna Balganesh, Irina D Manta, and Tess Wilkinson-Ryan · 2014
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks, 2015
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M. Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith · 2017
Earlier work this paper cites.
Shreya Shankar, Yoni Halpern, Eric Breck, James Atwood, Jimbo Wilson, and D Sculley · 2017
Earlier work this paper cites.
Artificial intelligence’s fair use crisis
Benjamin LW Sobel · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman · 2018
Earlier work this paper cites.
Does object recognition work for everyone?
Terrance De Vries, Ishan Misra, Changhan Wang, and Laurens Van der Maaten · 2019
Earlier work this paper cites.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Fair learning
Mark A Lemley and Bryan Casey · 2020
Earlier work this paper cites.
The new legal landscape for text mining and machine learning
Matthew J. Sag · 2020
Earlier work this paper cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano · 2020
Earlier work this paper cites.
The low-resource double bind: An empirical study of pruning for low-resource machine translation
Orevaoghene Ahia, Julia Kreutzer, and Sara Hooker · 2021
Earlier work this paper cites.
Jack Bandy and Nicholas Vincent · 2021
Earlier work this paper cites.
The problem of zombie datasets: A framework for deprecating datasets
Frances Corry, Hamsini Sridharan, Alexandra Sasha Luccioni, Mike Ananny, Jason Schultz, and Kate Crawford · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner · 2021
Earlier work this paper cites.
Memorization vs. generalization : Quantifying data leakage in NLP performance evaluation
Aparna Elangovan, Jiayuan He, and Karin Verspoor · 2021
Earlier work this paper cites.
Datahunter: A system for finding datasets based on scientific problem descriptions
Michael Färber and Ann-Kathrin Leisinger · 2021
Earlier work this paper cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford · 2021
Earlier work this paper cites.
Ai derivatives: The application to the derivative work right to literary and artistic productions of ai machines
Daniel J Gervais · 2021
Earlier work this paper cites.
Towards accountability for machine learning datasets: Practices from software engineering and infrastructure
Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell · 2021
Earlier work this paper cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, et al · 2021
Earlier work this paper cites.
Understanding gender and racial disparities in image recognition models
Rohan Mahadev and Anindya Chakravarti · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna · 2021
Earlier work this paper cites.
Does training ai violate copyright law?
Jenny Quang · 2021
Earlier work this paper cites.
Changing the world by changing the data
Anna Rogers · 2021
Earlier work this paper cites.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo · 2021
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al · 2021
Earlier work this paper cites.
Jerrold Soh · 2021
Earlier work this paper cites.
Gaia search tool
Spacerini · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Cited alongside, same era.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang · 2021
Cited alongside, same era.
Detoxifying language models risks marginalizing minority voices
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
Art and the science of generative ai
Ziv Epstein, Aaron Hertzmann, Laura Herman, Robert Mahari, Morgan R Frank, Matthew Groh, Hope Schroeder, Amy Smith, Memo Akten, Jessica Fjeld, et al · 2023
Closest in time.
Stanford human preferences dataset, 2023
Kawin Ethayarajh, Heidi Zhang, Yizhong Wang, and Dan Jurafsky · 2023
Closest in time.
Tweet by mosaic ml
Jonathan Frankle · 2023
Closest in time.
Koala: A dialogue model for academic research
Xinyang Geng, Arnav Gudibande, Hao Liu, Eric Wallace, Pieter Abbeel, Sergey Levine, and Dawn Song · 2023
Closest in time.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection, 2023
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu · 2023
Closest in time.
Foundation models and fair use
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A Lemley, and Percy Liang · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stella Biderman, Kieran Bicheno, and Leo Gao · 2022
Cited alongside, same era.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Cited alongside, same era.
Behavioral use licensing for responsible ai
Danish Contractor, Daniel McDuff, Julia Katherine Haines, Jenny Lee, Christopher Hines, Brent Hecht, Nicholas Vincent, and Hanlin Li · 2022
Cited alongside, same era.
Interactive model cards: A human-centered approach to model documentation
Anamaria Crisan, Margaret Drouhard, Jesse Vig, and Nazneen Rajani · 2022
Cited alongside, same era.
Sui generis database protection 2.0: judicial and legislative reforms
Estelle Derclaye and Martin Husovec · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Cited alongside, same era.
Closest in time.
Are chatgpt and other similar systems the modern lernaean hydras of ai?
Dimitrios Ioannidis, Jeremy Kepner, Andrew Bowne, and Harriet S Bryant · 2023
Closest in time.
Reforms: Reporting standards for machine learning based science
Sayash Kapoor, Emily F. Cantrell, Kenny Peng, Thanh Hien Pham, Christopher A. Bail, Odd Erik Gundersen, Jake M. Hofman, Jessica R. Hullman, Michael A. Lones, Momin M. Malik, Priyanka Nanayakkara, Russel A. Poldrack, Inioluwa Deborah Raji, Michael Roberts, Matthew J. Salganik, Marta Serra-Garcia, Brandon M Stewart, Gilles Vandewiele, and Arvind Narayanan · 2023
Closest in time.
The cot collection: Improving zero-shot and few-shot learning of language models via chain-of-thought fine-tuning, 2023
Seungone Kim, Se June Joo, Doyoung Kim, Joel Jang, Seonghyeon Ye, Jamin Shin, and Minjoon Seo · 2023
Closest in time.
Rodney Michael Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexandra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, Miles Crawford, Doug Downey, Jason Dunkelberger, Oren Etzioni, Rob Evans, Sergey Feldman, Joseph Gorney, David W. Graham, F.Q. Hu, Regan Huff, Daniel King, Sebastian Kohlmeier, Bailey Kuehl, Michael Langan, Daniel Lin, Haokun Liu, Kyle Lo, Jaron Lochner, Kelsey MacMillan, Tyler Murray, Christopher Newell, Smita Rao, Shaurya Rohatgi, Paul L Sayre, Zejiang Shen, Amanpreet Singh, Luca Soldaini, Shivashankar Subramanian, A. Tanaka, Alex D Wade, Linda M. Wagner, Lucy Lu Wang, Christopher Wilhelm, Caroline Wu, Jiangjiang Yang, Angele Zamarron, Madeleine van Zuylen, and Daniel S. Weld · 2023
Closest in time.
Openassistant conversations–democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, et al · 2023
Closest in time.
Longform: Optimizing instruction tuning for long text generation with corpus extraction, 2023
Abdullatif Köksal, Timo Schick, Anna Korhonen, and Hinrich Schütze · 2023
Closest in time.
Talkin”bout ai generation: Copyright and the generative-ai supply chain
Katherine Lee, A Feder Cooper, and James Grimmelmann · 2023
Closest in time.
Wizardcoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang · 2023
Closest in time.
The Law of Business Torts and Unfair Competition: Cases, Materials, and Problems
Colin P. Marks and Douglas K. Moll · 2023
Closest in time.
The dataset multiplicity problem: How unreliable data impacts predictions
Anna P. Meyer, Aws Albarghouthi, and Loris D’Antoni · 2023
Closest in time.
Silo language models: Isolating legal risk in a nonparametric datastore
Sewon Min, Suchin Gururangan, Eric Wallace, Hanna Hajishirzi, Noah A. Smith, and Luke Zettlemoyer · 2023
Closest in time.
Orca: Progressive learning from complex explanation traces of gpt-4, 2023
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah · 2023
Closest in time.
The open instruction generalist (oig) dataset
Huu Nguyen, Sameer Suri, Ken Tsui, and Christoph Schuhmann · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Gorilla: Large language model connected with massive apis, 2023
Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Closest in time.
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao · 2023
Closest in time.
Language model tokenizers introduce unfairness between languages
Aleksandar Petrov, Emanuele La Malfa, Philip HS Torr, and Adel Bibi · 2023
Closest in time.
On the challenges of using black-box apis for toxicity evaluation in research, 2023
Luiza Pozzobon, Beyza Ermis, Patrick Lewis, and Sara Hooker · 2023
Closest in time.
Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2023
Closest in time.
Copyright safety for generative ai
Matthew Sag · 2023
Closest in time.
Generative ai meets copyright
Pamela Samuelson · 2023
Closest in time.
Paul tremblay, mona awad vs. openai, inc., et al., 2023
Joseph R. Saveri, Cadio Zirpoli, Christopher K.L. Young, and Kathleen J. McMahon · 2023
Closest in time.
tasksource: A dataset harmonization framework for streamlined nlp multi-task learning and evaluation
Damien Sileo · 2023
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2023
Closest in time.
Protecting customers with generative AI indemnification, 2023
Neal Suggs and Phil Venables · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Sharegpt, 2023
Vercel · 2023
Closest in time.
Alphabet’s google and deepmind pause grudges, join forces to chase openai
Jon Victor and Amir Efrati · 2023
Closest in time.
Datafinder: Scientific dataset recommendation from natural language descriptions
Vijay Viswanathan, Luyu Gao, Tongshuang Wu, Pengfei Liu, and Graham Neubig · 2023
Closest in time.
Provable copyright protection for generative models
Nikhil Vyas, Sham Kakade, and Boaz Barak · 2023
Closest in time.
Lima: Less is more for alignment, 2023
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Closest in time.