Fetching the paper…
Reading the bibliography…
New capabilities in foundation models are owed in large part to massive, widely-sourced, and under-documented training data collections.
Measuring bias in contextualized word representations
Kurita, K., Vyas, N., Pareek, A., Black, A. W., and Tsvetkov, Y · 1906
Earlier work this paper cites.
Hidden digital watermarks in images
Hsu, C.-T. and Wu, J.-L · 1999
Earlier work this paper cites.
We’ve reached peak infographics. are you ready for what comes next?
Lupi, G · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Bender, E. M. and Friedman, B · 2018
Earlier work this paper cites.
The productivity j-curve: How intangibles complement general purpose technologies
Brynjolfsson, E., Rock, D., and Syverson, C · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Buolamwini, J. and Gebru, T · 2018
Earlier work this paper cites.
Humans forget, machines remember: Artificial intelligence and the right to be forgotten
Villaronga, E. F., Kieseberg, P., and Li, T · 2018
Earlier work this paper cites.
Mitigating unwanted biases with adversarial learning
Zhang, B. H., Lemoine, B., and Mitchell, M · 2018
Earlier work this paper cites.
Creativity, copyright, and close-knit communities: a case study of social norm formation and enforcement
Fiesler, C. and Bruckman, A. S · 2019
Earlier work this paper cites.
Model cards for model reporting
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., and Gebru, T · 2019
Earlier work this paper cites.
The emergence of deepfake technology: A review
Westerlund, M · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Forecasting extreme labor displacement: A survey of AI practitioners
Gruetzemacher, R., Paradice, D., and Lee, K. B · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Fair learning
Lemley, M. A. and Casey, B · 2020
Earlier work this paper cites.
Saving face: Investigating the ethical concerns of facial recognition auditing
Raji, I. D., Gebru, T., Mitchell, M., Buolamwini, J., Lee, J., and Denton, E · 2020
Earlier work this paper cites.
Bandy, J. and Vincent, N · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Datasheets for datasets help ML engineers notice and understand ethical issues in training data
Boyd, K. L · 2021
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Earlier work this paper cites.
The Data Nutrition Project
Data Nutrition Team · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., and Gardner, M · 2021
Earlier work this paper cites.
Datasheets for datasets
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Iii, H. D., and Crawford, K · 2021
Earlier work this paper cites.
A new AI lexicon: Open
Khan, M · 2021
Earlier work this paper cites.
Datasets: A community library for natural language processing
Lhoest, Q., del Moral, A. V., Jernite, Y., Thakur, A., von Platen, P., Patil, S., Chaumond, J., Drame, M., Plu, J., Tunstall, L., et al · 2021
Earlier work this paper cites.
Changing the world by changing the data
Rogers, A · 2021
Earlier work this paper cites.
Just what do you think you’re doing, Dave?’a checklist for responsible data use in NLP
Rogers, A., Baldwin, T., and Leins, K · 2021
Earlier work this paper cites.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes AI
Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., and Aroyo, L. M · 2021
Earlier work this paper cites.
Image representations learned with unsupervised pre-training contain human-like biases
Steed, R. and Caliskan, A · 2021
Earlier work this paper cites.
New frontiers: The origins and content of new work, 1940–2018
Autor, D., Chin, C., Salomons, A., and Seegmiller, B · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Earlier work this paper cites.
Pile of law: Learning responsible data filtering from the law and a 256gb open-source legal dataset
Henderson, P., Krass, M., Zheng, L., Guha, N., Manning, C. D., Jurafsky, D., and Ho, D · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Earlier work this paper cites.
Data governance in the age of large-scale data-driven language technology
Jernite, Y., Nguyen, H., Biderman, S., Rogers, A., Masoud, M., Danchev, V., Tan, S., Luccioni, A. S., Subramani, N., Johnson, I., et al · 2022
Earlier work this paper cites.
Provenance documentation to enable explainable and trustworthy AI: A literature review
Kale, A., Nguyen, T., Jr., F. C. H., Li, C., Zhang, J., and Ma, X · 2022
Cited alongside, same era.
Leakage and the reproducibility crisis in ML-based science
Kapoor, S. and Narayanan, A · 2022
Cited alongside, same era.
“this isn’t your data, friend”: Black Twitter as a case study on research ethics for public data
Klassen, S. and Fiesler, C · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al · 2022
Cited alongside, same era.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al · 2022
Cited alongside, same era.
How AI is being abused to create child sexual abuse imagery
Internet Watch Foundation · 2023
Later among the works it cites.
Donottrain: A metadata standard for indicating consent for machine learning
Ippolito, D. and Yu, Y. W · 2023
Later among the works it cites.
Defining best practices for opting out of ml training
Keller, P. and Warso, Z · 2023
Later among the works it cites.
A watermark for large language models
Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T · 2023
Later among the works it cites.
Talkin”bout ai generation: Copyright and the generative-AI supply chain
Lee, K., Cooper, A. F., and Grimmelmann, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Time waits for no one! analysis and challenges of temporal misalignment
Luu, K., Khashabi, D., Gururangan, S., Mandyam, K., and Smith, N. A · 2022
Cited alongside, same era.
Hello datasphere — towards a systems approach to data governance
Porciuncula, L. and De La Chapelle, B · 2022
Cited alongside, same era.
Data cards: Purposeful and transparent dataset documentation for responsible AI
Pushkarna, M., Zaldivar, A., and Kjartansson, O · 2022
Cited alongside, same era.
C2PA: the world’s first industry standard for content provenance (conference presentation)
Rosenthol, L · 2022
Cited alongside, same era.
The lawsuit that could rewrite the rules of AI copyright, 2022
Vincent, J · 2022
Cited alongside, same era.
Establishing data provenance for responsible artificial intelligence systems
Werder, K., Ramesh, B., and Zhang, R. S · 2022
Cited alongside, same era.
The New York Times got its content removed from one of the biggest AI training datasets. here’s how it did it
Barr, A. and Hays, K · 2023
Cited alongside, same era.
Lohr, S · 2023
Later among the works it cites.
Comments regarding artificial intelligence and copyright
Mahari, R., Longpre, S., Donewald, L., Polozov, A., Lipsitz, A., and Pentland, S · 2023
Later among the works it cites.
404 Media generative AI market analysis: People love to cum
Maiberg, E · 2023
Later among the works it cites.
OpenAI’s ChatGPT now has 100 million weekly active users
Malik, A · 2023
Later among the works it cites.
AI adoption in America: Who, what, and where
McElheran, K., Li, J. F., Brynjolfsson, E., Kroff, Z., Dinlersoz, E., Foster, L., and Zolas, N · 2023
Later among the works it cites.
Silo language models: Isolating legal risk in a nonparametric datastore
Min, S., Gururangan, S., Wallace, E., Hajishirzi, H., Smith, N. A., and Zettlemoyer, L · 2023
Later among the works it cites.
Generative AI companies must publish transparency reports
Narayanan, A. and Kapoor, S · 2023
Later among the works it cites.
Scalable extraction of training data from (production) language models
Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., and Lee, K · 2023
Later among the works it cites.
The Impact of AI on Developer Productivity: Evidence from GitHub Copilot, 2023
Peng, S., Kalliamvakou, E., Cihon, P., and Demirer, M · 2023
Later among the works it cites.
Stronger together: on the articulation of ethical charters, legal tools, and technical documentation in ML
Pistilli, G., Muñoz Ferrandis, C., Jernite, Y., and Mitchell, M · 2023
Later among the works it cites.
A survey of hallucination in large foundation models
Rawte, V., Sheth, A., and Das, A · 2023
Later among the works it cites.
Can AI-generated text be reliably detected?
Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S · 2023
Later among the works it cites.
Paul Tremblay, Mona Awad vs. OpenAI, Inc., et al., 2023
Saveri, J. R., Zirpoli, C., Young, C. K., and McMahon, K. J · 2023
Later among the works it cites.
The curse of recursion: Training on generated data makes models forget, 2023
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R · 2023
Later among the works it cites.
Towards expert-level medical question answering with large language models
Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Hou, L., Clark, K., Pfohl, S., Cole-Lewis, H., Neal, D., et al · 2023
Later among the works it cites.
Sarah Silverman sues OpenAI and Meta over copyright infringement, 2023
Small, Z · 2023
Later among the works it cites.
AI is no threat to traditional artists. but it is thrilling
Smee, S · 2023
Later among the works it cites.
A global digital compact — an open, free and secure digital future for all
United Nations · 2023
Later among the works it cites.
Artificial artificial artificial intelligence: Crowd workers widely use large language models for text production tasks, 2023
Veselovsky, V., Ribeiro, M. H., and West, R · 2023
Later among the works it cites.
Getty Images sues AI art generator Stable Diffusion in the US for copyright infringement — theverge.com
Vincent, J · 2023
Later among the works it cites.
Getty Images is suing the creators of AI art tool Stable Diffusion for scraping its content
Vincent, J · 2023
Later among the works it cites.
Market concentration implications of foundation models: The invisible hand of ChatGPT
Vipra, J. and Korinek, A · 2023
Later among the works it cites.
WGA negotiations—status as of may 1, 2023, May 2023
Writers Guild of America · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey, 2023
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., Yin, Z., Dou, S., Weng, R., Cheng, W., Zhang, Q., Qin, W., Zheng, Y., Qiu, X., Huang, X., and Gui, T · 2023
Later among the works it cites.
Machine unlearning: A survey, 2023
Xu, H., Zhu, T., Zhang, L., Zhou, W., and Yu, P. S · 2023
Later among the works it cites.
Watermarks in the sand: Impossibility of strong watermarking for generative models
Zhang, H., Edelman, B. L., Francati, D., Venturi, D., Ateniese, G., and Barak, B · 2023
Later among the works it cites.
Can large language models transform computational social science?
Ziems, C., Shaikh, O., Zhang, Z., Held, W., Chen, J., and Yang, D · 2023
Later among the works it cites.
Foundation model transparency reports
Bommasani, R., Klyman, K., Longpre, S., Xiong, B., Kapoor, S., Maslej, N., Narayanan, A., and Liang, P · 2024
Closest in time.
On the societal impact of open foundation models, 2024
Kapoor, S., Bommasani, R., Klyman, K., Longpre, S., Ramaswami, A., Cihon, P., Hopkins, A., Bankston, K., Biderman, S., Bogen, M., Chowdhury, R., Engler, A., Henderson, P., Jernite, Y., Lazar, S., Maffulli, S., Nelson, A., Pineau, J., Skowron, A., Song, D., Storchan, V., Zhang, D., Ho, D. E., Liang, P., and Narayanan, A · 2024
Closest in time.
Exploring the landscape of machine unlearning: A comprehensive survey and taxonomy, 2024
Shaik, T., Tao, X., Xie, H., Li, L., Zhu, X., and Li, Q · 2024
Closest in time.