Fetching the paper…
Reading the bibliography…
'Scale the model, scale the data, scale the compute' is the reigning sentiment in the world of generative AI today.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Viral hate: Containing its spread on the Internet
Foxman, A. H., and Wolf, C · 2013
Earlier work this paper cites.
Hate crimes in cyberspace
Citron, D. K · 2014
Earlier work this paper cites.
Gendertrolling: How misogyny went viral: How misogyny went viral
Mantilla, K · 2015
Earlier work this paper cites.
Stereotyping and bias in the flickr30k dataset
Van Miltenburg, E · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Davidson, T., Warmsley, D., Macy, M., and Weber, I · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M., Ali, M., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Understanding abuse: A typology of abusive language detection subtasks
Waseem, Z., Davidson, T., Warmsley, D., and Weber, I · 2017
Earlier work this paper cites.
Racial bias in hate speech and abusive language detection datasets
Davidson, T., Bhattacharya, D., and Weber, I · 2019
Earlier work this paper cites.
Hate speech detection: Challenges and solutions
MacAvaney, S., Yao, H.-R., Yang, E., Russell, K., Goharian, N., and Frieder, O · 2019
Earlier work this paper cites.
Actionable auditing: Investigating the impact of publicly naming biased performance results of commercial ai products
Raji, I. D., and Buolamwini, J · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Against scale: Provocations and resistances to scale thinking
Hanna, A., and Park, T. M · 2020
Earlier work this paper cites.
Detoxify
Hanu, L., and Unitary team · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
How we’ve taught algorithms to see identity: Constructing race and gender in image databases for facial analysis
Scheuerman, M. K., Wade, K., Lustig, C., and Brubaker, J. R · 2020
Earlier work this paper cites.
Velioglu, R., and Rose, J · 2020
Earlier work this paper cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Vidgen, B., Thrush, T., Waseem, Z., and Kiela, D · 2020
Earlier work this paper cites.
The grey hoodie project: Big tobacco, big tech, and the threat on academic integrity
Abdalla, M., and Abdalla, M · 2021
Earlier work this paper cites.
Persistent anti-muslim bias in large language models
Abid, A., Farooqi, M., and Zou, J · 2021
Earlier work this paper cites.
To" see" is to stereotype: Image tagging algorithms, gender recognition, and the accuracy-fairness trade-off
Barlas, P., Kyriakou, K., Guest, O., Kleanthous, S., and Otterbacher, J · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
The values encoded in machine learning research
Birhane, A., Kalluri, P., Card, D., Agnew, W., Dotan, R., and Bao, M · 2021
Earlier work this paper cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Birhane, A., Prabhu, V. U., and Kahembwe, E · 2021
Earlier work this paper cites.
On the genealogy of machine learning datasets: A critical history of imagenet
Denton, E., Hanna, A., Amironesei, R., Smart, A., and Nicole, H · 2021
Cited alongside, same era.
Documenting the english colossal clean crawled corpus
Dodge, J., Sap, M., Marasovic, A., Agnew, W., Ilharco, G., Groeneveld, D., and Gardner, M · 2021
Cited alongside, same era.
Datasheets for datasets
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Iii, H. D., and Crawford, K · 2021
Cited alongside, same era.
A systematic review of hate speech automatic detection using natural language processing
Jahan, M. S., and Oussalah, M · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T · 2021
Ai recognition of patient race in medical imaging: a modelling study
Gichoya, J. W., Banerjee, I., Bhimireddy, A. R., Burns, J. L., Celi, L. A., Chen, L.-C., Correa, R., Dullerud, N., Ghassemi, M., Huang, S.-C., et al · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
The bigscience roots corpus: A 1.6 tb composite multilingual dataset
Laurençon, H., Saulnier, L., Wang, T., Akiki, C., Villanova del Moral, A., Le Scao, T., Von Werra, L., Mou, C., González Ponferrada, E., Nguyen, H., et al · 2022
Later among the works it cites.
Shutterstock to integrate OpenAI’s DALL-E 2 and launch fund for contributor artists, 2022
Lomas, N · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reduced, reused and recycled: The life of a dataset in machine learning research
Koch, B., Denton, E., Hanna, A., and Foster, J. G · 2021
Cited alongside, same era.
Multimodal cyberbullying detection using capsule network with dynamic routing and deep convolutional neural network
Kumar, A., and Sachdeva, N · 2021
Cited alongside, same era.
Disentangling hate in online memes
Lee, R. K.-W., Cao, R., Fan, Z., Jiang, J., and Chong, W.-H · 2021
Cited alongside, same era.
What’s in the box? an analysis of undesirable content in the common crawl corpus
Luccioni, A. S., and Viviano, J. D · 2021
Cited alongside, same era.
Auditing algorithms: Understanding algorithmic systems from the outside in
Metaxa, D., Park, J. S., Robertson, R. E., Karahalios, K., Wilson, C., Hancock, J., Sandvig, C., et al · 2021
Cited alongside, same era.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Paullada, A., Raji, I. D., Bender, E. M., Denton, E., and Hanna, A · 2021
Cited alongside, same era.
Mitigating dataset harms requires stewardship: Lessons from 1000 papers
Peng, K., Mathur, A., and Narayanan, A · 2021
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Later among the works it cites.
Diffusion art or digital forgery? investigating data replication in diffusion models
Somepalli, G., Singla, V., Goldblum, M., Geiping, J., and Goldstein, T · 2022
Later among the works it cites.
Manifestations of xenophobia in ai systems, 2022
Tomasev, N., Maynard, J. L., and Gabriel, I · 2022
Later among the works it cites.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al · 2022
Later among the works it cites.
Extracting training data from diffusion models
Carlini, N., Hayes, J., Nasr, M., Jagielski, M., Sehwag, V., Tramer, F., Balle, B., Ippolito, D., and Wallace, E · 2023
Closest in time.
A literature survey on multimodal and multilingual automatic hate speech identification
Chhabra, A., and Vishwakarma, D. K · 2023
Closest in time.
Flooded with ai-generated images, some art communities ban them completely | ars technica
EDWARDS, B · 2023
Closest in time.
Are we on the cusp of a generative ai revolution?
Eric Sheridan, K. R · 2023
Closest in time.
Uncurated image-text datasets: Shedding light on demographic bias
Garcia, N., Hirota, Y., Wu, Y., and Nakashima, Y · 2023
Closest in time.
Generative ai is changing everything. but what’s left when the hype is gone?
Heaven, W. D · 2023
Closest in time.
The generative ai revolution has begun—how did we get here? | ars technica
Huang, H · 2023
Closest in time.
The generative ai revolution will enable anyone to create games | andreessen horowitz
Joshua Lu, R. G · 2023
Closest in time.
Stable diffusion copyright lawsuits could be a legal earthquake for ai | ars technica
LEE, T. B · 2023
Closest in time.
Stable bias: Analyzing societal representations in diffusion models
Luccioni, A. S., Akiki, C., Mitchell, M., and Jernite, Y · 2023
Closest in time.
Understanding hate speech
Nations, U · 2023
Closest in time.
Stable diffusion ai has mastered the female form - niche gamer
Orselli, B · 2023
Closest in time.
Evaluating the social impact of generative ai systems in systems and society
Solaiman, I., Talat, Z., Agnew, W., Ahmad, L., Baker, D., Blodgett, S. L., Daumé III, H., Dodge, J., Evans, E., Hooker, S., et al · 2023
Closest in time.
Scaling laws literature review, 2023
Villalobos, P · 2023
Closest in time.
On the de-duplication of laion-2b
Webster, R., Rabin, J., Simon, L., and Jurie, F · 2023
Closest in time.