Fetching the paper…
Reading the bibliography…
Whereas much of the success of the current generation of neural language models has been driven by increasingly large training corpora, relatively little research has been dedicated to analyzing these massive sources of textual data.
Gonen, H. and Goldberg, Y. (2019) · 1903
Earlier work this paper cites.
On measuring social biases in sentence encoders
May, C., Wang, A., Bordia, S., Bowman, S. R., and Rudinger, R. (2019) · 1903
Earlier work this paper cites.
Identifying and reducing gender bias in word-level language models
Bordia, S. and Bowman, S. R. (2019) · 1904
Earlier work this paper cites.
Manzini, T., Lim, Y. C., Tsvetkov, Y., and Black, A. W. (2019) · 1904
Earlier work this paper cites.
Gender bias in contextualized word embeddings
Zhao, J., Wang, T., Yatskar, M., Cotterell, R., Ordonez, V., and Chang, K.-W. (2019) · 1904
Earlier work this paper cites.
Tackling online abuse: A survey of automated abuse detection methods
Mishra, P., Yannakoudakis, H., and Shutova, E. (2019) · 1908
Earlier work this paper cites.
Release strategies and the social impacts of language models
Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., et al. (2019) · 1908
Earlier work this paper cites.
Temporal effects of unmoderated hate speech in gab
Mathew, B., Illendula, A., Saha, P., Sarkar, S., Goyal, P., and Mukherjee, A. (2019) · 1909
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N. (2019) · 1909
Earlier work this paper cites.
On the unintended social bias of training language generation models with data from local media
Florez, O. U. (2019) · 1911
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Webguard: Web based adult content detection and filtering system
Hammami, M., Chahir, Y., and Chen, L. (2003) · 2003
Earlier work this paper cites.
Deep learning models for multilingual hate speech detection
Aluru, S. S., Mathew, B., Saha, P., and Mukherjee, A. (2020) · 2004
Earlier work this paper cites.
Statistical and structural approaches to filtering internet pornography
Ho, W. H. and Watters, P. A. (2004) · 2004
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Nadeem, M., Bethke, A., and Reddy, S. (2020) · 2004
Earlier work this paper cites.
Directions in abusive language training data: Garbage in, garbage out
Vidgen, B. and Derczynski, L. (2020) · 2004
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 2005
Earlier work this paper cites.
Social biases in NLP models as barriers for persons with disabilities
Hutchinson, B., Prabhakaran, V., Denton, E., Webster, K., Zhong, Y., and Denuyl, S. (2020) · 2005
Earlier work this paper cites.
Pointer: Constrained text generation via insertion-based generative pre-training
Zhang, Y., Wang, G., Li, C., Gan, Z., Brockett, C., and Dolan, B. (2020) · 2005
Earlier work this paper cites.
Content-based text classifiers for pornographic web filtering
Polpinij, J., Chotthanom, A., Sibunruang, C., Chamchong, R., and Puangpronpitag, S. (2006) · 2006
Earlier work this paper cites.
Large scale image-based adult-content filtering
Rowley, H. A., Jing, Y., and Baluja, S. (2006) · 2006
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. (2020) · 2009
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al. (2020) · 2012
Earlier work this paper cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Paullada, A., Raji, I. D., Bender, E. M., Denton, E., and Hanna, A. (2020) · 2012
Earlier work this paper cites.
Exploratory analysis of a terabyte scale web corpus
Kolias, V., Anagnostopoulos, I., and Kayafas, E. (2014) · 2014
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D. (2015) · 2015
Cited alongside, same era.
Analysis of representation of sexuality on women’s and men’s pornographic websites
Shim, J. W., Kwon, M., and Cheng, H.-I. (2015) · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. (2015) · 2015
Cited alongside, same era.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T. (2016) · 2016
Know what you don’t know: Unanswerable questions for SQuAD
Rajpurkar, P., Jia, R., and Liang, P. (2018) · 2018
Later among the works it cites.
Adult content detection in videos with convolutional and recurrent neural networks
Wehrmann, J., Simões, G. S., Barros, R. C., and Cavalcante, V. F. (2018) · 2018
Later among the works it cites.
Learning gender-neutral word embeddings
Zhao, J., Zhou, Y., Li, Z., Wang, W., and Chang, K.-W. (2018) · 2018
Later among the works it cites.
Semeval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter
Basile, V., Bosco, C., Fersini, E., Debora, N., Patti, V., Pardo, F. M. R., Rosso, P., Sanguinetti, M., et al. (2019) · 2019
Later among the works it cites.
Exploring misogyny across the manosphere in reddit
Farrell, T., Fernandez, M., Novotny, J., and Alani, H. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Languagecrawl: A generic tool for building language models upon common-crawl
Roziewski, S. and Stokowiec, W. (2016) · 2016
Cited alongside, same era.
Hateful symbols or hateful people? predictive features for hate speech detection on Twitter
Waseem, Z. and Hovy, D. (2016) · 2016
Cited alongside, same era.
Deep learning for hate speech detection in tweets
Badjatiya, P., Gupta, S., Gupta, M., and Varma, V. (2017) · 2017
Cited alongside, same era.
Semantics derived automatically from language corpora contain human-like biases
Caliskan, A., Bryson, J. J., and Narayanan, A. (2017) · 2017
Cited alongside, same era.
Automated hate speech detection and the problem of offensive language
Davidson, T., Warmsley, D., Macy, M., and Weber, I. (2017) · 2017
Cited alongside, same era.
Latent constraints: Learning to generate conditionally from unconditional generative models
Engel, J., Hoffman, M., and Roberts, A. (2017) · 2017
Cited alongside, same era.
Bias in Wikipedia
Hube, C. (2017) · 2017
Cited alongside, same era.
Pornography and sexual violence
Foubert, J. D., Blanchard, W., Houston, M., and Williams, R. R. (2019) · 2019
Later among the works it cites.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
Ortiz Suarez, P. J., Sagot, B., and Romary, L. (2019) · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Later among the works it cites.
A transparent framework for evaluating unintended demographic bias in word embeddings
Sweeney, C. and Najafian, M. (2019) · 2019
Later among the works it cites.
Detection of depression-related posts in reddit social media forum
Tadesse, M. M., Lin, H., Xu, B., and Yang, L. (2019) · 2019
Later among the works it cites.
Understanding the overlap between cyberbullying and cyberhate perpetration: Moderating effects of toxic online disinhibition
Wachs, S., Wright, M. F., and Vazsonyi, A. T. (2019) · 2019
Later among the works it cites.
Worse than objects: The depiction of black women and men and their sexual relationship in pornography
Fritz, N., Malic, V., Paul, B., and Zhou, Y. (2020) · 2020
Later among the works it cites.
Web page classification: a survey of perspectives, gaps, and future directions
Hashemi, M. (2020) · 2020
Later among the works it cites.
Lessons from archives: strategies for collecting sociocultural data in machine learning
Jo, E. S. and Gebru, T. (2020) · 2020
Later among the works it cites.
Identifying sensitive URLs at web-scale
Matic, S., Iordanou, C., Smaragdakis, G., and Laoutaris, N. (2020) · 2020
Later among the works it cites.
CCNet: Extracting high quality monolingual datasets from web crawl data
Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A., and Grave, É. (2020) · 2020
Later among the works it cites.
Persistent Anti-Muslim Bias in Large Language Models
Abid, A., Farooqi, M., and Zou, J. (2021) · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big
Bender, E., Gebru, T., McMillan-Major, A., et al. (2021) · 2021
Closest in time.
Investigating gender bias in BERT
Bhardwaj, R., Majumder, N., and Poria, S. (2021) · 2021
Closest in time.
Quality at a glance: An audit of web-crawled multilingual datasets
Caswell, I., Kreutzer, J., Wang, L., Wahab, A., van Esch, D., Ulzii-Orshikh, N., Tapo, A., Subramani, N., Sokolov, A., Sikasote, C., et al. (2021) · 2021
Closest in time.
Characterizing (un) moderated textual data in social systems
de Lima, L. H. C., Reis, J., Melo, P., Murai, F., and Benevenuto, F. (2021) · 2021
Closest in time.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Schick, T., Udupa, S., and Schütze, H. (2021) · 2021
Closest in time.
Putting humans in the natural language processing loop: A survey
Wang, Z. J., Choi, D., Xu, S., and Yang, D. (2021) · 2021
Closest in time.
Indiviuals using the Internet
World Bank (2018) · 2021
Closest in time.