Fetching the paper…
Reading the bibliography…
Sophisticated language models such as OpenAI's GPT-3 can generate hateful text that targets marginalized groups.
Two decades of statistical language modeling: Where do we go from here?
Rosenfeld, R. (2000) · 2000
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. (2003) · 2003
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 2005
Earlier work this paper cites.
The radicalization risks of GPT-3 and advanced neural language models
McGuffie, K. and Newhouse, A. (2020) · 2009
Earlier work this paper cites.
Word representations: A simple and general method for semi-supervised learning
Turian, J., Ratinov, L., and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
The social impact of natural language processing
Hovy, D. and Spruit, S. L. (2016) · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? Predictive features for hate speech detection on Twitter
Waseem, Z. and Hovy, D. (2016) · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Davidson, T., Warmsley, D., Macy, M., and Weber, I. (2017) · 2017
Earlier work this paper cites.
A survey on hate speech detection using natural language processing
Schmidt, A. and Wiegand, M. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Cited alongside, same era.
Hatebusters: A Web Application for Actively Reporting YouTube Hate Speech
Anagnostou, A., Mollas, I., and Tsoumakas, G. (2018) · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Cited alongside, same era.
The gab hate corpus: A collection of 27k posts annotated for hate speech
Kennedy, B., Atari, M., Davani, A. M., Yeh, L., Omrani, A., Kim, Y., Coombs, K., Havaldar, S., Portillo-Wightman, G., Gonzalez, E., et al. (2018) · 2018
Cited alongside, same era.
Racial bias in hate speech and abusive language detection datasets
Davidson, T., Bhattacharya, D., and Weber, I. (2019) · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021) · 2021
Closest in time.
Government of Canada
Criminal Code (1985) · 2021
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Fedus, W., Zoph, B., and Shazeer, N. (2021) · 2021
Closest in time.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O. (2021) · 2021
Closest in time.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Cited alongside, same era.
The pushshift reddit dataset
Baumgartner, J., Zannettou, S., Keegan, B., Squire, M., and Blackburn, J. (2020) · 2020
Cited alongside, same era.
ETHOS: An Online Hate Speech Detection Dataset
Mollas, I., Chrysopoulou, Z., Karlos, S., and Tsoumakas, G. (2020) · 2020
Cited alongside, same era.
Schick, T., Udupa, S., and Schütze, H. (2021) · 2021
Closest in time.
Addressing hate speech with data science: An overview from computer science perspective
Srba, I., Lenzini, G., Pikuliak, M., and Pecar, S. (2021) · 2021
Closest in time.
Hateful conduct policy
Twitter (2021) · 2021
Closest in time.