Fetching the paper…
Reading the bibliography…
In this paper, we explore the feasibility of leveraging large language models (LLMs) to automate or otherwise assist human raters with identifying harmful content including hate speech, harassment, violent extremism, and election misinformation.
Pathways to violent extremism in the digital era
C. Edwards and L. Gribbon · 2013
Earlier work this paper cites.
What is a flag for? social media reporting tools and the vocabulary of complaint
K. Crawford and T. Gillespie · 2016
Earlier work this paper cites.
Characterizations of online harassment: Comparing policies across social media platforms
J. A. Pater, M. K. Kim, E. D. Mynatt, and C. Fiesler · 2016
Earlier work this paper cites.
Analyzing the targets of hate in online social media
L. Silva, M. Mondal, D. Correa, F. Benevenuto, and I. Weber · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
T. Davidson, D. Warmsley, M. Macy, and I. Weber · 2017
Earlier work this paper cites.
Update on the Global Internet Forum to Counter Terrorism
Google · 2017
Earlier work this paper cites.
Toxic comment classification challenge
Jigsaw · 2017
Earlier work this paper cites.
A measurement study of hate speech in social media
M. Mondal, L. A. Silva, and F. Benevenuto · 2017
Earlier work this paper cites.
A survey on automatic detection of hate speech in text
P. Fortuna and S. Nunes · 2018
Earlier work this paper cites.
Censored, suspended, shadowbanned: User interpretations of content moderation on social media platforms
S. Myers West · 2018
Earlier work this paper cites.
Anatomy of online hate: Developing a taxonomy and machine learning models for identifying and classifying hate in online news media
J. O. Salminen, H. Almerekhi, M. Milenkovic, S.-G. Jung, J. An, H. Kwak, and B. Jansen · 2018
Earlier work this paper cites.
On the origins of memes by means of fringe web communities
S. Zannettou, T. Caulfield, J. Blackburn, E. De Cristofaro, M. Sirivianos, G. Stringhini, and G. Suarez-Tangil · 2018
Earlier work this paper cites.
Rethinking the detection of child sexual abuse imagery on the internet
E. Bursztein, E. Clarke, M. DeLaune, D. M. Elifff, N. Hsu, L. Olson, J. Shehan, M. Thakur, K. Thomas, and T. Bright · 2019
Earlier work this paper cites.
Racial bias in hate speech and abusive language detection datasets
T. Davidson, D. Bhattacharya, and I. Weber · 2019
Earlier work this paper cites.
Predicting the type and target of offensive posts in social media
M. Zampieri, S. Malmasi, P. Nakov, S. Rosenthal, N. Farra, and R. Kumar · 2019
Earlier work this paper cites.
Algorithmic content moderation: Technical and political challenges in the automation of platform governance
R. Gorwa, R. Binns, and C. Katzenbach · 2020
Earlier work this paper cites.
Algorithmic content moderation: Technical and political challenges in the automation of platform governance
R. Gorwa, R. Binns, and C. Katzenbach · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
ETHOS: an online hate speech detection dataset
I. Mollas, Z. Chrysopoulou, S. Karlos, and G. Tsoumakas · 2020
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models, 2020
M. Nadeem, A. Bethke, and S. Reddy · 2020
Cited alongside, same era.
Prevention, disruption and deterrence of online child sexual exploitation and abuse
E. Quayle · 2020
Cited alongside, same era.
“at the end of the day facebook does what it wants” how users experience contesting algorithmic content moderation
K. Vaccaro, C. Sandvig, and K. Karahalios · 2020
Cited alongside, same era.
Using terms and conditions to apply fundamental rights to content moderation: Is article 12 dsa a paper tiger?
N. Appelman, J. P. Quintais, and R. Fahy · 2021
Cited alongside, same era.
Detecting hate speech with GPT-3
K.-L. Chiu, A. Collins, and R. Alexander · 2021
Cited alongside, same era.
Designing toxic content classification for a diversity of perspectives
Mumin: A large-scale multilingual multimodal fact-checked misinformation social network dataset
D. S. Nielsen and R. McConville · 2022
Later among the works it cites.
Ignore previous prompt: Attack techniques for language models
F. Perez and I. Ribeiro · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Later among the works it cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
S. Abdelnabi, K. Greshake, S. Mishra, C. Endres, T. Holz, and M. Fritz · 2023
Later among the works it cites.
Palm 2 for text
Google · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Kumar, P. G. Kelley, S. Consolvo, J. Mason, E. Bursztein, Z. Durumeric, K. Thomas, and M. Bailey · 2021
Cited alongside, same era.
Sustainable modular debiasing of language models
A. Lauscher, T. Lueken, and G. Glavaš · 2021
Cited alongside, same era.
Measuring the impact of covid-19 vaccine misinformation on vaccination intent in the uk and usa
S. Loomba, A. de Figueiredo, S. J. Piatek, K. de Graaf, and H. J. Larson · 2021
Cited alongside, same era.
An empirical survey of the effectiveness of debiasing techniques for pre-trained language models
N. Meade, E. Poole-Dayan, and S. Reddy · 2021
Cited alongside, same era.
A framework of severity for harmful content online
M. K. Scheuerman, J. A. Jiang, C. Fiesler, and J. R. Brubaker · 2021
Cited alongside, same era.
The psychological well-being of content moderators: the emotional labor of commercial moderation and avenues for improving support
M. Steiger, T. J. Bharucha, S. Venkatagiri, M. J. Riedl, and M. Lease · 2021
Cited alongside, same era.
Sok: Hate, harassment, and the changing landscape of online abuse
K. Thomas, D. Akhawe, M. Bailey, D. Boneh, E. Bursztein, S. Consolvo, N. Dell, Z. Durumeric, P. G. Kelley, D. Kumar, et al · 2021
Cited alongside, same era.
F. Huang, H. Kwak, and J. An · 2023
Later among the works it cites.
Personalizing content moderation on social media: User perspectives on moderation choices, interface design, and labor
S. Jhaver, A. Q. Zhang, Q. Z. Chen, N. Natarajan, R. Wang, and A. X. Zhang · 2023
Later among the works it cites.
Lost in the middle: How language models use long contexts
N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang · 2023
Later among the works it cites.
Automatic prompt optimization with “gradient descent” and beam search
R. Pryzant, D. Iter, J. Li, Y. T. Lee, C. Zhu, and M. Zeng · 2023
Later among the works it cites.
Sok: Content moderation in social media, from guidelines to enforcement, and research to practice
M. Singhal, C. Ling, P. Paudel, P. Thota, N. Kumarswamy, G. Stringhini, and S. Nilizadeh · 2023
Later among the works it cites.
Large language models as optimizers
C. Yang, X. Wang, Y. Lu, H. Liu, Q. V. Le, D. Zhou, and X. Chen · 2023
Later among the works it cites.
Tempera: Test-time prompt editing via reinforcement learning
T. Zhang, X. Wang, D. Zhou, D. Schuurmans, and J. E. Gonzalez · 2023
Later among the works it cites.
Large language models are human-level prompt engineers
Y. Zhou, A. I. Muresanu, Z. Han, K. Paster, S. Pitis, H. Chan, and J. Ba · 2023
Later among the works it cites.
Learn about policies across Google
Google · 2024
Closest in time.
Moderating text
Google · 2024
Closest in time.
You only prompt once: On the capabilities of prompt learning on large language models to tackle toxic content
X. He, S. Zannettou, Y. Shen, and Y. Zhang · 2024
Closest in time.
Policies
Meta · 2024
Closest in time.
Harm categories in Azure AI content safety
Microsoft · 2024
Closest in time.
Moderating new waves of online hate with chain-of-thought reasoning in large language models
N. Vishwamitra, K. Guo, F. T. Romit, I. Ondracek, L. Cheng, Z. Zhao, and H. Hu · 2024
Closest in time.
Using GPT-4 for content moderation
L. Weng, V. Goel, and A. Vallone · 2024
Closest in time.