Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have exploded in popularity due to their ability to perform a wide array of natural language tasks.
The Internet’s hidden rules: An empirical study of Reddit norm violations at micro, meso, and macro scales
Chandrasekharan, E.; Samory, M.; Jhaver, S.; Charvat, H.; Bruckman, A.; Lampe, C.; Eisenstein, J.; and Gilbert, E. 2018 · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Dixon, L.; Li, J.; Sorensen, J.; Thain, N.; and Vasserman, L. 2018 · 2018
Earlier work this paper cites.
Reddit rules! characterizing an ecosystem of governance
Fiesler, C.; Jiang, J.; McCann, J.; Frye, K.; and Brubaker, J. 2018 · 2018
Earlier work this paper cites.
Perspective API
Google Jigsaw. 2018 · 2018
Earlier work this paper cites.
Conversations gone awry: Detecting early signs of conversational failure
Zhang, J.; Chang, J. P.; Danescu-Niculescu-Mizil, C.; Dixon, L.; Hua, Y.; Thain, N.; and Taraborelli, D. 2018 · 2018
Earlier work this paper cites.
Crossmod: A cross-community learning-based system to assist reddit moderators
Chandrasekharan, E.; Gandhi, C.; Mustelier, M. W.; and Gilbert, E. 2019 · 2019
Earlier work this paper cites.
Does transparency in moderation really matter? User behavior after content removal explanations on reddit
Jhaver, S.; Bruckman, A.; and Gilbert, E. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 2020
Earlier work this paper cites.
”I run the world’s largest historical outreach project and it’s on a cesspool of a website.” Moderating a Public Scholarship Site on Reddit: A Case Study of r/AskHistorians
Gilbert, S. A. 2020 · 2020
Earlier work this paper cites.
Content moderation, AI, and the question of scale
Gillespie, T. 2020 · 2020
Earlier work this paper cites.
Toxicity Detection: Does Context Really Matter?
Pavlopoulos, J.; Sorensen, J.; Dixon, L.; Thain, N.; and Androutsopoulos, I. 2020 · 2020
Earlier work this paper cites.
Exploring antecedents and consequences of toxicity in online discussions: A case study on reddit
Xia, Y.; Zhu, H.; Lu, T.; Zhang, P.; and Gu, N. 2020 · 2020
Earlier work this paper cites.
Detecting hate speech with gpt-3
Chiu, K.-L.; Collins, A.; and Alexander, R. 2021 · 2021
Earlier work this paper cites.
Latent hatred: A benchmark for understanding implicit hate speech
ElSherief, M.; Ziems, C.; Muchlinski, D.; Anupindi, V.; Seybolt, J.; De Choudhury, M.; and Yang, D. 2021 · 2021
Earlier work this paper cites.
Mitigating racial biases in toxic language detection with an equity-based ensemble framework
Halevy, M.; Harris, C.; Bruckman, A.; Yang, D.; and Howard, A. 2021 · 2021
Earlier work this paper cites.
Designing Toxic Content Classification for a Diversity of Perspectives
Kumar, D.; Kelley, P. G.; Consolvo, S.; Mason, J.; Bursztein, E.; Durumeric, Z.; Thomas, K.; and Bailey, M. 2021 · 2021
Cited alongside, same era.
How Multi-Layered Moderation Influences Positive Commenting
OpenWeb. 2021 · 2021
Cited alongside, same era.
The structure of toxic conversations on Twitter
Saveski, M.; Roy, B.; and Roy, D. 2021 · 2021
Cited alongside, same era.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Schick, T.; Udupa, S.; and Schütze, H. 2021 · 2021
Cited alongside, same era.
Human-ai collaboration via conditional delegation: A case study of content moderation
Lai, V.; Carton, S.; Bhatnagar, R.; Liao, Q. V.; Zhang, Y.; and Tan, C. 2022 · 2022
Cited alongside, same era.
Conversational Resilience: Quantifying and Predicting Conversational Outcomes Following Adverse Events
Chatgpt outperforms crowd-workers for text-annotation tasks
Gilardi, F.; Alizadeh, M.; and Kubli, M. 2023 · 2023
Closest in time.
Hate raids on Twitch: Echoes of the past, new modalities, and implications for platform governance
Han, C.; Seering, J.; Kumar, D.; Hancock, J. T.; and Durumeric, Z. 2023 · 2023
Closest in time.
Twits, Toxic Tweets, and Tribal Tendencies: Trends in Politically Polarized Posts on Twitter
Hanley, H. W.; and Durumeric, Z. 2023 · 2023
Closest in time.
A trade-off-centered framework of content moderation
Jiang, J. A.; Nie, P.; Brubaker, J. R.; and Fiesler, C. 2023 · 2023
Closest in time.
Understanding the behaviors of toxic accounts on reddit
Kumar, D.; Hancock, J.; Thomas, K.; and Durumeric, Z. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lambert, C.; Rajagopal, A.; and Chandrasekharan, E. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
An In-depth Look at Gemini’s Language Abilities
Akter, S. N.; Yu, Z.; Muhamed, A.; Ou, T.; Bäuerle, A.; Cabrera, Á. A.; Dholakia, K.; Xiong, C.; and Neubig, G. 2023 · 2023
Cited alongside, same era.
Leveraging AI for democratic discourse: Chat interventions can improve online political conversations at scale
Argyle, L. P.; Bail, C. A.; Busby, E. C.; Gubler, J. R.; Howe, T.; Rytting, C.; Sorensen, T.; and Wingate, D. 2023 · 2023
Cited alongside, same era.
AppealMod: Shifting Effort from Moderators to Users Making Appeals
Atreja, S.; Im, J.; Resnick, P.; and Hemphill, L. 2023 · 2023
Cited alongside, same era.
https://the-decoder.com/gpt-4-has-a-trillion-parameters
Bastian, M. 2023 · 2023
Cited alongside, same era.
How is ChatGPT’s behavior changing over time?
Chen, L.; Zaharia, M.; and Zou, J. 2023 · 2023
Cited alongside, same era.
The Unsung Heroes of Facebook Groups Moderation: A Case Study of Moderation Practices and Tools
Kuo, T.; Hernani, A.; and Grossklags, J. 2023 · 2023
Closest in time.
Li, L.; Fan, L.; Atreja, S.; and Hemphill, L. 2023 · 2023
Closest in time.
Defaulting to boilerplate answers, they didn’t engage in a genuine conversation: Dimensions of Transparency Design in Creator Moderation
Ma, R.; and Kou, Y. 2023 · 2023
Closest in time.
How Good Is ChatGPT For Detecting Hate Speech In Portuguese?
Oliveira, A. S.; Cecote, T. C.; Silva, P. H.; Gertrudes, J. C.; Freitas, V. L.; and Luz, E. J. 2023 · 2023
Closest in time.
Prompt Engineering
OpenAI. 2023 · 2023
Closest in time.
Törnberg, P. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Closest in time.
A prompt pattern catalog to enhance prompt engineering with chatgpt
White, J.; Fu, Q.; Hays, S.; Sandborn, M.; Olea, C.; Gilbert, H.; Elnashar, A.; Spencer-Smith, J.; and Schmidt, D. C. 2023 · 2023
Closest in time.
You Only Prompt Once: On the Capabilities of Prompt Learning on Large Language Models to Tackle Toxic Content
He, X.; Zannettou, S.; Shen, Y.; and Zhang, Y. 2024 · 2024
Closest in time.