Fetching the paper…
Reading the bibliography…
We present a holistic approach to building a robust and useful natural language classification system for real-world content moderation.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Augmenting data with mixup for sentence classification: An empirical study
Guo, H.; Mao, Y.; and Zhang, R. 2019 · 1905
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Dinan, E.; Humeau, S.; Chintagunta, B.; and Weston, J. 2019 · 1908
Earlier work this paper cites.
Query by committee
Seung, H. S.; Opper, M.; and Sompolinsky, H. 1992 · 1992
Earlier work this paper cites.
Heterogeneous uncertainty sampling for supervised learning
Lewis, D. D.; and Catlett, J. 1994 · 1994
Earlier work this paper cites.
A sequential algorithm for training text classifiers
Lewis, D. D.; and Gale, W. A. 1994 · 1994
Earlier work this paper cites.
Committee-based sampling for training probabilistic classifiers
Dagan, I.; and Engelson, S. P. 1995 · 1995
Earlier work this paper cites.
Employing EM and Pool-Based Active Learning for Text Classification
McCallum, A.; and Nigam, K. 1998 · 1998
Earlier work this paper cites.
Less is more: Active learning with support vector machines
Schohn, G.; and Cohn, D. 2000 · 2000
Earlier work this paper cites.
Active hidden markov models for information extraction
Scheffer, T.; Decomain, C.; and Wrobel, S. 2001 · 2001
Earlier work this paper cites.
Data augmentation using pre-trained transformer models
Kumar, V.; Choudhary, A.; and Cho, E. 2020 · 2003
Earlier work this paper cites.
Deep learning models for multilingual hate speech detection
Aluru, S. S.; Mathew, B.; Saha, P.; and Mukherjee, A. 2020 · 2004
Earlier work this paper cites.
Active learning using pre-clustering
Nguyen, H. T.; and Smeulders, A. 2004 · 2004
Earlier work this paper cites.
A Large-Scale Semi-Supervised Dataset for Offensive Language Identification
Rosenthal, S.; Atanasova, P.; Karadzhov, G.; Zampieri, M.; and Nakov, P. 2020 · 2004
Earlier work this paper cites.
Reducing labeling effort for structured prediction tasks
Culotta, A.; and McCallum, A. 2005 · 2005
Earlier work this paper cites.
Active learning to recognize multiple types of plankton
Luo, T.; Kramer, K.; Goldgof, D. B.; Hall, L. O.; Samson, S.; Remsen, A.; Hopkins, T.; and Cohn, D. 2005 · 2005
Earlier work this paper cites.
Active feedback in ad hoc information retrieval
Shen, X.; and Zhai, C. 2005 · 2005
Earlier work this paper cites.
Analysis of Representations for Domain Adaptation
Ben-David, S.; Blitzer, J.; Crammer, K.; and Pereira, F. C. 2006 · 2006
Earlier work this paper cites.
Domain Adaptation with Structural Correspondence Learning
Blitzer, J.; McDonald, R. T.; and Pereira, F. C. 2006 · 2006
Earlier work this paper cites.
Batch mode active learning and its application to medical image classification
Hoi, S. C.; Jin, R.; Zhu, J.; and Lyu, M. R. 2006 · 2006
Earlier work this paper cites.
ETHOS: an online hate speech detection dataset
Mollas, I.; Chrysopoulou, Z.; Karlos, S.; and Tsoumakas, G. 2020 · 2006
Earlier work this paper cites.
Toxicity detection: Does context really matter?
Pavlopoulos, J.; Sorensen, J.; Dixon, L.; Thain, N.; and Androutsopoulos, I. 2020 · 2006
Earlier work this paper cites.
Neural Unsupervised Domain Adaptation in NLP—A Survey
Ramponi, A.; and Plank, B. 2020 · 2006
Earlier work this paper cites.
Incorporating diversity and density in active learning for relevance feedback
Xu, Z.; Akella, R.; and Zhang, Y. 2007 · 2007
Earlier work this paper cites.
Domain Adaptation with Multiple Sources
Mansour, Y.; Mohri, M.; and Rostamizadeh, A. 2008 · 2008
Earlier work this paper cites.
An analysis of active learning strategies for sequence labeling tasks
Settles, B.; and Craven, M. 2008 · 2008
Earlier work this paper cites.
A theory of learning from different domains
Ben-David, S.; Blitzer, J.; Crammer, K.; Kulesza, A.; Pereira, F. C.; and Vaughan, J. W. 2009 · 2009
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S.; Gururangan, S.; Sap, M.; Choi, Y.; and Smith, N. A. 2020 · 2009
Earlier work this paper cites.
Shen, D.; Zheng, M.; Shen, Y.; Qu, Y.; and Chen, W. 2020 · 2009
Earlier work this paper cites.
Subjective Natural Language Problems: Motivations, Applications, Characterizations, and Implications
Ovesdotter Alm, C. 2011 · 2011
Earlier work this paper cites.
Fairness through Awareness
Dwork, C.; Hardt, M.; Pitassi, T.; Reingold, O.; and Zemel, R. 2012 · 2012
Earlier work this paper cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Vidgen, B.; Thrush, T.; Waseem, Z.; and Kiela, D. 2020 · 2012
Earlier work this paper cites.
Locate the Hate: Detecting Tweets against Blacks
Kwok, I.; and Wang, Y. 2013 · 2013
Cited alongside, same era.
Intriguing properties of neural networks
Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I. J.; and Fergus, R. 2013 · 2013
Cited alongside, same era.
Generative Adversarial Nets
Goodfellow, I. J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A. C.; and Bengio, Y. 2014 · 2014
Cited alongside, same era.
Unsupervised Domain Adaptation by Backpropagation
Ganin, Y.; and Lempitsky, V. S. 2015 · 2015
Cited alongside, same era.
Explaining and Harnessing Adversarial Examples
Goodfellow, I. J.; Shlens, J.; and Szegedy, C. 2015 · 2015
Cited alongside, same era.
Do not have enough data? Deep learning to the rescue!
Anaby-Tavor, A.; Carmeli, B.; Goldbraich, E.; Kantor, A.; Kour, G.; Shlomov, S.; Tepper, N.; and Zwerdling, N. 2020 · 2020
Later among the works it cites.
TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification
Barbieri, F.; Camacho-Collados, J.; Espinosa Anke, L.; and Neves, L. 2020 · 2020
Later among the works it cites.
Machine learning techniques for the detection of inappropriate erotic content in text
Barrientos, G. M.; Alaiz-Rodríguez, R.; González-Castro, V.; and Parnell, A. C. 2020 · 2020
Later among the works it cites.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
Ribeiro, M. T.; Wu, T.; Guestrin, C.; and Singh, S. 2020 · 2020
Later among the works it cites.
Advanced active learning strategies for object detection
Schmidt, S.; Rao, Q.; Tatsch, J.; and Knoll, A. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abusive Language Detection in Online User Content
Nobata, C.; Tetreault, J.; Thomas, A.; Mehdad, Y.; and Chang, Y. 2016 · 2016
Cited alongside, same era.
Are You a Racist or Am I Seeing Things? Annotator Influence on Hate Speech Detection on Twitter
Waseem, Z. 2016 · 2016
Cited alongside, same era.
A survey of transfer learning
Weiss, K. R.; Khoshgoftaar, T. M.; and Wang, D. 2016 · 2016
Cited alongside, same era.
Wasserstein Generative Adversarial Networks
Arjovsky, M.; Chintala, S.; and Bottou, L. 2017 · 2017
Cited alongside, same era.
Automated hate speech detection and the problem of offensive language
Davidson, T.; Warmsley, D.; Macy, M.; and Weber, I. 2017 · 2017
Cited alongside, same era.
Deep bayesian active learning with image data
Gal, Y.; Islam, R.; and Ghahramani, Z. 2017 · 2017
Cited alongside, same era.
Counterfactual Fairness
Kusner, M. J.; Loftus, J.; Russell, C.; and Silva, R. 2017 · 2017
Cited alongside, same era.
Vidgen, B.; and Derczynski, L. 2020 · 2020
Later among the works it cites.
Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations
Davani, A. M.; Díaz, M.; and Prabhakaran, V. 2021 · 2021
Later among the works it cites.
SimCSE: Simple Contrastive Learning of Sentence Embeddings
Gao, T.; Yao, X.; and Chen, D. 2021 · 2021
Later among the works it cites.
Dynabench: Rethinking Benchmarking in NLP
Kiela, D.; Bartolo, M.; Nie, Y.; Kaushik, D.; Geiger, A.; Wu, Z.; Vidgen, B.; Prasad, G.; Singh, A.; Ringshia, P.; Ma, Z.; Thrush, T.; Riedel, S.; Waseem, Z.; Stenetorp, P.; Jia, R.; Bansal, M.; Potts, C.; and Williams, A. 2021 · 2021
Later among the works it cites.
https://partnershiponai.org/paper/responsible-sourcing-considerations/
PAI. 2021 · 2021
Later among the works it cites.
On Releasing Annotator-Level Labels and Information in Datasets
Prabhakaran, V.; Mostafazadeh Davani, A.; and Diaz, M. 2021 · 2021
Later among the works it cites.
HateCheck: Functional Tests for Hate Speech Detection Models
Röttger, P.; Vidgen, B.; Nguyen, D.; Waseem, Z.; Margetts, H.; and Pierrehumbert, J. 2021 · 2021
Later among the works it cites.
Generating datasets with pretrained language models
Schick, T.; and Schütze, H. 2021 · 2021
Later among the works it cites.
Practical Transformer-based Multilingual Text Classification
Wang, C.; and Banko, M. 2021 · 2021
Later among the works it cites.
Want To Reduce Labeling Cost? GPT-3 Can Help
Wang, S.; Liu, Y.; Xu, Y.; Zhu, C.; and Zeng, M. 2021a · 2021
Later among the works it cites.
Ethical and social risks of harm from language models
Weidinger, L.; Mellor, J.; Rauh, M.; Griffin, C.; Uesato, J.; Huang, P.-S.; Cheng, M.; Glaese, M.; Balle, B.; Kasirzadeh, A.; et al. 2021 · 2021
Later among the works it cites.
Challenges in detoxifying language models
Welbl, J.; Glaese, A.; Uesato, J.; Dathathri, S.; Mellor, J.; Hendricks, L. A.; Anderson, K.; Kohli, P.; Coppin, B.; and Huang, P.-S. 2021 · 2021
Later among the works it cites.
Towards generalisable hate speech detection: a review on obstacles and solutions
Yin, W.; and Zubiaga, A. 2021 · 2021
Later among the works it cites.
GPT3Mix: Leveraging large-scale language models for text augmentation
Yoo, K. M.; Park, D.; Kang, J.; Lee, S.-W.; and Park, W. 2021 · 2021
Later among the works it cites.
Double Perturbation: On the Robustness of Robustness and Counterfactual Bias Evaluation
Zhang, C.; Zhao, J.; Zhang, H.; Chang, K.-W.; and Hsieh, C.-J. 2021 · 2021
Later among the works it cites.
LaMDA: Language Models for Dialog Applications
Cohen, A. D.; Roberts, A.; Molina, A.; Butryna, A.; Jin, A.; Kulshreshtha, A.; Hutchinson, B.; Zevenbergen, B.; Aguera-Arcas, B. H.; ching Chang, C.; Cui, C.; Du, C.; Adiwardana, D. D. F.; Chen, D.; Lepikhin, D. D.; Chi, E. H.; Hoffman-John, E.; Cheng, H.-T.; Lee, H.; Krivokon, I.; Qin, J.; Hall, J.; Fenton, J.; Soraker, J.; Meier-Hellstern, K.; Olson, K.; Aroyo, L. M.; Bosma, M. P.; Pickett, M. J.; Menegali, M. A.; Croak, M.; Díaz, M.; Lamm, M.; Krikun, M.; Morris, M. R.; Shazeer, N.; Le, Q. V.; Bernstein, R.; Rajakumar, R.; Kurzweil, R.; Thoppilan, R.; Zheng, S.; Bos, T.; Duke, T.; Doshi, T.; Prabhakaran, V.; Rusch, W.; Li, Y.; Huang, Y.; Zhou, Y.; Xu, Y.; and Chen, Z. 2022 · 2022
Closest in time.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Hartvigsen, T.; Gabriel, S.; Palangi, H.; Sap, M.; Ray, D.; and Kamar, E. 2022 · 2022
Closest in time.
Perspective API
Jigsaw. ???? · 2022
Closest in time.
Toxic Comment Classification Challenge
Jigsaw. 2018 · 2022
Closest in time.
A new generation of perspective api: Efficient multilingual character-level transformers
Lees, A.; Tran, V. Q.; Tay, Y.; Sorensen, J.; Gupta, J.; Metzler, D.; and Vasserman, L. 2022 · 2022
Closest in time.
DALL·E 2 Preview - Risks and Limitations
Mishkin, P.; Ahmad, L.; Brundage, M.; Krueger, G.; and Sastry, G. 2022 · 2022
Closest in time.
Red teaming language models with language models
Perez, E.; Huang, S.; Song, F.; Cai, T.; Ring, R.; Aslanides, J.; Glaese, A.; McAleese, N.; and Irving, G. 2022 · 2022
Closest in time.
Building Better Moderator Tools
Reddit. 2022 · 2022
Closest in time.
The Four Rs of Responsibility, Part 1: Removing Harmful Content
YouTube. 2019 · 2022
Closest in time.
Adversarial Training for High-Stakes Reliability
Ziegler, D. M.; Nix, S.; Chan, L.; Bauman, T.; Schmidt-Nielsen, P.; Lin, T.; Scherlis, A.; Nabeshima, N.; Weinstein-Raun, B.; de Haas, D.; et al. 2022 · 2022
Closest in time.
Domain-Adversarial Training of Neural Networks
Ganin, Y.; Ustinova, E.; Ajakan, H.; Germain, P.; Larochelle, H.; Laviolette, F.; Marchand, M.; and Lempitsky, V. 2016 · 2030
Closest in time.