Fetching the paper…
Reading the bibliography…
An important issue with Large Language Models (LLMs) is their undesired ability to generate toxic language.
Understanding interobserver agreement: the kappa statistic
Viera, A. J., Garrett, J. M., et al · 2005
Earlier work this paper cites.
Alignment for advanced machine learning systems
Taylor, J., Yudkowsky, E., LaVictoire, P., and Critch, A · 2016
Earlier work this paper cites.
Toxic comment classification challenge, 2017
Adams, C., Sorensen, J., Elliott, J., Dixon, L., McDonald, M., Nithum, and Cukierski, W · 2017
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Davidson, T., Warmsley, D., Macy, M., and Weber, I · 2017
Earlier work this paper cites.
Deceiving google’s perspective api built for detecting toxic comments
Hosseini, H., Kannan, S., Zhang, B., and Poovendran, R · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Radford, A., Jozefowicz, R., and Sutskever, I · 2017
Earlier work this paper cites.
Hurtlex: A multilingual lexicon of words to hurt
Bassignana, E., Basile, V., Patti, V., et al · 2018
Earlier work this paper cites.
Large scale crowdsourcing and characterization of twitter abusive behavior
Founta, A., Djouvas, C., Chatzakou, D., Leontiadis, I., Blackburn, J., Stringhini, G., Vakali, A., Sirivianos, M., and Kourtellis, N · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Evaluating the underlying gender bias in contextualized word embeddings
Basta, C. R. S., Ruiz Costa-Jussà, M., and Casas Manzanares, N · 2019
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R · 2019
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R · 2019
Earlier work this paper cites.
Measuring bias in contextualized word representations
Kurita, K., Vyas, N., Pareek, A., Black, A. W., and Tsvetkov, Y · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
May, C., Wang, A., Bordia, S., Bowman, S. R., and Rudinger, R · 2019
Earlier work this paper cites.
Multilingual and multi-aspect hate speech analysis
Ousidhoum, N., Lin, Z., Zhang, H., Song, Y., and Yeung, D.-Y · 2019
Earlier work this paper cites.
Socialiqa: Commonsense reasoning about social interactions
Sap, M., Rashkin, H., Chen, D., LeBras, R., and Choi, Y · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
Predicting the type and target of offensive posts in social media
Zampieri, M., Malmasi, S., Nakov, P., Rosenthal, S., Farra, N., and Kumar, R · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Cited alongside, same era.
Gender bias in contextualized word embeddings
Zhao, J., Wang, T., Yatskar, M., Cotterell, R., Ordonez, V., and Chang, K.-W · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Cited alongside, same era.
Language models are few-shot learners
HateCheck: Functional tests for hate speech detection models
Röttger, P., Vidgen, B., Nguyen, D., Waseem, Z., Margetts, H., and Pierrehumbert, J · 2021
Later among the works it cites.
Challenges in detoxifying language models
Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S · 2021
Later among the works it cites.
Fudge: Controlled text generation with future discriminators
Yang, K. and Klein, D · 2021
Later among the works it cites.
The cringe loss: Learning what language not to model
Adolphs, L., Gao, T., Xu, J., Shuster, K., Sukhbaatar, S., and Weston, J · 2022
Later among the works it cites.
Holy $#!t: Are popular toxicity models simply profanity detectors?, 2022
Chen, E · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Cited alongside, same era.
Gedi: Generative discriminator guided sequence generation
Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F · 2020
Cited alongside, same era.
Stereoset: Measuring stereotypical bias in pretrained language models
Nadeem, M., Bethke, A., and Reddy, S · 2020
Cited alongside, same era.
Crows-pairs: A challenge dataset for measuring social biases in masked language models
Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Recipes for safety in open-domain chatbots
Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Dinan, E · 2020
Cited alongside, same era.
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Later among the works it cites.
Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models
Delobelle, P., Tokpo, E., Calders, T., and Berendt, B · 2022
Later among the works it cites.
Ju, D., Xu, J., Boureau, Y.-L., and Weston, J · 2022
Later among the works it cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Later among the works it cites.
A robustly optimized BMRC for aspect sentiment triplet extraction
Liu, S., Li, K., and Li, Z · 2022
Later among the works it cites.
Measuring harmful sentence completion in language models for LGBTQIA+ individuals
Nozza, D., Bianchi, F., Lauscher, A., and Hovy, D · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Ignore previous prompt: Attack techniques for language models
Perez, F. and Ribeiro, I · 2022
Later among the works it cites.
Self-conditioning pre-trained language models
Suau, X., Zappella, L., and Apostoloff, N · 2022
Later among the works it cites.
Falcon-40B: an open large language model with state-of-the-art performance
Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, E., Heslow, D., Launay, J., Malartic, Q., Noune, B., Pannier, B., and Penedo, G · 2023
Later among the works it cites.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Later among the works it cites.
Pretraining language models with human preferences
Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Fundamental limitations of alignment in large language models
Wolf, Y., Wies, N., Levine, Y., and Shashua, A · 2023
Later among the works it cites.