Fetching the paper…
Reading the bibliography…
In the pursuit of developing Large Language Models (LLMs) that adhere to societal standards, it is imperative to detect the toxicity in the generated text.
The aggression questionnaire
A. H. Buss and M. Perry. 1992 · 1992
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
The brief aggression questionnaire: Psychometric and behavioral evidence for an efficient measure of trait aggression
Gregory D Webster, C Nathan DeWall, Richard S Pond Jr, Timothy Deckman, Peter K Jonason, Bonnie M Le, Austin Lee Nichols, Tatiana Orozco Schember, Laura C Crysel, Benjamin S Crosier, et al. 2014 · 2014
Earlier work this paper cites.
chrF++: words helping character n-grams
Maja Popović. 2017 · 2017
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2019 · 2019
Earlier work this paper cites.
Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them
Hila Gonen and Yoav Goldberg. 2019a · 2019
Earlier work this paper cites.
Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them
Hila Gonen and Yoav Goldberg. 2019b · 2019
Earlier work this paper cites.
The global landscape of ai ethics guidelines
Anna Jobin, Marcello Ienca, and Effy Vayena. 2019 · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Nurse is closer to woman than surgeon? mitigating gender-biased proximities in word embeddings
Vaibhav Kumar, Tenzin Singhay Bhotia, Vaibhav Kumar, and Tanmoy Chakraborty. 2020 · 2020
Earlier work this paper cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg. 2020 · 2020
Earlier work this paper cites.
Comet: A neural framework for mt evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Earlier work this paper cites.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed H. Chi, and Slav Petrov. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Earlier work this paper cites.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021 · 2021
Earlier work this paper cites.
Aligning {ai} with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Sustainable modular debiasing of language models
Anne Lauscher, Tobias Lueken, and Goran Glavaš. 2021 · 2021
Earlier work this paper cites.
UMIC: An unreferenced metric for image captioning via contrastive learning
Hwanhee Lee, Seunghyun Yoon, Franck Dernoncourt, Trung Bui, and Kyomin Jung. 2021 · 2021
Earlier work this paper cites.
Evaluating the robustness of neural language models to input perturbations
Milad Moradi and Matthias Samwald. 2021 · 2021
Earlier work this paper cites.
Sample selection for fair and robust training
Yuji Roh, Kangwook Lee, Steven Euijong Whang, and Changho Suh. 2021 · 2021
Earlier work this paper cites.
Evaluating gender bias in natural language inference
Shanya Sharma, Manan Dey, and Koustuv Sinha. 2021 · 2021
Earlier work this paper cites.
Introducing CAD: the contextual abuse dataset
Bertie Vidgen, Dong Nguyen, Helen Margetts, Patricia Rossini, and Rebekah Tromble. 2021 · 2021
Earlier work this paper cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Cited alongside, same era.
Entropy-based attention regularization frees unintended bias mitigation from lists
Giuseppe Attanasio, Debora Nozza, Dirk Hovy, and Elena Baralis. 2022 · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022 · 2022
Cited alongside, same era.
Debiasing pretrained text encoders by paying attention to paying attention
Yacine Gaci, Boualem Benatallah, Fabio Casati, and Khalid Benabdeslem. 2022a · 2022
Cited alongside, same era.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2023 · 2023
Later among the works it cites.
LLM-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models
Yen-Ting Lin and Yun-Nung Chen. 2023 · 2023
Later among the works it cites.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
Debiasing should be good and bad: Measuring the consistency of debiasing techniques in language models
Robert Morabito, Jad Kabbara, and Ali Emami. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Cited alongside, same era.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2022 · 2022
Cited alongside, same era.
Prosocialdialog: A prosocial backbone for conversational agents
Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, and Maarten Sap. 2022 · 2022
Cited alongside, same era.
Probing classifiers are unreliable for concept removal and detection
Abhinav Kumar, Chenhao Tan, and Amit Sharma. 2022 · 2022
Cited alongside, same era.
ParaDetox: Detoxification with parallel data
Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022 · 2022
Cited alongside, same era.
Who is GPT-3? an exploration of personality, values and demographics
Marilù Miotto, Nicola Rossberg, and Bennett Kleinberg. 2022 · 2022
Cited alongside, same era.
Choose your lenses: Flaws in gender bias evaluation
Hadas Orgad and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Gianluca Nogara, Francesco Pierri, Stefano Cresci, Luca Luceri, Petter Törnberg, and Silvia Giordano. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Three ways of using large language models to evaluate chat
Ondřej Plátek, Vojtěch Hudeček, Patricia Schmidtová, Mateusz Lango, and Ondřej Dušek. 2023 · 2023
Later among the works it cites.
On the challenges of using black-box APIs for toxicity evaluation in research
Luiza Pozzobon, Beyza Ermis, Patrick Lewis, and Sara Hooker. 2023 · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
Later among the works it cites.
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2023 · 2023
Later among the works it cites.
Causes and cures for interference in multilingual translation
Uri Shaham, Maha Elbayad, Vedanuj Goswami, Omer Levy, and Shruti Bhosale. 2023 · 2023
Later among the works it cites.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models
Google Gemini Team. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
Do-not-answer: A dataset for evaluating safeguards in llms
Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. 2023 · 2023
Later among the works it cites.
Large language models as optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. 2023 · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Later among the works it cites.
Rethinking machine ethics – can llms perform moral reasoning through the lens of moral theories?
Jingyan Zhou, Minda Hu, Junan Li, Xiaoying Zhang, Xixin Wu, Irwin King, and Helen Meng. 2023 · 2023
Later among the works it cites.
Mp2d: An automated topic shift dialogue generation framework leveraging knowledge graphs
Yerin Hwang, Yongil Kim, Yunah Jang, Jeesoo Bang, Hyunkyung Bae, and Kyomin Jung. 2024 · 2024
Closest in time.
LifeTox: Unveiling implicit toxicity in life advice
Minbeom Kim, Jahyun Koo, Hwanhee Lee, Joonsuk Park, Hwaran Lee, and Kyomin Jung. 2024a · 2024
Closest in time.
Fine-grained gender control in machine translation with large language models
Minwoo Lee, Hyukhun Koh, Minsung Kim, and Kyomin Jung. 2024 · 2024
Closest in time.
Llm theory of mind and alignment: Opportunities and risks
Winnie Street. 2024 · 2024
Closest in time.
Mitigating biases for instruction-following language models via bias neurons elimination
Nakyeong Yang, Taegwan Kang, Stanley Jungkyu Choi, Honglak Lee, and Kyomin Jung. 2024 · 2024
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2086
Closest in time.