Fetching the paper…
Reading the bibliography…
Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Conversational ai: Social and ethical considerations
Elayne Ruane, Abeba Birhane, and Anthony Ventresque · 2019
Earlier work this paper cites.
Modeling annotator perspective and polarized opinions to improve hate speech detection
Sohail Akhtar, Valerio Basile, and Viviana Patti · 2020
Earlier work this paper cites.
Identifying and measuring annotator bias based on annotators’ demographic characteristics
Hala Al Kuwatly, Maximilian Wich, and Georg Groh · 2020
Earlier work this paper cites.
Toxicity detection: Does context really matter?
John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain, and Ion Androutsopoulos · 2020
Earlier work this paper cites.
Open-domain conversational agents: Current progress, open problems, and future directions, 2020
Stephen Roller, Y-Lan Boureau, Jason Weston, Antoine Bordes, Emily Dinan, Angela Fan, David Gunning, Da Ju, Margaret Li, Spencer Poff, Pratik Ringshia, Kurt Shuster, Eric Michael Smith, Arthur Szlam, Jack Urbanek, and Mary Williamson · 2020
Earlier work this paper cites.
Investigating annotator bias with a graph-based approach
Maximilian Wich, Hala Al Kuwatly, and Georg Groh · 2020
Earlier work this paper cites.
Sohail Akhtar, Valerio Basile, and Viviana Patti · 2021
Earlier work this paper cites.
Ground-truth, whose truth? – examining the challenges with annotating toxic text datasets, 2021
Kofi Arhin, Ioana Baldini, Dennis Wei, Karthikeyan Natesan Ramamurthy, and Moninder Singh · 2021
Earlier work this paper cites.
Toward a perspectivist turn in ground truthing for predictive computing, 2021
Valerio Basile, Federico Cabitza, Andrea Campagner, and Michael Fell · 2021
Earlier work this paper cites.
Toward a perspectivist turn in ground truthing for predictive computing
Valerio Basile, Federico Cabitza, Andrea Campagner, and Michael Fell · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Emily Denton, Mark Díaz, Ian Kivlichan, Vinodkumar Prabhakaran, and Rachel Rosen · 2021
Earlier work this paper cites.
Anticipating safety issues in e2e conversational ai: Framework and tooling
Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser · 2021
Earlier work this paper cites.
Mitigating racial biases in toxic language detection with an equity-based ensemble framework
Matan Halevy, Camille Harris, Amy Bruckman, Diyi Yang, and Ayanna Howard · 2021
Earlier work this paper cites.
Assessing biases, relaxing moralism: On ground-truthing practices in machine learning design and application
Florian Jaton · 2021
Cited alongside, same era.
Offensive, aggressive, and hate speech analysis: From data-centric to human-centered approach
Jan Kocoń, Alicja Figas, Marcin Gruza, Daria Puchalska, Tomasz Kajdanowicz, and Przemysław Kazienko · 2021
Cited alongside, same era.
Learning personal human biases and representations for subjective tasks in natural language processing
Jan Kocoń, Marcin Gruza, Julita Bielaniewicz, Damian Grimling, Kamil Kanclerz, Piotr Miłkowski, and Przemysław Kazienko · 2021
Cited alongside, same era.
Agreeing to disagree: Annotating offensive language datasets with annotators’ disagreement
Elisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini, and Sara Tonelli · 2021
Cited alongside, same era.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee · 2021
Unsolved problems in ml safety, 2022
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt · 2022
Later among the works it cites.
Mitigating toxic degeneration with empathetic data: Exploring the relationship between toxicity and empathy
Allison Lahnala, Charles Welch, Béla Neuendorf, and Lucie Flek · 2022
Later among the works it cites.
Red teaming language models with language models, 2022
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Later among the works it cites.
Detecting unintended social bias in toxic language datasets, 2022
Nihar Sahoo, Himanshu Gupta, and Pushpak Bhattacharyya · 2022
Later among the works it cites.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection, 2022
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On releasing annotator-level labels and information in datasets
Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz · 2021
Cited alongside, same era.
Two contrasting data annotation paradigms for subjective nlp tasks
Paul Röttger, Bertie Vidgen, Dirk Hovy, and Janet B Pierrehumbert · 2021
Cited alongside, same era.
Process for adapting language models to society (palms) with values-targeted datasets
Irene Solaiman and Christy Dennison · 2021
Cited alongside, same era.
Investigating annotator bias in abusive language datasets
Maximilian Wich, Christian Widmer, Gerhard Hagerer, and Georg Groh · 2021
Cited alongside, same era.
Teach me to explain: A review of datasets for explainable natural language processing
Sarah Wiegreffe and Ana Marasovic · 2021
Cited alongside, same era.
Recipes for safety in open-domain chatbots, 2021
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan · 2021
Cited alongside, same era.
A literature review of textual hate speech detection methods and datasets
Fatimah Alkomah and Xiaogang Ma · 2022
Cited alongside, same era.
Why so toxic? measuring and triggering toxic behavior in open-domain chatbots
Wai Man Si, Michael Backes, Jeremy Blackburn, Emiliano De Cristofaro, Gianluca Stringhini, Savvas Zannettou, and Yang Zhang · 2022
Later among the works it cites.
On the safety of conversational models: Taxonomy, dataset, and benchmark, 2022
Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang · 2022
Later among the works it cites.
Lamda: Language models for dialog applications, 2022
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar akaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le · 2022
Later among the works it cites.
Toxicity detection sensitive to conversational context
Alexandros Xenos, John Pavlopoulos, Ion Androutsopoulos, Lucas Dixon, Jeffrey Sorensen, and Léo Laugier · 2022
Later among the works it cites.
Adversarial training for high-stakes reliability, 2022
Daniel M. Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel de Haas, Buck Shlegeris, and Nate Thomas · 2022
Later among the works it cites.
A drop of ink may make a million think: The spread of false information in large language models
Ning Bian, Peilin Liu, Xianpei Han, Hongyu Lin, Yaojie Lu, Ben He, and Le Sun · 2023
Closest in time.
Recent advances towards safe, responsible, and moral dialogue systems: A survey, 2023
Jiawen Deng, Hao Sun, Zhexin Zhang, Jiale Cheng, and Minlie Huang · 2023
Closest in time.
Is chatgpt better than human annotators? potential and limitations of chatgpt in explaining implicit hate speech
Fan Huang, Haewoon Kwak, and Jisun An · 2023
Closest in time.
Why don’t you do it right? analysing annotators’ disagreement in subjective tasks
Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Ježek · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Closest in time.