Fetching the paper…
Reading the bibliography…
Models trained on large unlabeled corpora of human interactions will learn patterns and mimic behaviors therein, which include offensive or otherwise toxic behavior and unwanted biases.
Semeval-2019 task 6: Identifying and categorizing offensive language in social media (offenseval)
Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019 · 1903
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 1904
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 1908
Earlier work this paper cites.
A crowd-based evaluation of abuse response strategies in conversational agents
Amanda Cercas Curry and Verena Rieser. 2019 · 1909
Earlier work this paper cites.
CTRL: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019 · 1909
Earlier work this paper cites.
Does gender matter? towards fairness in dialogue systems
Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2019 · 1910
Earlier work this paper cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019 · 1910
Earlier work this paper cites.
Queens are powerful too: Mitigating gender bias in dialogue generation
Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, and Jason Weston. 2019a · 1911
Earlier work this paper cites.
Eric Michael Smith, Diana Gonzalez-Rico, Emily Dinan, and Y-Lan Boureau. 2019 · 1911
Earlier work this paper cites.
DialoGPT: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2019 · 1911
Earlier work this paper cites.
Plug and play language models: a simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019 · 1912
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020 · 2001
Earlier work this paper cites.
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020 · 2001
Earlier work this paper cites.
Reliability in content analysis: Some common misconceptions and recommendations
Klaus Krippendorff. 2004 · 2004
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M Smith, et al. 2020 · 2004
Earlier work this paper cites.
Language (technology) is power: A critical survey of" bias" in nlp
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2005
Earlier work this paper cites.
Stupid computer! abuse and social identities
Antonella De Angeli and Rollo Carpenter. 2005 · 2005
Earlier work this paper cites.
Multi-dimensional gender bias classification
Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, and Adina Williams. 2020 · 2005
Earlier work this paper cites.
Cyberbullying detection with fairness constraints
Oguzhan Gencoglu. 2020 · 2005
Earlier work this paper cites.
Chat as expected: Learning to manipulate black-box neural dialogue models
Haochen Liu, Zhiwei Wang, Tyler Derr, and Jiliang Tang. 2020 · 2005
Earlier work this paper cites.
Demoting racial bias in hate speech detection
Mengzhou Xia, Anjalie Field, and Yulia Tsvetkov. 2020 · 2005
Earlier work this paper cites.
Marcos Zampieri, Preslav Nakov, Sara Rosenthal, Pepa Atanasova, Georgi Karadzhov, Hamdy Mubarak, Leon Derczynski, Zeses Pitenis, and Çağrı Çöltekin. 2020 · 2006
Earlier work this paper cites.
I hate you! disinhibition with virtual partners
Antonella De Angeli and Sheryl Brahnam. 2008 · 2008
Cited alongside, same era.
Neural generation meets real people: Towards emotionally engaging mixed-initiative conversations
Ashwin Paranjape, Abigail See, Kathleen Kenealy, Haojun Li, Amelia Hardy, Peng Qi, Kaushik Ram Sadagopan, Nguyet Minh Phu, Dilara Soylu, and Christopher D Manning. 2020 · 2008
Cited alongside, same era.
Detecting and classifying malevolent dialogue responses: Taxonomy, data and methodology
Yangjun Zhang, Pengjie Ren, and Maarten de Rijke. 2020 · 2008
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Sam Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2009
Cited alongside, same era.
Polite dialogue generation without parallel data
Tong Niu and Mohit Bansal. 2018 · 2018
Later among the works it cites.
Reducing gender bias in abusive language detection
Ji Ho Park, Jamin Shin, and Pascale Fung. 2018 · 2018
Later among the works it cites.
Fighting offensive language on social media with unsupervised text style transfer
Cicero Nogueira dos Santos, Igor Melnyk, and Inkit Padhi. 2018 · 2018
Later among the works it cites.
Engaging image chat: Modeling personality in grounded dialogue
Kurt Shuster, Samuel Humeau, Antoine Bordes, and Jason Weston. 2018 · 2018
Later among the works it cites.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alon Halevy, Cristian Canton Ferrer, Hao Ma, Umut Ozertem, Patrick Pantel, Marzieh Saeidi, Fabrizio Silvestri, and Ves Stoyanov. 2020 · 2009
Cited alongside, same era.
Gedi: Generative discriminator guided sequence generation
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2020 · 2009
Cited alongside, same era.
Judgment of the humanness of an interlocutor is in the eye of the beholder
Catherine L Lortie and Matthieu J Guitton. 2011 · 2011
Cited alongside, same era.
Real conversations with artificial intelligence: A comparison between human–human online conversations and human–chatbot conversations
Jennifer Hill, W Randolph Ford, and Ingrid G Farreras. 2015 · 2015
Cited alongside, same era.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
Automatic detection of hate speech in text: an overview of the topic and dataset annotation with hierarchical classes
Paula Cristina Teixeira Fortuna. 2017 · 2017
Cited alongside, same era.
Detecting the hate code on social media
Rijul Magu, Kshitij Joshi, and Jiebo Luo. 2017 · 2017
Cited alongside, same era.
ParlAI: A dialog research software platform
Alexander Miller, Will Feng, Dhruv Batra, Antoine Bordes, Adam Fisch, Jiasen Lu, Devi Parikh, and Jason Weston. 2017a · 2017
Cited alongside, same era.
Later among the works it cites.
Should an agent be ignoring it? a study of verbal abuse types and conversational agents’ response styles
Hyojin Chin and Mun Yong Yi. 2019 · 2019
Later among the works it cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019b · 2019
Later among the works it cites.
A unified deep learning architecture for abuse detection
Antigoni Maria Founta, Despoina Chatzakou, Nicolas Kourtellis, Jeremy Blackburn, Athena Vakali, and Ilias Leontiadis. 2019 · 2019
Later among the works it cites.
ACUTE-EVAL: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li, Jason Weston, and Stephen Roller. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019 · 2019
Later among the works it cites.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A Smith. 2019 · 2019
Later among the works it cites.
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019 · 2019
Later among the works it cites.
Studying generalisability across abusive language detection datasets
Steve Durairaj Swamy, Anupam Jamatia, and Björn Gambäck. 2019 · 2019
Later among the works it cites.
Challenges and frontiers in abusive content detection
Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts. 2019 · 2019
Later among the works it cites.
Covert hate speech: white nationalists and dog whistle communication on twitter
Prashanth Bhat and Ofra Klein. 2020 · 2020
Closest in time.
I feel offended, don’t be abusive! implicit/explicit messages in offensive and abusive language
Tommaso Caselli, Valerio Basile, Jelena Mitrović, Inga Kartoziya, and Michael Granitzer. 2020 · 2020
Closest in time.
Empathy is all you need: How a conversational agent should respond to verbal abuse
Hyojin Chin, Lebogang Wame Molefi, and Mun Yong Yi. 2020 · 2020
Closest in time.
Proceedings of the Second Workshop on Trolling, Aggression and Cyberbullying . European Language Resources Association (ELRA), Marseille, France
Ritesh Kumar, Atul Kr. Ojha, Bornini Lahiri, Marcos Zampieri, Shervin Malmasi, Vanessa Murdock, and Daniel Kadar, editors. 2020 · 2020
Closest in time.
Offensive language detection explained
Julian Risch, Robin Ruff, and Ralf Krestel. 2020 · 2020
Closest in time.
Neural text generation with unlikelihood training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020 · 2020
Closest in time.