Fetching the paper…
Reading the bibliography…
Warning: this paper contains content that maybe offensive or upsetting.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2010
Earlier work this paper cites.
Robust training under linguistic adversity
Yitong Li, Trevor Cohn, and Timothy Baldwin. 2017 · 2017
Earlier work this paper cites.
Why we should have seen that coming: Comments on microsoft’s tay “experiment,” and wider implications
M.J. Wolf, K.W. Miller, and F.S. Grodzinsky. 2017 · 2017
Earlier work this paper cites.
Adversarial attacks and defences: A survey
Anirban Chakraborty, Manaar Alam, Vishal Dey, Anupam Chattopadhyay, and Debdeep Mukhopadhyay. 2018 · 2018
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2018 · 2018
Earlier work this paper cites.
Ethical challenges in data-driven dialogue systems
Peter Henderson, Koustuv Sinha, Nicolas Angelard-Gontier, Nan Rosemary Ke, Genevieve Fried, Ryan Lowe, and Joelle Pineau. 2018 · 2018
Earlier work this paper cites.
Adversarial over-sensitivity and over-stability strategies for dialogue models
Tong Niu and Mohit Bansal. 2018 · 2018
Earlier work this paper cites.
Conversations gone awry: Detecting early signs of conversational failure
Justine Zhang, Jonathan Chang, Cristian Danescu-Niculescu-Mizil, Lucas Dixon, Yiqing Hua, Dario Taraborelli, and Nithum Thain. 2018 · 2018
Earlier work this paper cites.
Detecting toxicity triggers in online discussions
Hind Almerekhi, Haewoon Kwak, Bernard J. Jansen, and Joni Salminen. 2019 · 2019
Cited alongside, same era.
Evaluating and enhancing the robustness of dialogue systems: A case study on a negotiation agent
Minhao Cheng, Wei Wei, and Cho-Jui Hsieh. 2019 · 2019
Cited alongside, same era.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Towards Controllable Biases in Language Generation
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2020 · 2020
Later among the works it cites.
Poisoning attacks on algorithmic fairness
D Solans, B Biggio, and C Castillo. 2021 · 2020
Later among the works it cites.
Word-level textual adversarial attacking as combinatorial optimization
Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu, and Maosong Sun. 2020 · 2020
Later among the works it cites.
Just say no: Analyzing the stance of neural dialogue generation in offensive contexts
Ashutosh Baheti, Maarten Sap, Alan Ritter, and Mark Riedl. 2021 · 2021
Later among the works it cites.
Anticipating safety issues in E2E Conversational AI: Framework and tooling
Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2020
Cited alongside, same era.
Reevaluating adversarial examples in natural language
John Morris, Eli Lifland, Jack Lanchantin, Yangfeng Ji, and Yanjun Qi. 2020 · 2020
Cited alongside, same era.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Cited alongside, same era.
Attributing fair decisions with attention interventions
Ninareh Mehrabi, Umang Gupta, Fred Morstatter, Greg Ver Steeg, and Aram Galstyan. 2021a
Cited in the paper.
Exacerbating algorithmic bias through fairness attacks
Ninareh Mehrabi, Muhammad Naveed, Fred Morstatter, and Aram Galstyan. 2021b
Cited in the paper.
Enhancing neural models with vulnerability via adversarial attack
Rong Zhang, Qifei Zhou, Bo An, Weiping Li, Tong Mo, and Bo Wu. 2020a
Cited in the paper.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li. 2020b
Cited in the paper.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Later among the works it cites.
Contrasting human- and machine-generated word-level adversarial examples for text classification
Maximilian Mozes, Max Bartolo, Pontus Stenetorp, Bennett Kleinberg, and Lewis Griffin. 2021 · 2021
Later among the works it cites.
Local explanation of dialogue response generation
Yi-Lin Tuan, Connor Pryor, Wenhu Chen, Lise Getoor, and William Yang Wang. 2021 · 2021
Later among the works it cites.