Fetching the paper…
Reading the bibliography…
Understanding what constitutes safe text is an important issue in natural language processing and can often prevent the deployment of models deemed harmful and unsafe.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
The ontogeny of common sense
Lynd Forguson and Alison Gopnik. 1988 · 1988
Earlier work this paper cites.
Using common sense to recognize cultural differences
Junia Anacleto, Henry Lieberman, Aparecido de Carvalho, Vânia Néris, Muriel Godoi, Marie Tsutsumi, Jose Espinosa, Américo Talarico, and Silvia Zem-Mascarenhas. 2006 · 2006
Earlier work this paper cites.
“got you!”: Automatic vandalism detection in Wikipedia with web-based shallow syntactic-semantic modeling
William Yang Wang and Kathleen McKeown. 2010 · 2010
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Commonsense knowledge base completion
Xiang Li, Aynaz Taheri, Lifu Tu, and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Internet use, risks and online behaviour: The view of internet users with intellectual disabilities and their caregivers
Esther Chiner, Marcos Gómez-Puerta, and María Cristina Cardona-Moltó. 2017 · 2017
Earlier work this paper cites.
Patient and consumer safety risks when using conversational assistants for medical information: an observational study of siri, alexa, and google assistant
Timothy W Bickmore, Ha Trinh, Stefan Olafsson, Teresa K O’Leary, Reza Asadi, Nathaniel M Rickles, and Ricardo Cruz. 2018 · 2018
Earlier work this paper cites.
Detecting gang-involved escalation on social media using context
Serina Chang, Ruiqi Zhong, Ethan Adams, Fei-Tzin Lee, Siddharth Varia, Desmond Patton, William Frey, Chris Kedzie, and Kathy McKeown. 2018 · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Comet-atomic 2020: On symbolic and neural commonsense knowledge graphs
Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi. 2021 · 2020
Cited alongside, same era.
Social work thinking for ux and ai design
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021 · 2021
Later among the works it cites.
On hallucination and predictive uncertainty in conditional language generation
Yijun Xiao and William Yang Wang. 2021 · 2021
Later among the works it cites.
Ethical-advice taker: Do language models understand natural language interventions?
Jieyu Zhao, Daniel Khashabi, Tushar Khot, Ashish Sabharwal, and Kai-Wei Chang. 2021 · 2021
Later among the works it cites.
SafetyKit: First aid for measuring safety in open-domain conversational systems
Emily Dinan, Gavin Abercrombie, A. Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2022 · 2022
Closest in time.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Desmond Upton Patton. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Latent hatred: A benchmark for understanding implicit hate speech
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021 · 2021
Cited alongside, same era.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Cited alongside, same era.
Delphi: Towards machine ethics and norms
Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Maxwell Forbes, Jon Borchardt, Jenny Liang, Oren Etzioni, Maarten Sap, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022 · 2022
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022 · 2022
Closest in time.
Skillbot: Identifying risky content for children in alexa skills
Tu Le, Danny Yuxing Huang, Noah Apthorpe, and Yuan Tian. 2022 · 2022
Closest in time.
Mitigating covertly unsafe text within natural language systems
Alex Mei, Anisha Kabir, Sharon Levy, Melanie Subbiah, Emily Allaway, John Judge, Desmond Patton, Bruce Bimber, Kathleen McKeown, and William Yang Wang. 2022 · 2022
Closest in time.
‘beach’ to ‘bitch’: Inadvertent unsafe transcription of kids’ content on youtube
Krithika Ramesh, Ashiqur R. KhudaBukhsh, and Sumeet Kumar. 2022 · 2022
Closest in time.
On the safety of conversational models: Taxonomy, dataset, and benchmark
Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang. 2022 · 2022
Closest in time.
ANLIzing the adversarial natural language inference dataset
Adina Williams, Tristan Thrush, and Douwe Kiela. 2022 · 2022
Closest in time.
The moral integrity corpus: A benchmark for ethical dialogue systems
Caleb Ziems, Jane Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022 · 2022
Closest in time.