Fetching the paper…
Reading the bibliography…
Large pretrained language models can easily produce toxic or biased content, which is prohibitive for practical use.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019 · 1909
Earlier work this paper cites.
Say what i want: Towards the dark side of neural dialogue models
Haochen Liu, Tyler Derr, Zitao Liu, and Jiliang Tang. 2019 · 1909
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M Smith, et al. 2020 · 2004
Earlier work this paper cites.
Chat as expected: Learning to manipulate black-box neural dialogue models
Haochen Liu, Zhiwei Wang, Tyler Derr, and Jiliang Tang. 2020 · 2005
Earlier work this paper cites.
The radicalization risks of gpt-3 and advanced neural language models
Kris McGuffie and Alex Newhouse. 2020 · 2009
Earlier work this paper cites.
Recipes for safety in open-domain chatbots
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2010
Earlier work this paper cites.
Hatecheck: Functional tests for hate speech detection models
Paul Röttger, Bertram Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet B Pierrehumbert. 2020 · 2012
Earlier work this paper cites.
Learning from tay’s introduction
Peter Lee. 2016 · 2016
Earlier work this paper cites.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, and Piotr Bojanowski Tomas Mikolov. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2018 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Detecting egregious responses in neural sequence-to-sequence models
Tianxing He and James Glass. 2018 · 2018
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Cited alongside, same era.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019 · 2019
Cited alongside, same era.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2019 · 2019
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Later among the works it cites.
Dexperts: Decoding-time controlled text generation with experts and anti-experts
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A Smith, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Later among the works it cites.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
“nice try, kiddo”: Investigating ad hominems in dialogue responses
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2021 · 2021
Later among the works it cites.
Automatically exposing problems with neural dialog models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2020
Cited alongside, same era.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Towards controllable biases in language generation
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2020 · 2020
Cited alongside, same era.
Dialogpt: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B Dolan. 2020 · 2020
Cited alongside, same era.
Assessing political prudence of open-domain chatbots
Yejin Bang, Nayeon Lee, Etsuko Ishii, Andrea Madotto, and Pascale Fung. 2021 · 2021
Cited alongside, same era.
Plato-2: Towards building an open-domain chatbot via curriculum learning
Siqi Bao, Huang He, Fan Wang, Hua Wu, Haifeng Wang, Wenquan Wu, Zhen Guo, Zhibin Liu, and Xinchao Xu. 2021 · 2021
Cited alongside, same era.
Dian Yu and Kenji Sagae. 2021 · 2021
Later among the works it cites.
Exploring prompt-based few-shot learning for grounded dialog generation
Chujie Zheng and Minlie Huang. 2021 · 2021
Later among the works it cites.
Eva: An open-domain chinese dialogue system with large-scale generative pre-training
Hao Zhou, Pei Ke, Zheng Zhang, Yuxian Gu, Yinhe Zheng, Chujie Zheng, Yida Wang, Chen Henry Wu, Hao Sun, Xiaocong Yang, Bosi Wen, Xiaoyan Zhu, Minlie Huang, and Jie Tang. 2021 · 2021
Later among the works it cites.
Cold: A benchmark for chinese offensive language detection
Jiawen Deng, Jingyan Zhou, Hao Sun, Fei Mi, and Minlie Huang. 2022 · 2022
Closest in time.
Eva2.0: Investigating open-domain chinese dialogue systems with large-scale pre-training
Yuxian Gu, Jiaxin Wen, Hao Sun, Yi Song, Pei Ke, Chujie Zheng, Zheng Zhang, Jianzhu Yao, Xiaoyan Zhu, Jie Tang, and Minlie Huang. 2022 · 2022
Closest in time.
Pangu-bot: Efficient generative dialogue pre-training from pre-trained language model
Fei Mi, Yitong Li, Yulong Zeng, Jingyan Zhou, Yasheng Wang, Chuanfei Xu, Lifeng Shang, Xin Jiang, Shiqi Zhao, and Qun Liu. 2022 · 2022
Closest in time.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Closest in time.
On the safety of conversational models: Taxonomy, dataset, and benchmark
Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang. 2022 · 2022
Closest in time.
Towards identifying social bias in dialog systems: Frame, datasets, and benchmarks
Jingyan Zhou, Jiawen Deng, Fei Mi, Yitong Li, Yasheng Wang, Minlie Huang, Xin Jiang, Qun Liu, and Helen Meng. 2022 · 2022
Closest in time.
Can you put it all together: Evaluating conversational agents’ ability to blend skills
Eric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston, and Y-Lan Boureau. 2020 · 2030
Closest in time.