Improving alignment of dialogue agents via targeted human judgements
Original
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al. 2022 · 2022
Later among the works it cites.
Eva2.0: Investigating open-domain chinese dialogue systems with large-scale pre-training
Yuxian Gu, Jiaxin Wen, Hao Sun, Yi Song, Pei Ke, Chujie Zheng, Zheng Zhang, Jianzhu Yao, Xiaoyan Zhu, Jie Tang, and Minlie Huang. 2022 · 2022
Later among the works it cites.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
Later among the works it cites.
ProsocialDialog: A prosocial backbone for conversational agents
Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, and Maarten Sap. 2022 · 2022
Later among the works it cites.
SafeText: A benchmark for exploring physical safety in language models
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, and William Yang Wang. 2022 · 2022
Later among the works it cites.
You don’t know my favorite color: Preventing dialogue representations from revealing speakers’ private personas
Haoran Li, Yangqiu Song, and Lixin Fan. 2022a · 2022
Later among the works it cites.
ParaDetox: Detoxification with parallel data
Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022 · 2022
Later among the works it cites.
Robust conversational agents against imperceptible toxicity triggers
Ninareh Mehrabi, Ahmad Beirami, Fred Morstatter, and Aram Galstyan. 2022 · 2022
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue
OpenAI. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Original
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Later among the works it cites.
Ignore previous prompt: Attack techniques for language models
Original
Fábio Perez and Ian Ribeiro. 2022 · 2022
Later among the works it cites.
Why so toxic? measuring and triggering toxic behavior in open-domain chatbots
Wai Man Si, Michael Backes, Jeremy Blackburn, Emiliano De Cristofaro, Gianluca Stringhini, Savvas Zannettou, and Yang Zhang. 2022 · 2022
Later among the works it cites.
On the safety of conversational models: Taxonomy, dataset, and benchmark
Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang. 2022 · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. 2022 · 2022
Later among the works it cites.
Toxicity detection with generative prompt-based inference
Yau-Shian Wang and Yingshan Chang. 2022 · 2022
Later among the works it cites.
Leashing the inner demons: Self-detoxification for language models
Canwen Xu, Zexue He, Zhankui He, and Julian McAuley. 2022 · 2022
Later among the works it cites.
Constructing highly inductive contexts for dialogue safety through controllable reverse generation
Zhexin Zhang, Jiale Cheng, Hao Sun, Jiawen Deng, Fei Mi, Yasheng Wang, Lifeng Shang, and Minlie Huang. 2022 · 2022
Later among the works it cites.
The moral integrity corpus: A benchmark for ethical dialogue systems
Caleb Ziems, Jane Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022 · 2022
Later among the works it cites.
Enhancing offensive language detection with data augmentation and knowledge distillation
Jiawen Deng, Zhuang Chen, Hao Sun, Zhexin Zhang, Jincenzi Wu, Satoshi Nakagawa, Fuji Ren, and Minlie Huang. 2023 · 2023
Closest in time.
Toxicity in chatgpt: Analyzing persona-assigned language models
Original
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023 · 2023
Closest in time.
From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023 · 2023
Closest in time.
Critic: Large language models can self-correct with tool-interactive critiquing
Original
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023 · 2023
Closest in time.
An overview of catastrophic ai risks
Original
Dan Hendrycks, Mantas Mazeika, and Thomas Woodside. 2023 · 2023
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023 · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Original
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023 · 2023
Closest in time.
Interpretability dreams
Chris Olah. 2023 · 2023
Closest in time.
Safety assessment of chinese large language models
Original
Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Original
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Shepherd: A critic for language model generation
Original
Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan, Sean O’Brien, Ramakanth Pasunuru, Jane Dwivedi-Yu, Olga Golovneva, Luke Zettlemoyer, Maryam Fazel-Zarandi, and Asli Celikyilmaz. 2023 · 2023
Closest in time.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Original
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2023 · 2023
Closest in time.
Chbias: Bias evaluation and mitigation of chinese conversational language models
Original
Jiaxu Zhao, Meng Fang, Zijing Shi, Yitong Li, Ling Chen, and Mykola Pechenizkiy. 2023 · 2023
Closest in time.
Bbq: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2086
Closest in time.