Fetching the paper…
Reading the bibliography…
The common toxicity and societal bias in contents generated by large language models (LLMs) necessitate strategies to reduce harm.
The philosophy of moral development: Moral stages and the idea of justice , volume 1
Lawrence Kohlberg. 1921 · 1921
Earlier work this paper cites.
The moral judgment of the child
Jean Piaget. 1934 · 1934
Earlier work this paper cites.
On a test of whether one of two random variables is stochastically larger than the other
Henry B Mann and Donald R Whitney. 1947 · 1947
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
Perspective taking: Imagining how another feels versus imaging how you would feel
C Daniel Batson, Shannon Early, and Giovanni Salvarani. 1997 · 1997
Earlier work this paper cites.
Emotional development and emotional intelligence: Educational implications
Peter Ed Salovey and David J Sluyter. 1997 · 1997
Earlier work this paper cites.
Working with emotional intelligence
Daniel Goleman. 1998 · 1998
Earlier work this paper cites.
Perspective-taking: decreasing stereotype expression, stereotype accessibility, and in-group favoritism
Adam D Galinsky and Gordon B Moskowitz. 2000 · 2000
Earlier work this paper cites.
Perspective taking and prejudice reduction: The mediational role of empathy arousal and situational attributions
Theresa K Vescio, Gretchen B Sechrist, and Matthew P Paolucci. 2003 · 2003
Earlier work this paper cites.
The bar-on model of emotional-social intelligence (esi) 1
Reuven Bar-On. 2006 · 2006
Earlier work this paper cites.
The neural substrate of human empathy: effects of perspective-taking and cognitive appraisal
Claus Lamm, C Daniel Batson, and Jean Decety. 2007 · 2007
Earlier work this paper cites.
Perspective taking combats automatic expressions of racial bias
Andrew R Todd, Galen V Bodenhausen, Jennifer A Richeson, and Adam D Galinsky. 2011 · 2011
Earlier work this paper cites.
Two forms of perspective taking: Imagining how another feels and imagining how you would feel
C Daniel Batson. 2012 · 2012
Earlier work this paper cites.
Vader: A parsimonious rule-based model for sentiment analysis of social media text
Clayton Hutto and Eric Gilbert. 2014 · 2014
Earlier work this paper cites.
Perspective-taking as a strategy for improving intergroup relations: Evidence, mechanisms, and qualifications
Andrew R Todd and Adam D Galinsky. 2014 · 2014
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Empathy: A social psychological approach
Mark H Davis. 2018 · 2018
Earlier work this paper cites.
Content analysis: An introduction to its methodology
Klaus Krippendorff. 2018 · 2018
Earlier work this paper cites.
Identifying and reducing gender bias in word-level language models
Shikha Bordia and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
Reducing gender bias in word-level language models with a gender-equalizing loss function
Yusu Qian, Urwa Muaz, Ben Zhang, and Jae Won Hyun. 2019 · 2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019 · 2019
Earlier work this paper cites.
Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology
Ran Zmigrod, Sabrina J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel. 2020 · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Reducing sentiment bias in language models via counterfactual evaluation
Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, and Pushmeet Kohli. 2020 · 2020
Earlier work this paper cites.
Does gender matter? towards fairness in dialogue systems
Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang. 2020 · 2020
Earlier work this paper cites.
Gender bias in neural natural language processing
Kaiji Lu, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta. 2020 · 2020
Earlier work this paper cites.
Reducing non-normative text generation from language models
Xiangyu Peng, Siyan Li, Spencer Frazier, and Mark Riedl. 2020 · 2020
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Earlier work this paper cites.
Reducing gender bias in neural machine translation as a domain adaptation problem
Danielle Saunders and Bill Byrne. 2020 · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
Text detoxification using large pre-trained neural models
David Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva, Olga Kozlova, Nikita Semenov, and Alexander Panchenko. 2021 · 2021
Cited alongside, same era.
Bold: Dataset and metrics for measuring biases in open-ended language generation
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021 · 2021
Cited alongside, same era.
Gedi: Generative discriminator guided sequence generation
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2021 · 2021
Cited alongside, same era.
Towards understanding and mitigating social biases in language models
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021 · 2021
Cited alongside, same era.
Queer people are people first: Deconstructing sexual identity stereotypes in large language models
Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. 2023 · 2023
Later among the works it cites.
A survey for in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2023 · 2023
Later among the works it cites.
The capacity for moral self-correction in large language models
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al. 2023 · 2023
Later among the works it cites.
Gender-tuning: Empowering fine-tuning for debiasing pre-trained language models
Somayeh Ghanbarzadeh, Yan Huang, Hamid Palangi, Radames Cruz Moreno, and Hamed Khanpour. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DExperts: Decoding-time controlled text generation with experts and anti-experts
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021 · 2021
Cited alongside, same era.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
“nice try, kiddo”: Investigating ad hominems in dialogue responses
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2021a · 2021
Cited alongside, same era.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021 · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021 · 2021
Cited alongside, same era.
Challenges in detoxifying language models
Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021 · 2021
Cited alongside, same era.
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023 · 2023
Later among the works it cites.
Large language models can self-improve
Jiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2023a · 2023
Later among the works it cites.
Chatgpt for good? on opportunities and challenges of large language models for education
Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, et al. 2023 · 2023
Later among the works it cites.
Biastestgpt: Using chatgpt for social bias testing of language models
Rafal Kocielnik, Shrimai Prabhumoye, Vivian Zhang, Roy Jiang, R. Michael Alvarez, and Anima Anandkumar. 2023 · 2023
Later among the works it cites.
On the intersection of self-correction and trust in language models
Satyapriya Krishna. 2023 · 2023
Later among the works it cites.
Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models
Tiffany H Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepaño, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, et al. 2023 · 2023
Later among the works it cites.
Self-detoxifying language models via toxification reversal
Chak Leong, Yi Cheng, Jiashuo Wang, Jian Wang, and Wenjie Li. 2023 · 2023
Later among the works it cites.
Large language models understand and can be enhanced by emotional stimuli
Cheng Li, Jindong Wang, Yixuan Zhang, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Fang Luo, Qiang Yang, and Xing Xie. 2023 · 2023
Later among the works it cites.
Yucheng Li. 2023 · 2023
Later among the works it cites.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. 2023 · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023 · 2023
Later among the works it cites.
Selfcheck: Using llms to zero-shot check their own step-by-step reasoning
Ning Miao, Yee Whye Teh, and Tom Rainforth. 2023 · 2023
Later among the works it cites.
Chatgpt: A large-scale generative model for open-domain chat
OpenAI. 2023 · 2023
Later among the works it cites.
OpenAI et al. 2023 · 2023
Later among the works it cites.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang. 2023 · 2023
Later among the works it cites.
On the challenges of using black-box apis for toxicity evaluation in research
Luiza Amador Pozzobon, Beyza Ermis, Patrick Lewis, and Sara Hooker. 2023 · 2023
Later among the works it cites.
Large language model alignment: A survey
Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023 · 2023
Later among the works it cites.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Noah Shinn, Beck Labash, and Ashwin Gopinath. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Effective long-context scaling of foundation models
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al. 2023 · 2023
Later among the works it cites.
Adept: A debiasing prompt framework
Ke Yang, Charles Yu, Yi R Fung, Manling Li, and Heng Ji. 2023 · 2023
Later among the works it cites.
A survey of controllable text generation using transformer-based pre-trained language models
Hanqing Zhang, Haolin Song, Shaoyu Li, Ming Zhou, and Dawei Song. 2023 · 2023
Later among the works it cites.
Mil-decoding: Detoxifying language models at token-level via multiple instance learning
Xu Zhang and Xiaojun Wan. 2023 · 2023
Later among the works it cites.
Exploring ai ethics of chatgpt: A diagnostic analysis
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing. 2023 · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023 · 2024
Closest in time.
Self-debiasing large language models: Zero-shot recognition and reduction of stereotypes
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Tong Yu, Hanieh Deilamsalehy, Ruiyi Zhang, Sungchul Kim, and Franck Dernoncourt. 2024 · 2024
Closest in time.
You only prompt once: On the capabilities of prompt learning on large language models to tackle toxic content
Xinlei He, Savvas Zannettou, Yun Shen, and Yang Zhang. 2023 · 2024
Closest in time.
Empathy and the right to be an exception: What llms can and cannot do
William Kidder, Jason D’Cruz, and Kush R Varshney. 2024 · 2024
Closest in time.
Zaijing Li, Gongwei Chen, Rui Shao, Dongmei Jiang, and Liqiang Nie. 2024 · 2024
Closest in time.
Parameter-efficient detoxification with contrastive decoding
Tong Niu, Caiming Xiong, Semih Yavuz, and Yingbo Zhou. 2024 · 2024
Closest in time.