Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are powerful tools with the potential to benefit society immensely, yet, they have demonstrated biases that perpetuate societal inequalities.
Towards understanding and mitigating social biases in language models
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021 · 2021
Earlier work this paper cites.
Demographic-aware language model fine-tuning as a bias mitigation technique
Aparna Garimella, Rada Mihalcea, and Akhash Amarnath. 2022 · 2022
Earlier work this paper cites.
Debiasing the pre-trained language model through fine-tuning the downstream tasks
Somayeh Ghanbarzadeh, Yan Huang, Hamid Palangi, Radames Cruz Moreno, and Hamed Khanpour. 2022 · 2022
Earlier work this paper cites.
Mitigating gender bias in distilled language models via counterfactual role reversal
Umang Gupta, Jwala Dhamala, Varun Kumar, Apurv Verma, Yada Pruksachatkun, Satyapriya Krishna, Rahul Gupta, Kai-Wei Chang, Greg Ver Steeg, and Aram Galstyan. 2022 · 2022
Earlier work this paper cites.
Przemyslaw Joniak and Akiko Aizawa. 2022 · 2022
Earlier work this paper cites.
How gender debiasing affects internal model representations, and why it matters
Hadas Orgad, Seraphina Goldfarb-Tarrant, and Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Don’t just clean it, proxy clean it: Mitigating bias by proxy in pre-trained models
Swetasudha Panda, Ari Kobren, Michael Wick, and Qinlan Shen. 2022 · 2022
Earlier work this paper cites.
Fairness in machine learning: Detecting and removing gender bias in language models
Carson Sue, Adam Miyauchi, Kunal S Kasodekar, Sai Prathik Mandyala, Priyal Padheriya, and Aesha Shah. 2022 · 2022
Earlier work this paper cites.
A robust bias mitigation procedure based on the stereotype content model
Eddie L Ungless, Amy Rafferty, Hrichika Nag, and Björn Ross. 2022 · 2022
Earlier work this paper cites.
Llm-deliberation: Evaluating llms with interactive multi-agent negotiation games
Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Schönherr, and Mario Fritz. 2023 · 2023
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Earlier work this paper cites.
Emergent cooperation and strategy adaptation in multi-agent systems: An extended coevolutionary theory with llms
I de Zarzà, J de Curtò, Gemma Roig, Pietro Manzoni, and Carlos T Calafate. 2023 · 2023
Earlier work this paper cites.
Breaking the bias: Gender fairness in llms using prompt engineering and in-context learning
Satyam Dwivedi, Sanjukta Ghosh, and Shivam Dwivedi. 2023 · 2023
Earlier work this paper cites.
Bias assessment and mitigation in llm-based code generation
Dong Huang, Qingwen Bu, Jie Zhang, Xiaofei Xie, Junjie Chen, and Heming Cui. 2023 · 2023
Cited alongside, same era.
Smart-llm: Smart multi-agent robot task planning using large language models
Shyam Sundar Kannan, Vishnunandan LN Venkatesh, and Byung-Cheol Min. 2023 · 2023
Cited alongside, same era.
A reinforcement learning approach to mitigating stereotypical biases in language models
Mohammed Rameez Qureshi, Luis Galárraga, and Miguel Couceiro. 2023 · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang. 2023 · 2023
Cited alongside, same era.
Rlrf:reinforcement learning from reflection through debates as feedback for bias mitigation in llms
Llm-based multi-agent systems for software engineering: Vision and the road ahead
Junda He, Christoph Treude, and David Lo. 2024 · 2024
Closest in time.
Jen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang, Wenxuan Wang, Youliang Yuan, Wenxiang Jiao, Xing Wang, Zhaopeng Tu, and Michael R Lyu. 2024 · 2024
Closest in time.
Evaluating gender bias in large language models via chain-of-thought prompting
Masahiro Kaneko, Danushka Bollegala, Naoaki Okazaki, and Timothy Baldwin. 2024 · 2024
Closest in time.
Prometheus 2: An open source language model specialized in evaluating other language models
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ruoxi Cheng, Haoxuan Ma, Shuirong Cao, and Tianyu Shi. 2024 · 2024
Cited alongside, same era.
Few-shot fairness: Unveiling llm’s potential for fairness-aware classification
Garima Chhikara, Anurag Sharma, Kripabandhu Ghosh, and Abhijnan Chakraborty. 2024 · 2024
Cited alongside, same era.
Prompting fairness: Learning prompts for debiasing large language models
Andrei-Victor Chisca, Andrei-Cristian Rad, and Camelia Lemnaru. 2024 · 2024
Cited alongside, same era.
Axolotl: Fairness through assisted self-debiasing of large language model outputs
Sana Ebrahimi, Kaiwen Chen, Abolfazl Asudeh, Gautam Das, and Nick Koudas. 2024 · 2024
Cited alongside, same era.
Cognitive bias in high-stakes decision-making with llms
Jessica Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zexue He. 2024 · 2024
Cited alongside, same era.
Chenhao Fang, Xiaohan Li, Zezhong Fan, Jianpeng Xu, Kaushiki Nag, Evren Korpeoglu, Sushant Kumar, and Kannan Achan. 2024 · 2024
Cited alongside, same era.
Evaluating the efficacy of prompting techniques for debiasing language model outputs (student abstract)
Shaz Furniturewala, Surgan Jandial, Abhinav Java, Simra Shahid, Pragyan Banerjee, Balaji Krishnamurthy, Sumit Bhatia, and Kokil Jaidka. 2024 · 2024
Cited alongside, same era.
Large language model based multi-agents: A survey of progress and challenges
Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang. 2024 · 2024
Cited alongside, same era.
Zhongkun Liu, Zheng Chen, Mengqi Zhang, Zhaochun Ren, Zhumin Chen, and Pengjie Ren. 2024 · 2024
Closest in time.
Fairness-guided few-shot prompting for large language models
Huan Ma, Changqing Zhang, Yatao Bian, Lemao Liu, Zhirui Zhang, Peilin Zhao, Shu Zhang, Huazhu Fu, Qinghua Hu, and Bingzhe Wu. 2024 · 2024
Closest in time.
Llm-guided counterfactual data generation for fairer ai
Ashish Mishra, Gyanaranjan Nayak, Suparna Bhattacharya, Tarun Kumar, Arpit Shah, and Martin Foltin. 2024 · 2024
Closest in time.
Agentcoord: Visually exploring coordination strategy for llm-based multi-agent collaboration
Bo Pan, Jiaying Lu, Ke Wang, Li Zheng, Zhen Wen, Yingchaojie Feng, Minfeng Zhu, and Wei Chen. 2024 · 2024
Closest in time.
Navigating complexity: Orchestrated problem solving with multi-agent llms
Sumedh Rasal and EJ Hauer. 2024 · 2024
Closest in time.
Simulating human strategic behavior: Comparing single and multi-agent llms
Karthik Sreedhar and Lydia Chilton. 2024 · 2024
Closest in time.
Llm-based multi-agent reinforcement learning: Current and future directions
Chuanneng Sun, Songjun Huang, and Dario Pompili. 2024 · 2024
Closest in time.
Autodefense: Multi-agent llm defense against jailbreak attacks
Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qingyun Wu. 2024 · 2024
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2086
Closest in time.