Fetching the paper…
Reading the bibliography…
The widespread adoption of large language models (LLMs) underscores the urgent need to ensure their fairness.
Mass. Gen. Laws ch. 234A, 2016
Massachusetts general laws · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, Paul F, Leike, Jan, Brown, Tom, Martic, Miljan, Legg, Shane, and Amodei, Dario · 2017
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, Tom B., Mann, Benjamin, Ryder, Nick, Subbiah, Melanie, Kaplan, Jared, Dhariwal, Prafulla, Neelakantan, Arvind, Shyam, Pranav, Sastry, Girish, Askell, Amanda, Agarwal, Sandhini, Herbert-Voss, Ariel, Krueger, Gretchen, Henighan, Tom, Child, Rewon, Ramesh, Aditya, Ziegler, Daniel M., Wu, Jeffrey, Winter, Clemens, Hesse, Christopher, Chen, Mark, Sigler, Eric, Litwin, Mateusz, Gray, Scott, Chess, Benjamin, Clark, Jack, Berner, Christopher, McCandlish, Sam, Radford, Alec, Sutskever, Ilya, and Amodei, Dario · 2020
Earlier work this paper cites.
Unqovering stereotyping biases via underspecified questions
Li, Tao, Khot, Tushar, Khashabi, Daniel, Sabharwal, Ashish, and Srikumar, Vivek · 2020
Earlier work this paper cites.
Persistent anti-muslim bias in large language models, 2021
Abid, Abubakar, Farooqi, Maheen, and Zou, James · 2021
Earlier work this paper cites.
Collecting a large-scale gender bias dataset for coreference resolution and machine translation
Levy, Shahar, Lazar, Koren, and Stanovsky, Gabriel · 2021
Earlier work this paper cites.
Bbq: A hand-built bias benchmark for question answering
Parrish, Alicia, Chen, Angelica, Nangia, Nikita, Padmakumar, Vishakh, Phang, Jason, Thompson, Jana, Htut, Phu Mon, and Bowman, Samuel R · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Yuntao, Jones, Andy, Ndousse, Kamal, Askell, Amanda, Chen, Anna, DasSarma, Nova, Drain, Dawn, Fort, Stanislav, Ganguli, Deep, Henighan, Tom, et al · 2022
Earlier work this paper cites.
Measuring harmful sentence completion in language models for lgbtqia+ individuals
Nozza, Debora, Bianchi, Federico, Lauscher, Anne, Hovy, Dirk, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, Long, Wu, Jeffrey, Jiang, Xu, Almeida, Diogo, Wainwright, Carroll, Mishkin, Pamela, Zhang, Chong, Agarwal, Sandhini, Slama, Katarina, Ray, Alex, et al · 2022
Earlier work this paper cites.
Perturbation augmentation for fairer nlp
Qian, Rebecca, Ross, Candace, Fernandes, Jude, Smith, Eric, Kiela, Douwe, and Williams, Adina · 2022
Earlier work this paper cites.
Bloom: A 176b-parameter open-access multilingual language model
Scao, Teven Le, Fan, Angela, Akiki, Christopher, Pavlick, Ellie, Ilić, Suzana, Hesslow, Daniel, Castagné, Roman, Luccioni, Alexandra Sasha, Yvon, François, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, Jason, Wang, Xuezhi, Schuurmans, Dale, Bosma, Maarten, Xia, Fei, Chi, Ed, Le, Quoc V, Zhou, Denny, et al · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Bubeck, Sébastien, Chandrasekaran, Varun, Eldan, Ronen, Gehrke, Johannes, Horvitz, Eric, Kamar, Ece, Lee, Peter, Lee, Yin Tat, Li, Yuanzhi, Lundberg, Scott, Nori, Harsha, Palangi, Hamid, Ribeiro, Marco Tulio, and Zhang, Yi · 2023
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chan, Chi-Min, Chen, Weize, Su, Yusheng, Yu, Jianxuan, Xue, Wei, Zhang, Shanghang, Fu, Jie, and Liu, Zhiyuan · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, Aakanksha, Narang, Sharan, Devlin, Jacob, Bosma, Maarten, Mishra, Gaurav, Roberts, Adam, Barham, Paul, Chung, Hyung Won, Sutton, Charles, Gehrmann, Sebastian, et al · 2023
Cited alongside, same era.
Fft: Towards harmlessness evaluation and analysis for llms with factuality, fairness, toxicity
Li, Yunqi and Zhang, Yongfeng · 2023
Later among the works it cites.
Encouraging divergent thinking in large language models through multi-agent debate, 2023
Liang, Tian, He, Zhiwei, Jiao, Wenxiang, Wang, Xing, Wang, Yan, Wang, Rui, Yang, Yujiu, Tu, Zhaopeng, and Shi, Shuming · 2023
Later among the works it cites.
Training socially aligned language models on simulated social interactions, 2023
Liu, Ruibo, Yang, Ruixin, Jia, Chenyan, Zhang, Ge, Zhou, Denny, Dai, Andrew M., Yang, Diyi, and Vosoughi, Soroush · 2023
Later among the works it cites.
Do llms possess a personality? making the mbti test an amazing evaluation for large language models, 2023
Pan, Keyu and Zeng, Yawen · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, Rafael, Sharma, Archit, Mitchell, Eric, Ermon, Stefano, Manning, Christopher D, and Finn, Chelsea · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cui, Shiyao, Zhang, Zhenyu, Chen, Yilong, Zhang, Wenyuan, Liu, Tianyun, Wang, Siqi, and Liu, Tingwen · 2023
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Du, Yilun, Li, Shuang, Torralba, Antonio, Tenenbaum, Joshua B, and Mordatch, Igor · 2023
Cited alongside, same era.
Improving language model negotiation with self-play and in-context learning from ai feedback, 2023
Fu, Yao, Peng, Hao, Khot, Tushar, and Lapata, Mirella · 2023
Cited alongside, same era.
Bias and fairness in large language models: A survey
Gallegos, Isabel O, Rossi, Ryan A, Barrow, Joe, Tanjim, Md Mehrab, Kim, Sungchul, Dernoncourt, Franck, Yu, Tong, Zhang, Ruiyi, and Ahmed, Nesreen K · 2023
Cited alongside, same era.
Trustgpt: A benchmark for trustworthy and responsible large language models
Huang, Yue, Zhang, Qihui, Sun, Lichao, et al · 2023
Cited alongside, same era.
Top 70 controversial debate topics for critical thinkers in 2023, 2023
Jane, Ng · 2023
Cited alongside, same era.
225 trending social issues topics for academic writing, 2023
Jay, Cooper · 2023
Cited alongside, same era.
Mistral 7b, 2023
Jiang, Albert Q., Sablayrolles, Alexandre, Mensch, Arthur, Bamford, Chris, Chaplot, Devendra Singh, de las Casas, Diego, Bressand, Florian, Lengyel, Gianna, Lample, Guillaume, Saulnier, Lucile, Lavaud, Lélio Renard, Lachaux, Marie-Anne, Stock, Pierre, Scao, Teven Le, Lavril, Thibaut, Wang, Thomas, Lacroix, Timothée, and Sayed, William El · 2023
Cited alongside, same era.
Later among the works it cites.
In-context impersonation reveals large language models’ strengths and biases
Salewski, Leonard, Alaniz, Stephan, Rio-Torto, Isabel, Schulz, Eric, and Akata, Zeynep · 2023
Later among the works it cites.
Multi-agent collaboration: Harnessing the power of intelligent llm agents, 2023
Talebirad, Yashar and Nadiri, Amirhossein · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, Hugo, Martin, Louis, Stone, Kevin, Albert, Peter, Almahairi, Amjad, Babaei, Yasmine, Bashlykov, Nikolay, Batra, Soumya, Bhargava, Prajjwal, Bhosale, Shruti, et al · 2023
Later among the works it cites.
Biasasker: Measuring the bias in conversational ai system
Wan, Yuxuan, Wang, Wenxuan, He, Pinjia, Gu, Jiazhen, Bai, Haonan, and Lyu, Michael R · 2023
Later among the works it cites.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
Zhuo, Terry Yue, Huang, Yujin, Chen, Chunyang, and Xing, Zhenchang · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, Andy, Wang, Zifan, Kolter, J Zico, and Fredrikson, Matt · 2023
Later among the works it cites.
Llm harmony: Multi-agent communication for problem solving, 2024
Rasal, Sumedh · 2024
Closest in time.