Fetching the paper…
Reading the bibliography…
Bias in LLMs can harm user experience and societal outcomes.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Gender bias in contextualized word embeddings
Zhao, J.; Wang, T.; Yatskar, M.; Cotterell, R.; Ordonez, V.; and Chang, K.-W. 2019 · 1904
Earlier work this paper cites.
Mitigating gender bias in natural language processing: Literature review
Sun, T.; Gaut, A.; Tang, S.; Huang, Y.; ElSherief, M.; Zhao, J.; Mirza, D.; Belding, E.; Chang, K.-W.; and Wang, W. Y. 2019 · 1906
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M.; Stiennon, N.; Wu, J.; Brown, T. B.; Radford, A.; Amodei, D.; Christiano, P.; and Irving, G. 2019 · 1909
Earlier work this paper cites.
Language (technology) is power: A critical survey of" bias" in nlp
Blodgett, S. L.; Barocas, S.; Daumé III, H.; and Wallach, H. 2020 · 2005
Earlier work this paper cites.
Reinforcement learning from simultaneous human and MDP reward
Knox, W. B.; and Stone, P. 2012 · 2012
Earlier work this paper cites.
Guided Policy Search
Levine, S.; and Koltun, V. 2013 · 2013
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T.; Chang, K.-W.; Zou, J. Y.; Saligrama, V.; and Kalai, A. T. 2016 · 2016
Earlier work this paper cites.
Learning to Learn from Weak Supervision by Full Supervision
Dehghani, M.; Severyn, A.; Rothe, S.; and Kamps, J. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
Legal summarization for multi-role debate dialogue via controversy focus mining and multi-task learning
Duan, X.; Zhang, Y.; Yuan, L.; Zhou, X.; Liu, X.; Wang, T.; Wang, R.; Zhang, Q.; Sun, C.; and Wu, F. 2019 · 2019
Earlier work this paper cites.
Towards the systematic reporting of the energy and carbon footprints of machine learning
Henderson, P.; Hu, J.; Romoff, J.; Brunskill, E.; Jurafsky, D.; and Pineau, J. 2020 · 2020
Earlier work this paper cites.
Learning to summarize from human feedback
Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; and Christiano, P. 2020 · 2020
Earlier work this paper cites.
Group multi-role assignment with conflicting roles and agents
Zhu, H. 2020 · 2020
Earlier work this paper cites.
A General Language Assistant as a Laboratory for Alignment
Amanda, A.; Bai, Y.; Chen, A.; Drain, D.; Deep, G.; Tom, H.; Jones, A.; Joseph, N.; Mann, B.; DasSarma, N.; Nelson, E.; Zac, H.-D.; Hernandez, D.; Kernion, J.; Kamal, N.; Olsson, C.; Amodei, D.; Brown, T.; Clark, J.; Sam, M.; Olah, C.; and Kaplan, J. 2021 · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S. 2021 · 2021
Earlier work this paper cites.
Large image datasets: A pyrrhic win for computer vision?
Birhane, A.; and Prabhu, V. U. 2021 · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. 2022 · 2022
Cited alongside, same era.
Large language models can self-improve
Huang, J.; Gu, S. S.; Hou, L.; Wu, Y.; Wang, X.; Yu, H.; and Han, J. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
Fairness reprogramming
Zhang, G.; Zhang, Y.; Zhang, Y.; Fan, W.; Li, Q.; Liu, S.; and Chang, S. 2022 · 2022
Cited alongside, same era.
Evaluating and mitigating discrimination in language model decisions
Tamkin, A.; Askell, A.; Lovitt, L.; Durmus, E.; Joseph, N.; Kravec, S.; Nguyen, K.; Kaplan, J.; and Ganguli, D. 2023 · 2023
Later among the works it cites.
Can Large Language Models Really Improve by Self-critiquing Their Own Plans?
Valmeekam, K.; Marquez, M.; and Kambhampati, S. 2023 · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Xi, Z.; Chen, W.; Guo, X.; He, W.; Ding, Y.; Hong, B.; Zhang, M.; Wang, J.; Jin, S.; Zhou, E.; et al. 2023 · 2023
Later among the works it cites.
Exploring collaboration mechanisms for llm agents: A social psychology view
Zhang, J.; Xu, X.; and Deng, S. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zong, C.; Yan, Y.; Lu, W.; Huang, E.; Shao, J.; and Zhuang, Y. 2024 · 2022
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chan, C.-M.; Chen, W.; Su, Y.; Yu, J.; Xue, W.; Zhang, S.; Fu, J.; and Liu, Z. 2023 · 2023
Cited alongside, same era.
Improving language model negotiation with self-play and in-context learning from ai feedback
Fu, Y.; Peng, H.; Khot, T.; and Lapata, M. 2023 · 2023
Cited alongside, same era.
The capacity for moral self-correction in large language models
Ganguli, D.; Askell, A.; Schiefer, N.; Liao, T. I.; Lukošiūtė, K.; Chen, A.; Goldie, A.; Mirhoseini, A.; Olsson, C.; Hernandez, D.; et al. 2023 · 2023
Cited alongside, same era.
Händler, T. 2023 · 2023
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Huang, J.; Chen, X.; Mishra, S.; Zheng, H. S.; Yu, A. W.; Song, X.; and Zhou, D. 2023 · 2023
Cited alongside, same era.
Towards mitigating LLM hallucination via self reflection
Ji, Z.; Yu, T.; Xu, Y.; Lee, N.; Ishii, E.; and Fung, P. 2023 · 2023
Cited alongside, same era.
Zhou, Z.; Song, J.; Yao, K.; Shu, Z.; and Ma, L. 2023 · 2023
Later among the works it cites.
LLM Multi-Agent Systems: Challenges and Open Problems
Han, S.; Zhang, Q.; Yao, Y.; Jin, W.; Xu, Z.; and He, C. 2024 · 2024
Closest in time.
Facilitating Holistic Evaluations with LLMs: Insights from Scenario-Based Experiments
Ishida, T. 2024 · 2024
Closest in time.
Kim, K.; Lee, S.; Huang, K.-H.; Chan, H. P.; Li, M.; and Ji, H. 2024 · 2024
Closest in time.
Lee, K.; Hwang, D.; Park, S.; Jang, Y.; and Lee, M. 2024 · 2024
Closest in time.
Debatrix: Multi-dimensinal Debate Judge with Iterative Chronological Analysis Based on LLM
Liang, J.; Ye, R.; Han, M.; Lai, R.; Zhang, X.; Huang, X.; and Wei, Z. 2024 · 2024
Closest in time.
Lu, L.-C.; Chen, S.-J.; Pai, T.-M.; Yu, C.-H.; Lee, H.-y.; and Sun, S.-H. 2024 · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. 2024 · 2024
Closest in time.
Self-alignment of large language models via multi-agent social simulation
Pang, X.; Tang, S.; Ye, R.; Xiong, Y.; Zhang, B.; Wang, Y.; and Chen, S. 2024 · 2024
Closest in time.
A Critical Evaluation of AI Feedback for Aligning Large Language Models
Sharma, A.; Keh, S.; Mitchell, E.; Finn, C.; Arora, K.; and Kollar, T. 2024 · 2024
Closest in time.
Self-contrast: Better reflection through inconsistent solving perspectives
Zhang, W.; Shen, Y.; Wu, L.; Peng, Q.; Wang, J.; Zhuang, Y.; and Lu, W. 2024 · 2024
Closest in time.
BBQ: A hand-built bias benchmark for question answering
Parrish, A.; Chen, A.; Nangia, N.; Padmakumar, V.; Phang, J.; Thompson, J.; Htut, P. M.; and Bowman, S. 2022 · 2086
Closest in time.