Fetching the paper…
Reading the bibliography…
The surge in popularity of large language models has given rise to concerns about biases that these models could learn from humans.
Carroll JB (1964) Language and Thought. Prentice-Hall, Englewood Cliffs, N.J
1964
Earlier work this paper cites.
Tajfel H, Billig MG, Bundy RP, et al (1971) Social categorization and intergroup behaviour. European Journal of Social Psychology 1(2):149–178. 10.1002/ejsp.2420010202 , URL https://onlinelibrary.wiley.com/doi/10.1002/ejsp.2420010202
1971
Earlier work this paper cites.
Tajfel H, Turner JC (1979) An integrativ etheory of intergroup conflict. In: Austin WG, Worchel S (eds) Psychology of Intergroup Relations. Nelson-Hall, Chicago
1979
Earlier work this paper cites.
Van der Dennen JM (1987) Ethnocentrism and in-group/out-group differentiation: A review and interpretation of the literature. The sociobiology of ethnocentrism pp 1–47
1987
Earlier work this paper cites.
Turner JC, Hogg MA, Oakes PJ, et al (1987) Rediscovering the social group: A self-categorization theory. Basil Blackwell, Oxford
1987
Earlier work this paper cites.
Hogg MA, Abrams D (1988) Social Identifications: A Social Psychology of Intergroup Relations and Group Processes
1988
Earlier work this paper cites.
Maass A, Salvi D, Arcuri L, et al (1989) Language use in intergroup contexts: the linguistic intergroup bias. Journal of Personality and Social Psychology 57(6):981–993. 10.1037//0022-3514.57.6.981
1989
Earlier work this paper cites.
Perdue CW, Dovidio JF, Michael B. Gurtman, et al (1990) Us and Them: Social Categorization and the Process of Intergroup Bias. Journal of Personality and Social Psychology 59(3):475–486. 10.1037/0022-3514.59.3.475
1990
Earlier work this paper cites.
Fiedler K, Semin GR, Finkenauer C (1993) The Battle of Words Between Gender Groups: A Language-Based Approach to Intergroup Processes. Human Communication Research 19(3):409–441. 10.1111/j.1468-2958.1993.tb00308.x , URL https://doi.org/10.1111/j.1468-2958.1993.tb00308.x
1993
Earlier work this paper cites.
Maass A, Milesi A, Zabbini S, et al (1995) Linguistic intergroup bias: differential expectancies or in-group protection? Journal of Personality and Social Psychology 68(1):116–126. 10.1037//0022-3514.68.1.116
1995
Earlier work this paper cites.
Kaplan J, McCandlish S, Henighan T, et al (2020) Scaling laws for neural language models. 2001.08361
2001
Earlier work this paper cites.
Viki GT, Winchester L, Titshall L, et al (2006) Beyond secondary emotions: The infrahumanization of outgroups using human–related and animal–related words. Social Cognition 24(6):753–775
2006
Earlier work this paper cites.
Pinter B, Greenwald AG (2011) A comparison of minimal group induction procedures. Group Processes & Intergroup Relations 14(1):81–98. 10.1177/1368430210375251 , URL http://journals.sagepub.com/doi/10.1177/1368430210375251
2011
Earlier work this paper cites.
Iyengar S, Sood G, Lelkes Y (2012) Affect, Not Ideology: A Social Identity Perspective on Polarization. The Public Opinion Quarterly 76(3):405–431. 10.1093/poq/nfs038 , URL https://www.jstor.org/stable/41684577
2012
Earlier work this paper cites.
Hutto C, Gilbert E (2014) VADER: A Parsimonious Rule-Based Model for Sentiment Analysis of Social Media Text. Proceedings of the International AAAI Conference on Web and Social Media 8(1):216–225. 10.1609/icwsm.v8i1.14550 , URL https://ojs.aaai.org/index.php/ICWSM/article/view/14550
2014
Earlier work this paper cites.
Abramowitz AI, Webster S (2016) The rise of negative partisanship and the nationalization of u.s. elections in the 21st century. Electoral Studies 41:12–22. 10.1016/j.electstud.2015.11.001
2015
Earlier work this paper cites.
Bolukbasi T, Chang KW, Zou JY, et al (2016) Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings. In: Advances in Neural Information Processing Systems, vol 29. Curran Associates, Inc., URL https://proceedings.neurips.cc/paper_files/paper/2016/hash/a486cd07e4ac3d270571622f4f316ec5-Abstract.html
2016
Earlier work this paper cites.
Caliskan A, Bryson JJ, Narayanan A (2017) Semantics derived automatically from language corpora contain human-like biases. Science 356(6334):183–186. 10.1126/science.aal4230 , URL https://www.science.org/doi/full/10.1126/science.aal4230
2017
Earlier work this paper cites.
Garg N, Schiebinger L, Jurafsky D, et al (2018) Word embeddings quantify 100 years of gender and ethnic stereotypes. Proceedings of the National Academy of Sciences 115(16):E3635–E3644. 10.1073/pnas.1720347115 , URL https://www.pnas.org/doi/10.1073/pnas.1720347115
2018
Earlier work this paper cites.
Mackie DM, Smith ER (2018) Intergroup Emotions Theory: Production, Regulation, and Modification of Group-Based Emotions. In: Advances in Experimental Social Psychology, vol 58. Elsevier, p 1–69, 10.1016/bs.aesp.2018.03.001 , URL https://linkinghub.elsevier.com/retrieve/pii/S0065260118300121
2018
Earlier work this paper cites.
Bordia S, Bowman SR (2019) Identifying and Reducing Gender Bias in Word-Level Language Models. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Student Research Workshop. Association for Computational Linguistics, Minneapolis, Minnesota, pp 7–15, 10.18653/v1/N19-3002 , URL https://aclanthology.org/N19-3002
2019
Earlier work this paper cites.
Gokaslan A, Cohen V (2019) OpenWebText Corpus. URL http://Skylion007.github.io/OpenWebTextCorpus
2019
Earlier work this paper cites.
Iyengar S, Lelkes Y, Levendusky M, et al (2019) The Origins and Consequences of Affective Polarization in the United States. Annual Review of Political Science 22(1):129–146. 10.1146/annurev-polisci-051117-073034 , URL https://www.annualreviews.org/doi/10.1146/annurev-polisci-051117-073034
2019
Earlier work this paper cites.
Radford A, Wu J, Child R, et al (2019) Language models are unsupervised multitask learners. URL https://openai.com/research/better-language-models
2019
Earlier work this paper cites.
Brown TB, Mann B, Ryder N, et al (2020) Language Models are Few-Shot Learners. In: Larochelle H, Ranzato M, Hadsell R, et al (eds) Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html
2020
Earlier work this paper cites.
Gao L, Biderman S, Black S, et al (2020) The pile: An 800gb dataset of diverse text for language modeling. 2101.00027
2020
Earlier work this paper cites.
Holtzman A, Buys J, Du L, et al (2020) The Curious Case of Neural Text Degeneration. In: International Conference on Learning Representations, URL https://openreview.net/forum?id=rygGQyrFvH
2020
Cited alongside, same era.
Raffel C, Shazeer N, Roberts A, et al (2020) Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. J Mach Learn Res 21(1)
2020
Cited alongside, same era.
Wolf T, Debut L, Sanh V, et al (2020) Transformers: State-of-the-Art Natural Language Processing. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics, Online, pp 38–45, 10.18653/v1/2020.emnlp-demos.6 , URL https://aclanthology.org/2020.emnlp-demos.6
2020
Cited alongside, same era.
Abid A, Farooqi M, Zou J (2021) Persistent anti-muslim bias in large language models. In: Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pp 298–306
2021
Cited alongside, same era.
Ivison H, Wang Y, Pyatkin V, et al (2023) Camels in a changing climate: Enhancing lm adaptation with tulu 2. 2311.10702
2023
Closest in time.
Iyer S, Lin XV, Pasunuru R, et al (2023) Opt-iml: Scaling language model instruction meta learning through the lens of generalization. 2212.12017
2023
Closest in time.
Jakesch M, Bhat A, Buschek D, et al (2023) Co-Writing with Opinionated Language Models Affects Users’ Views. In: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. ACM, Hamburg Germany, pp 1–15, 10.1145/3544548.3581196 , URL https://dl.acm.org/doi/10.1145/3544548.3581196
2023
Closest in time.
Jentzsch S, Kersting K (2023) ChatGPT is fun, but it is not funny! Humor is still challenging Large Language Models. In: Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis. Association for Computational Linguistics, Toronto, Canada, pp 325–340, 10.18653/v1/2023.wassa-1.29 , URL https://aclanthology.org/2023.wassa-1.29
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ahn J, Oh A (2021) Mitigating language-dependent ethnic bias in BERT. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, pp 533–549, 10.18653/v1/2021.emnlp-main.42 , URL https://aclanthology.org/2021.emnlp-main.42
2021
Cited alongside, same era.
Bender EM, Gebru T, McMillan-Major A, et al (2021) On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In: Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. ACM, Virtual Event Canada, pp 610–623, 10.1145/3442188.3445922 , URL https://dl.acm.org/doi/10.1145/3442188.3445922
2021
Cited alongside, same era.
Rathje S, Van Bavel JJ, van der Linden S (2021) Out-group animosity drives engagement on social media. Proceedings of the National Academy of Sciences 118(26):e2024292118. 10.1073/pnas.2024292118 , URL https://pnas.org/doi/full/10.1073/pnas.2024292118
2021
Cited alongside, same era.
Bai Y, Jones A, Ndousse K, et al (2022) Training a helpful and harmless assistant with reinforcement learning from human feedback. 2204.05862
2022
Cited alongside, same era.
Boyd RL, Ashokkumar A, Seraj S, et al (2022) The development and psychometric properties of LIWC-22. University of Texas at Austin, Austin, TX, URL https://www.liwc.app/static/documents/LIWC-22%20Manual%20-%20Development%20and%20Psychometrics.pdf
2022
Cited alongside, same era.
Caron G, Srivastava S (2022) Identifying and manipulating the personality traits of language models. 2212.10276
2022
Cited alongside, same era.
Chung HW, Hou L, Longpre S, et al (2022) Scaling instruction-finetuned language models. 2210.11416
2022
Cited alongside, same era.
Dettmers T, Lewis M, Belkada Y, et al (2022) GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale. In: Koyejo S, Mohamed S, Agarwal A, et al (eds) Advances in Neural Information Processing Systems, vol 35. Curran Associates, Inc., pp 30318–30332, URL https://proceedings.neurips.cc/paper_files/paper/2022/file/c3ba4962c05c49636d4c6206a97e9c8a-Paper-Conference.pdf
2022
Cited alongside, same era.
2023
Closest in time.
Jiang AQ, Sablayrolles A, Mensch A, et al (2023) Mistral 7b. 2310.06825
2023
Closest in time.
Milmo D (2023) Chatgpt reaches 100 million users two months after launch. URL https://www.theguardian.com/technology/2023/feb/02/chatgpt-100-million-users-open-ai-fastest-growing-app
2023
Closest in time.
Muennighoff N, Wang T, Sutawika L, et al (2023) Crosslingual generalization through multitask finetuning. 2211.01786
2023
Closest in time.
Park JS, O’Brien JC, Cai CJ, et al (2023) Generative Agents: Interactive Simulacra of Human Behavior. In: In the 36th Annual ACM Symposium on User Interface Software and Technology (UIST ’23). Association for Computing Machinery, New York, NY, USA, UIST ’23
2023
Closest in time.
Poddar R, Sinha R, Naaman M, et al (2023) AI Writing Assistants Influence Topic Choice in Self-Presentation. In: Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, CHI EA ’23, pp 1–6, 10.1145/3544549.3585893 , URL https://dl.acm.org/doi/10.1145/3544549.3585893
2023
Closest in time.
Taori R, Gulrajani I, Zhang T, et al (2023) Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca , URL https://github.com/tatsu-lab/stanford_alpaca
2023
Closest in time.
Tunstall L, Beeching E, Lambert N, et al (2023) Zephyr: Direct distillation of lm alignment. 2310.16944
2023
Closest in time.
Wang G, Cheng S, Zhan X, et al (2023) Openchat: Advancing open-source language models with mixed-quality data. arXiv preprint arXiv:230911235
2023
Closest in time.
Webb T, Holyoak KJ, Lu H (2023) Emergent analogical reasoning in large language models. Nature Human Behaviour pp 1–16. 10.1038/s41562-023-01659-w , URL https://www.nature.com/articles/s41562-023-01659-w
2023
Closest in time.
2023
Closest in time.
Zheng L, Chiang WL, Sheng Y, et al (2023) Lmsys-chat-1m: A large-scale real-world llm conversation dataset. 2309.11998
2023
Closest in time.
Zhu B, Frick E, Wu T, et al (2023) Starling-7b: Improving llm helpfulness & harmlessness with rlaif
2023
Closest in time.
Anil C, Durmus E, Sharma M, et al (2024) Many-shot jailbreaking. URL https://www.anthropic.com/research/many-shot-jailbreaking
2024
Closest in time.
Groeneveld D, Beltagy I, Walsh P, et al (2024) Olmo: Accelerating the science of language models. 2402.00838
2024
Closest in time.
Jiang AQ, Sablayrolles A, Roux A, et al (2024) Mixtral of experts. 2401.04088
2024
Closest in time.
Kosinski M (2024) Evaluating large language models in theory of mind tasks. 2302.02083
2024
Closest in time.
Laban P, Murakhovs’ka L, Xiong C, et al (2024) Are you sure? challenging llms leads to performance drops in the flipflop experiment. 2311.08596
2024
Closest in time.
Microsoft (2024) Global online safety survey results. URL https://www.microsoft.com/en-us/DigitalSafety/research/global-online-safety-survey
2024
Closest in time.
OpenAI, Achiam J, Adler S, et al (2024) Gpt-4 technical report. 2303.08774
2024
Closest in time.
Sharma M, Tong M, Korbak T, et al (2024) Towards understanding sycophancy in language models. In: The Twelfth International Conference on Learning Representations, URL https://openreview.net/forum?id=tvhaxkMKAn
2024
Closest in time.
Team G, Mesnard T, Hardin C, et al (2024) Gemma: Open models based on gemini research and technology. 2403.08295
2024
Closest in time.
Zhao W, Ren X, Hessel J, et al (2024) Wildchat: 1m chatgpt interaction logs in the wild. 2405.01470
2024
Closest in time.