Fetching the paper…
Reading the bibliography…
AI systems often exhibit political bias, influencing users' opinions and decisions.
Dissenting opinion in abrams v. united states
Holmes, O. W · 1919
Earlier work this paper cites.
Concurring opinion in Whitney v. California
Brandeis, L. D · 1927
Earlier work this paper cites.
Universal Declaration of Human Rights, 1948
United Nations · 1948
Earlier work this paper cites.
Report on the university’s role in political and social action, 1967
Kalven Committee · 1967
Earlier work this paper cites.
The Morality of Freedom
Raz, J · 1986
Earlier work this paper cites.
Political Liberalism
Rawls, J · 1993
Earlier work this paper cites.
The reflexive self: A sociological perspective
Falk, R. and Miller, N · 1998
Earlier work this paper cites.
Confirmation bias: A ubiquitous phenomenon in many guises
Nickerson, R. S · 1998
Earlier work this paper cites.
Motivated skepticism in the evaluation of political beliefs
Taber, C. S. and Lodge, M · 2006
Earlier work this paper cites.
The impossibility of political neutrality
Iwasa, N · 2010
Earlier work this paper cites.
Singapore Statement on Research Integrity, 2010
World Conference on Research Integrity · 2010
Earlier work this paper cites.
The Filter Bubble: What the Internet Is Hiding from You
Pariser, E · 2011
Earlier work this paper cites.
Neutrality — a survivor?
Gavouneli, M · 2012
Earlier work this paper cites.
Ranking digital rights: Advancing freedom of expression and privacy on the internet, 2013
Ranking Digital Rights · 2013
Earlier work this paper cites.
Understanding the determinants of political ideology: Implications of structural complexity
Feldman, S. and Johnston, C · 2014
Earlier work this paper cites.
Introduction , pp. 1–21
Merrill, R. and Weinstock, D · 2014
Earlier work this paper cites.
SPJ code of ethics, 2014
Society of Professional Journalists · 2014
Earlier work this paper cites.
Explanatory preferences shape learning and inference
Lombrozo, T · 2016
Earlier work this paper cites.
A Three-Decade Retrospective on the Hostile Media Effect
Perloff, R. M · 2018
Earlier work this paper cites.
Less than you think: Prevalence and predictors of fake news dissemination on facebook
Guess, A., Nagler, J., and Tucker, J · 2019
Earlier work this paper cites.
Information overload in the information age: a review of the literature from business administration, business psychology, and related disciplines with a bibliometric approach and framework development
Roetzel, P. G · 2019
Earlier work this paper cites.
Reflective practice in the art and science of counselling: A scoping review
Taylor, D · 2020
Earlier work this paper cites.
Post-hoc interpretability for neural NLP: A survey
Madsen, A., Reddy, S., and Chandar, A. P. S · 2021
Earlier work this paper cites.
American politics in two dimensions: Partisan and ideological identities versus anti-establishment orientations
Uscinski, J. E., Enders, A. M., Seelig, M. I., Klofstad, C. A., Funchion, J. R., Everett, C., Wuchty, S., Premaratne, K., and Murthi, M. N · 2021
Earlier work this paper cites.
Censorship of online encyclopedias: Implications for NLP models
Yang, E. and Roberts, M. E · 2021
Earlier work this paper cites.
Bothsiderism
Aikin, S. F. and Casey, J. P · 2022
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI feedback, 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., and Kaplan, J · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Comment on artificial intelligence discussion in “the future of AI with Yann LeCun and Lex Fridman”, September 2022
LeCun, Y · 2022
Earlier work this paper cites.
The digital divide: A review and future research agenda
Lythreatis, S., Singh, S. K., and El-Kassar, A.-N · 2022
Earlier work this paper cites.
The myth of race-neutral policy, June 2022
Maye, A. A · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Invariant language modeling
Peyrard, M., Ghotra, S., Josifoski, M., Agarwal, V., Patra, B., Carignan, D., Kiciman, E., Tiwary, S., and West, R · 2022
Earlier work this paper cites.
Fairness-aware adversarial perturbation towards bias mitigation for deployed deep models
Wang, Z., Dong, X., Xue, H., Zhang, Z., Chiu, W., Wei, T., and Ren, K · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E. H., Xia, F., Le, Q., and Zhou, D · 2022
Earlier work this paper cites.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., Haas, J., Legassick, S., Irving, G., and Gabriel, I · 2022
Cited alongside, same era.
Fairness reprogramming
Zhang, G., Zhang, Y., Zhang, Y., Fan, W., Li, Q., Liu, S., and Chang, S · 2022
Cited alongside, same era.
Conceptualizing the relationship between AI explanations and user agency, 2023
Adenuga, I. and Dodge, J · 2023
Cited alongside, same era.
From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models
Feng, S., Park, C. Y., Liu, Y., and Tsvetkov, Y · 2023
Cited alongside, same era.
Beyond bias and compliance: Towards individual agency and plurality of ethics in ai, 2023
Gilbert, T. K., Brozek, M. W., and Brozek, A · 2023
Differences in misinformation sharing can lead to politically asymmetric sanctions
Mosleh, M., Yang, Q., Zaman, T., Pennycook, G., and Rand, D. G · 2024
Later among the works it cites.
GPT-4o mini (OpenAI), 2024
OpenAI · 2024
Later among the works it cites.
Toward democracy levels for AI
Ovadya, A., Thorburn, L., Redman, K., Devine, F., Milli, S., Revel, M., and Kasirzadeh, A · 2024
Later among the works it cites.
Toward democracy levels for AI
Ovadya, A., Thorburn, L., Redman, K., Devine, F., Milli, S., Revel, M., Konya, A., and Kasirzadeh, A · 2024
Later among the works it cites.
Whose side are you on? investigating the political stance of large language models
Pit, P., Ma, X., Conway, M., Chen, Q., Bailey, J., Pit, H., Keo, P., Diep, W., and Jiang, Y.-G · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
News summarization and evaluation in the era of GPT-3, 2023
Goyal, T., Li, J. J., and Durrett, G · 2023
Cited alongside, same era.
The AI regulatory alignment problem
Guha, N., Lawrence, C. M., Gailmard, L. A., Rodolfa, K. T., Surani, F., Bommasani, R., Raji, I. D., Cuéllar, M.-F., Honigsberg, C., Liang, P., and Ho, D. E · 2023
Cited alongside, same era.
Rationalization for explainable NLP: a survey
Gurrapu, S., Kulkarni, A., Huang, L., Lourentzou, I., Freeman, L. J., and Batarseh, F. A · 2023
Cited alongside, same era.
Bad actor, good advisor: Exploring the role of large language models in fake news detection
Hu, B., Sheng, Q., Cao, J., Shi, Y., Li, Y., Wang, D., and Qi, P · 2023
Cited alongside, same era.
Llama Guard: LLM-based input-output safeguard for human-AI conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Chen, B., Sun, R., Wang, Y., and Yang, Y · 2023
Cited alongside, same era.
The dark side of ChatGPT: Legal and ethical challenges from stochastic parrots and hallucination
Li, Z. M · 2023
Cited alongside, same era.
Hidden persuaders: LLMs’ political leaning and their influence on voters
Potter, Y., Lai, S., Kim, J., Evans, J., and Song, D · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Later among the works it cites.
How artificial intelligence can influence elections: Analyzing the large language models (LLMs) political bias
Rotaru, G.-C., Anagnoste, S., and Oancea, V.-M · 2024
Later among the works it cites.
XSTest: A test suite for identifying exaggerated safety behaviours in large language models
Röttger, P., Kirk, H., Vidgen, B., Attanasio, G., Bianchi, F., and Hovy, D · 2024
Later among the works it cites.
A roadmap to pluralistic alignment
Sorensen, T., Moore, J., Fisher, J., Gordon, M., Mireshghallah, N., Rytting, C. M., Ye, A., Jiang, L., Lu, X., Dziri, N., Althoff, T., and Choi, Y · 2024
Later among the works it cites.
The political compass, 2024
The Political Compass · 2024
Later among the works it cites.
Do-not-answer: Evaluating safeguards in LLMs
Wang, Y., Li, H., Han, X., Nakov, P., and Baldwin, T · 2024
Later among the works it cites.
The art of refusal: A survey of abstention in large language models
Wen, B., Yao, J., Feng, S., Xu, C., Tsvetkov, Y., Howe, B., and Wang, L. L · 2024
Later among the works it cites.
Conspiracy theories in united states politics, 2024
Wikipedia · 2024
Later among the works it cites.
Unpacking political bias in large language models: Insights across topic polarization
Yang, K., Li, H., Chu, Y., Lin, Y., Peng, T.-Q., and Liu, H · 2024
Later among the works it cites.
Benchmarking large language models for news summarization
Zhang, T., Ladhak, F., Durmus, E., Liang, P., McKeown, K., and Hashimoto, T. B · 2024
Later among the works it cites.
LLaMA 3.3-70B (Meta), 2024
AI@Meta · 2025
Closest in time.
Claude 3.5 Haiku (Anthropic), 2024
Anthropic · 2025
Closest in time.
Our Approach to User Safety, 2025b
Anthropic · 2025
Closest in time.
System Prompts - Anthropic Documentation, 2025
Anthropic · 2025
Closest in time.
The 2024 Foundation Model Transparency Index, 2025
Bommasani, R., Klyman, K., Kapoor, S., Longpre, S., Xiong, B., Maslej, N., and Liang, P · 2025
Closest in time.
Gemini Flash (Google DeepMind), 2024a
DeepMind, G · 2025
Closest in time.
Gemini Pro (Google DeepMind), 2024b
DeepMind, G · 2025
Closest in time.
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning, 2025
DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., Xue, B., Wang, B., Wu, B., Feng, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., Dai, D., Chen, D., Ji, D., Li, E., Lin, F., Dai, F., Luo, F., Hao, G., Chen, G., Li, G., Zhang, H., Bao, H., Xu, H., Wang, H., Ding, H., Xin, H., Gao, H., Qu, H., Li, H., Guo, J., Li, J., Wang, J., Chen, J., Yuan, J., Qiu, J., Li, J., Cai, J. L., Ni, J., Liang, J., Chen, J., Dong, K., Hu, K., Gao, K., Guan, K., Huang, K., Yu, K., Wang, L., Zhang, L., Zhao, L., Wang, L., Zhang, L., Xu, L., Xia, L., Zhang, M., Zhang, M., Tang, M., Li, M., Wang, M., Li, M., Tian, N., Huang, P., Zhang, P., Wang, Q., Chen, Q., Du, Q., Ge, R., Zhang, R., Pan, R., Wang, R., Chen, R. J., Jin, R. L., Chen, R., Lu, S., Zhou, S., Chen, S., Ye, S., Wang, S., Yu, S., Zhou, S., Pan, S., Li, S. S., Zhou, S., Wu, S., Ye, S., Yun, T., Pei, T., Sun, T., Wang, T., Zeng, W., Zhao, W., Liu, W., Liang, W., Gao, W., Yu, W., Zhang, W., Xiao, W. L., An, W., Liu, X., Wang, X., Chen, X., Nie, X., Cheng, X., Liu, X., Xie, X., Liu, X., Yang, X., Li, X., Su, X., Lin, X., Li, X. Q., Jin, X., Shen, X., Chen, X., Sun, X., Wang, X., Song, X., Zhou, X., Wang, X., Shan, X., Li, Y. K., Wang, Y. Q., Wei, Y. X., Zhang, Y., Xu, Y., Li, Y., Zhao, Y., Sun, Y., Wang, Y., Yu, Y., Zhang, Y., Shi, Y., Xiong, Y., He, Y., Piao, Y., Wang, Y., Tan, Y., Ma, Y., Liu, Y., Guo, Y., Ou, Y., Wang, Y., Gong, Y., Zou, Y., He, Y., Xiong, Y., Luo, Y., You, Y., Liu, Y., Zhou, Y., Zhu, Y. X., Xu, Y., Huang, Y., Li, Y., Zheng, Y., Zhu, Y., Ma, Y., Tang, Y., Zha, Y., Yan, Y., Ren, Z. Z., Ren, Z., Sha, Z., Fu, Z., Xu, Z., Xie, Z., Zhang, Z., Hao, Z., Ma, Z., Yan, Z., Wu, Z., Gu, Z., Zhu, Z., Liu, Z., Li, Z., Xie, Z., Song, Z., Pan, Z., Huang, Z., Xu, Z., Zhang, Z., and Zhang, Z · 2025
Closest in time.
Biased AI can influence political decision-making
Fisher, J., Feng, S., Aron, R., Richardson, T., Choi, Y., Fisher, D. W., Pan, J., Tsvetkov, Y., and Reinecke, K · 2025
Closest in time.
Ethical concerns mount as ai takes bigger decision-making role
Gazette, H · 2025
Closest in time.
From distributional to overton pluralism: Investigating large language model alignment
Lake, T., Choi, E., and Durrett, G · 2025
Closest in time.
Investigating bias in LLM-based bias detection: Disparities between LLMs and human perception
Lin, L., Wang, L., Guo, J., and Wong, K.-F · 2025
Closest in time.
The chaos at OpenAI is a death knell for AI self-regulation
Lostri, E., Rozenshtein, A. Z., and Sharma, C · 2025
Closest in time.
How do people react to political bias in generative artificial intelligence (AI)?
Messer, U · 2025
Closest in time.
OLMo, T., Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., Lambert, N., Schwenk, D., Tafjord, O., Anderson, T., Atkinson, D., Brahman, F., Clark, C., Dasigi, P., Dziri, N., Guerquin, M., Ivison, H., Koh, P. W., Liu, J., Malik, S., Merrill, W., Miranda, L. J. V., Morrison, J., Murray, T., Nam, C., Pyatkin, V., Rangapur, A., Schmitz, M., Skjonsberg, S., Wadden, D., Wilhelm, C., Wilson, M., Zettlemoyer, L., Farhadi, A., Smith, N. A., and Hajishirzi, H · 2025
Closest in time.
Moderation API, 2025
OpenAI · 2025
Closest in time.
LLMs’ potential influences on our democracy: Challenges and opportunities
Potter, Y., Choi, Y., Rand, D., and Song, D · 2025
Closest in time.
Wikipedia: Neutral Point of View – Wikipedia, The Free Encyclopedia, 2025
Wikipedia contributors · 2025
Closest in time.