Fetching the paper…
Reading the bibliography…
Large language models (LLMs) exhibit cognitive biases -- systematic tendencies of irrational decision-making, similar to those seen in humans.
Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty
Amos Tversky and Daniel Kahneman · 1974
Earlier work this paper cites.
The framing of decisions and the psychology of choice
Amos Tversky and Daniel Kahneman · 1981
Earlier work this paper cites.
On the conflict between logic and belief in syllogistic reasoning
JSBT Evans, Julie L Barston, and Paul Pollard · 1983
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
An empirical comparison of instance attribution methods for nlp
Pouya Pezeshkpour, Sarthak Jain, Byron C. Wallace, and Sameer Singh · 2021
Earlier work this paper cites.
Towards a comprehensive understanding and accurate evaluation of societal biases in pre-trained transformers
Andrew Silva, Pradyumna Tambwekar, and M. Gombolay · 2021
Earlier work this paper cites.
Using cognitive psychology to understand gpt-3
Marcel Binz and Eric Schulz · 2022
Earlier work this paper cites.
Language models show human-like content effects on reasoning
Ishita Dasgupta, Andrew K Lampinen, Stephanie CY Chan, Antonia Creswell, Dharshan Kumaran, James L McClelland, and Felix Hill · 2022
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Earlier work this paper cites.
Improving gender fairness of pre-trained language models without catastrophic forgetting
Zahra Fatemi, Chen Xing, Wenhao Liu, and Caiming Xiong · 2023
Earlier work this paper cites.
Bias and fairness in large language models: A survey
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md. Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen Ahmed · 2023
Earlier work this paper cites.
Camels in a changing climate: Enhancing lm adaptation with tulu 2
Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A Smith, Iz Beltagy, et al · 2023
Earlier work this paper cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Earlier work this paper cites.
The flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al · 2023
Cited alongside, same era.
Biases in large language models: Origins, inventory, and discussion
Roberto Navigli, Simone Conia, and Björn Ross · 2023
Cited alongside, same era.
Gabriel Stanovsky, Tomasz Limisiewicz, David Marevcek, and Bar Iluz · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Comparing rationality between large language models and humans: Insights and open questions
Yuhe Ke, Rui Yang, Sui An Lie, Taylor Xin Yi Lim, Hairil Rizal Bin Abdullah, Daniel Shu Wei Ting, and Nan Liu · 2024
Later among the works it cites.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang · 2024
Later among the works it cites.
A comprehensive evaluation of cognitive biases in llms
Simon Malberg, Roman Poletukhin, Carolin M Schuster, and Georg Groh · 2024
Later among the works it cites.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dana Alsagheer, Rabimba Karanjai, Nour Diallo, Weidong Shi, Yang Lu, Suha Beydoun, and Qiaoning Zhang · 2024
Cited alongside, same era.
Generalization v.s. memorization: Tracing language models’ capabilities back to pretraining data
Antonis Antoniades, Xinyi Wang, Yanai Elazar, Alfonso Amayuelas, Alon Albalak, Kexun Zhang, and William Yang Wang · 2024
Cited alongside, same era.
Jeremie Bogaert, Marie-Catherine de Marneffe, Antonin Descampe, Louis Escouflaire, Cedrick Fairon, and Francois-Xavier Standaert · 2024
Cited alongside, same era.
Agr: Age group fairness reward for bias mitigation in llms
Shuirong Cao, Ruoxi Cheng, and Zhiqiang Wang · 2024
Cited alongside, same era.
Target-aware language modeling via granular data sampling
Ernie Chang, Pin-Jie Lin, Yang Li, Changsheng Zhao, Daeil Kim, Rastislav Rabatin, Zechun Liu, Yangyang Shi, and Vikas Chandra · 2024
Cited alongside, same era.
Fairness in large language models: A taxonomic survey
Zhibo Chu, Zichong Wang, and Wenbin Zhang · 2024
Cited alongside, same era.
Cognitive bias in decision-making with LLMs
Jessica Maria Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zexue He · 2024
Cited alongside, same era.
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al · 2024
Cited alongside, same era.
Samuel Schmidgall, Carl Harris, Ime Essien, Daniel Olshvang, Tawsifur Rahman, Ji Woong Kim, Rojin Ziaei, Jason Eshraghian, Peter Abadir, and Rama Chellappa · 2024
Later among the works it cites.
Cbeval: A framework for evaluating and interpreting cognitive biases in llms
Ammar Shaikh, Raj Abhijit Dandekar, Sreedath Panat, and Rajat Dandekar · 2024
Later among the works it cites.
Lora vs full fine-tuning: An illusion of equivalence
Reece Shuttleworth, Jacob Andreas, Antonio Torralba, and Pratyusha Sharma · 2024
Later among the works it cites.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2024
Later among the works it cites.
Language models are susceptible to incorrect patient self-diagnosis in medical applications
Rojin Ziaei and Samuel Schmidgall · 2024
Later among the works it cites.
Neuroplasticity and corruption in model mechanisms: A case study of indirect object identification
Vishnu Kabir Chhabra, Ding Zhu, and Mohammad Mahdi Khalili · 2025
Closest in time.
The impact of initialization on lora finetuning dynamics
Soufiane Hayou, Nikhil Ghosh, and Bin Yu · 2025
Closest in time.
Wildframe: Comparing framing in humans and llms on naturally occurring texts
Gili Lior, Liron Nacchace, and Gabriel Stanovsky · 2025
Closest in time.
Cognitive debiasing large language models for decision-making
Yougang Lyu, Shijie Ren, Yue Feng, Zihan Wang, Zhumin Chen, Zhaochun Ren, and Maarten de Rijke · 2025
Closest in time.
Towards understanding fine-tuning mechanisms of llms via circuit analysis
Xu Wang, Yan Hu, Wenyu Du, Reynold Cheng, Benyou Wang, and Difan Zou · 2025
Closest in time.