Fetching the paper…
Reading the bibliography…
We present a large-scale evaluation of 30 cognitive biases in 20 state-of-the-art large language models (LLMs) under various decision-making scenarios.
A constant error in psychological ratings
Edward Lee Thorndike. 1920 · 1920
Earlier work this paper cites.
A method of estimating plane vulnerability based on damage of survivors
Abraham Wald. 1943 · 1943
Earlier work this paper cites.
Bandwagon, snob, and veblen effects in the theory of consumers’ demand
H. Leibenstein. 1950 · 1950
Earlier work this paper cites.
The relationship between the judged desirability of a trait and the probability that the trait will be endorsed
Allen L Edwards. 1953 · 1953
Earlier work this paper cites.
The social desirability variable in personality assessment and research
Allen L Edwards. 1957 · 1957
Earlier work this paper cites.
A new scale of social desirability independent of psychopathology
Douglas P Crowne and David Marlowe. 1960 · 1960
Earlier work this paper cites.
On the failure to eliminate hypotheses in a conceptual task
P. C. Wason. 1960 · 1960
Earlier work this paper cites.
Relative and absolute strength of response as a function of frequency of reinforcement
Richard J Herrnstein. 1961 · 1961
Earlier work this paper cites.
A theory of psychological reactance
Jack W Brehm. 1966 · 1966
Earlier work this paper cites.
Reasoning
Peter C. Wason. 1966 · 1966
Earlier work this paper cites.
The attribution of attitudes
Edward E Jones and Victor A Harris. 1967 · 1967
Earlier work this paper cites.
Conservatism in human information processing (excerpted)
Ward Edwards. 1982 · 1968
Earlier work this paper cites.
Negativity in evaluations
David E Kanouse and L Reid Hanson Jr. 1972 · 1972
Earlier work this paper cites.
Availability: A heuristic for judging frequency and probability
Amos Tversky and Daniel Kahneman. 1973 · 1973
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Amos Tversky and Daniel Kahneman. 1974 · 1974
Earlier work this paper cites.
The illusion of control
Ellen J Langer. 1975 · 1975
Earlier work this paper cites.
Self-serving biases in the attribution of causality: Fact or fiction?
Dale T Miller and Michael Ross. 1975 · 1975
Earlier work this paper cites.
The effects of automobile safety regulation
Sam Peltzman. 1975 · 1975
Earlier work this paper cites.
Observer bias: A stringent test of behavior engulfing the field
Melvin Snyder and Arthur Frankel. 1976 · 1976
Earlier work this paper cites.
Egotism and attribution
Melvin L Snyder, Walter G Stephan, and David Rosenfield. 1976 · 1976
Earlier work this paper cites.
Knee-deep in the big muddy: a study of escalating commitment to a chosen course of action
Barry M. Staw. 1976 · 1976
Earlier work this paper cites.
The halo effect: Evidence for unconscious alteration of judgments
Richard E Nisbett and Timothy D Wilson. 1977 · 1977
Earlier work this paper cites.
The intuitive psychologist and his shortcomings: Distortions in the attribution process
Lee Ross. 1977 · 1977
Earlier work this paper cites.
Self-serving biases in the attribution process: A reexamination of the fact or fiction question
Gifford W Bradley. 1978 · 1978
Earlier work this paper cites.
The effect of imagining an event on expectations for the event: An interpretation in terms of the availability heuristic
John S Carroll. 1978 · 1978
Earlier work this paper cites.
Hypothesis testing in social judgment
Mark Snyder and William Swann. 1978 · 1978
Earlier work this paper cites.
Some additional evidence on survival biases
Ray Ball and Ross Watts. 1979 · 1979
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
Daniel Kahneman and Amos Tversky. 1979 · 1979
Earlier work this paper cites.
Toward a positive theory of consumer choice
Richard Thaler. 1980 · 1980
Earlier work this paper cites.
Ubiquitous halo
William H Cooper. 1981 · 1981
Earlier work this paper cites.
The escalation of commitment to a course of action
Barry M. Staw. 1981 · 1981
Earlier work this paper cites.
The framing of decisions and the psychology of choice
Amos Tversky and Daniel Kahneman. 1981 · 1981
Earlier work this paper cites.
The halo effect revisited: Forewarned is not forearmed
Christopher G Wetzel, Timothy D Wilson, and James Kort. 1981 · 1981
Earlier work this paper cites.
The Psychology of Interpersonal Relations
F. Heider. 1982 · 1982
Earlier work this paper cites.
Intuitive prediction: Biases and corrective procedures , page 414–421
Daniel Kahneman and Amos Tversky. 1982 · 1982
Earlier work this paper cites.
Investigating the not invented here (nih) syndrome: A look at the performance, tenure, and communication patterns of 50 r & d project groups
Ralph Katz and Thomas J Allen. 1982 · 1982
Earlier work this paper cites.
The theory of risk homeostasis: implications for safety and health
Gerald JS Wilde. 1982 · 1982
Earlier work this paper cites.
Extensional versus intuitive reasoning: The conjunction fallacy in probability judgment
Amos Tversky and Daniel Kahneman. 1983 · 1983
Earlier work this paper cites.
Priming and frequency estimation: A strict test of the availability heuristic
Adele Gabrielcik and Russell H Fazio. 1984 · 1984
Earlier work this paper cites.
Practical implications of the matching law
Joel Myerson and Sandra Hale. 1984 · 1984
Earlier work this paper cites.
A direct study of halo effect
Sheldon J Lachman and Alan R Bass. 1985 · 1985
Earlier work this paper cites.
The disposition to sell winners too early and ride losers too long: Theory and evidence
Hersh Shefrin and Meir Statman. 1985 · 1985
Earlier work this paper cites.
Mental accounting and consumer choice
Richard Thaler. 1985 · 1985
Earlier work this paper cites.
Fairness and the assumptions of economics
Daniel Kahneman, Jack L Knetsch, and Richard H Thaler. 1986 · 1986
Earlier work this paper cites.
Heuristics and biases in diagnostic reasoning: Ii. congruence, information, and certainty
Jonathan Baron, Jane Beattie, and John C Hershey. 1988 · 1988
Earlier work this paper cites.
The availability heuristic and perceived risk
Valerie S. Folkes. 1988 · 1988
Earlier work this paper cites.
Status quo bias in decision making
William Samuelson and Richard Zeckhauser. 1988 · 1988
Earlier work this paper cites.
Discount rates inferred from decisions: An experimental study
Uri Benzion, Amnon Rapoport, and Joseph Yagil. 1989 · 1989
Earlier work this paper cites.
The optimism bias and traffic accident risk perception
David M. DeJoy. 1989 · 1989
Earlier work this paper cites.
The endowment effect and evidence of nonreversible indifference curves
Jack L Knetsch. 1989 · 1989
Earlier work this paper cites.
Optimistic biases about personal risks
Neil D. Weinstein. 1989 · 1989
Earlier work this paper cites.
Hindsight: Biased judgments of past events after the outcomes are known
Scott A Hawkins and Reid Hastie. 1990 · 1990
Earlier work this paper cites.
Experimental tests of the endowment effect and the coase theorem
Daniel Kahneman, Jack L Knetsch, and Richard H Thaler. 1990 · 1990
Earlier work this paper cites.
Bounded rationality
Herbert A Simon. 1990 · 1990
Earlier work this paper cites.
The hindsight bias: A meta-analysis
Jay J.J Christensen-Szalanski and Cynthia Fobian Willham. 1991 · 1991
Earlier work this paper cites.
Anomalies: The endowment effect, loss aversion, and status quo bias
Daniel Kahneman, Jack L. Knetsch, and Richard H. Thaler. 1991 · 1991
Earlier work this paper cites.
Loss Aversion in Riskless Choice: A Reference-Dependent Model*
Amos Tversky and Daniel Kahneman. 1991 · 1991
Earlier work this paper cites.
Hyperbolic discounting
George Ainslie and Nicholas Haslam. 1992 · 1992
Earlier work this paper cites.
Survivorship bias in performance studies
Stephen J Brown, William Goetzmann, Roger G Ibbotson, and Stephen A Ross. 1992 · 1992
Earlier work this paper cites.
Mental accounting and categorization
Pamela W Henderson and Robert A Peterson. 1992 · 1992
Earlier work this paper cites.
Memory accessibility and probability judgments: an experimental evaluation of the availability heuristic
Colin MacLeod and Lynlee Campbell. 1992 · 1992
Earlier work this paper cites.
Teachers’ ratings of disruptive behaviors: The influence of halo effects
Howard Abikoff, Mary Courtney, William E Pelham, and Harold S Koplewicz. 1993 · 1993
Earlier work this paper cites.
Self-esteem and self-serving biases in reactions to positive and negative events: An integrative review
Bruce Blaine and Jennifer Crocker. 1993 · 1993
Earlier work this paper cites.
Reinvestment decisions by entrepreneurs: Rational decision-making or escalation of commitment?
Anne M McCarthy, F David Schoorman, and Arnold C Cooper. 1993 · 1993
Earlier work this paper cites.
New evidence about the existence of a bandwagon effect in the opinion formation process
Richard Nadeau, Edouard Cloutier, and J.-H. Guay. 1993 · 1993
Earlier work this paper cites.
Exploring the" planning fallacy": Why people underestimate their task completion times
Roger Buehler, Dale Griffin, and Michael Ross. 1994 · 1994
Earlier work this paper cites.
Construction of activity duration and time management potential
Christopher DB Burt and Simon Kemp. 1994 · 1994
Earlier work this paper cites.
Fairness in simple bargaining experiments
Robert Forsythe, Joel L Horowitz, Nathan E Savin, and Martin Sefton. 1994 · 1994
Earlier work this paper cites.
An advanced test of theory of mind: Understanding of story characters’ thoughts and feelings by able autistic, mentally handicapped, and normal children and adults
Francesca GE Happé. 1994 · 1994
Earlier work this paper cites.
Cultural variation in unrealistic optimism: Does the west feel more vulnerable than the east?
Steven J Heine and Darrin R Lehman. 1995 · 1995
Earlier work this paper cites.
Varieties of confirmation bias
Joshua Klayman. 1995 · 1995
Earlier work this paper cites.
Brand equity: The halo effect measure
Lance Leuthesser, Chiranjeev Kohli, and Katrin Harich. 1995 · 1995
Earlier work this paper cites.
Sufficient grounds for optimism?: The relationship between perceived controllability and optimistic bias
Peter Harris. 1996 · 1996
Earlier work this paper cites.
Golden eggs and hyperbolic discounting
David Laibson. 1997 · 1997
Earlier work this paper cites.
Negative information weighs more heavily on the brain: the negativity bias in evaluative categorizations
Tiffany A Ito, Jeff T Larsen, N Kyle Smith, and John T Cacioppo. 1998 · 1998
Cited alongside, same era.
All frames are not created equal: A typology and critical analysis of framing effects
Irwin P. Levin, Sandra L. Schneider, and Gary J. Gaeth. 1998 · 1998
Cited alongside, same era.
Confirmation bias: A ubiquitous phenomenon in many guises
Raymond Nickerson. 1998 · 1998
Cited alongside, same era.
The disposition effect in securities trading: An experimental analysis
Martin Weber and Colin F Camerer. 1998 · 1998
Cited alongside, same era.
Self-threat magnifies the self-serving bias: A meta-analytic integration
W Keith Campbell and Constantine Sedikides. 1999 · 1999
Cited alongside, same era.
Dictionary of Psychology
Mike Cardwell. 1999 · 1999
Can an algorithm reduce the perceived bias of news? testing the effect of machine attribution on news readers’ evaluations of bias, anthropomorphism, and credibility
T Franklin Waddell. 2019 · 2019
Later among the works it cites.
“everything is perfect, and we have no problems”: detecting and limiting social desirability bias in qualitative research
Nicole Bergen and Ronald Labonté. 2020 · 2020
Later among the works it cites.
Detoxify
Laura Hanu and Unitary team. 2020 · 2020
Later among the works it cites.
Confirmation bias in the utilization of others’ opinion strength
Heather Kappes, Ann Harvey, Terry Lohrenz, Pendleton Montague, and Tali Sharot. 2020 · 2020
Later among the works it cites.
Designing to debias: Measuring and reducing public managers’ anchoring bias
Rosanna Nagtegaal, Lars Tummers, Mirko Noordegraaf, and Victor Bekkers. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Advances in research on mental accounting and reason-based choice
Ran Kivetz. 1999 · 1999
Cited alongside, same era.
Mental accounting matters
Richard H Thaler. 1999 · 1999
Cited alongside, same era.
Illusions of control: How we overestimate our personal influence
Suzanne C Thompson. 1999 · 1999
Cited alongside, same era.
Risky business: safety regulations, risk compensation, and individual behavior
James Hedlund. 2000 · 2000
Cited alongside, same era.
Optimism bias and student debt
Hamish GW Seaward and Simon Kemp. 2000 · 2000
Cited alongside, same era.
Loss aversion equilibrium
Jonathan Shalev. 2000 · 2000
Cited alongside, same era.
Arleen Salles, Kathinka Evers, and Michele Farisco. 2020 · 2020
Later among the works it cites.
Survivorship bias
Dirk M Elston. 2021 · 2021
Later among the works it cites.
Influence of the fundamental attribution error on perceptions of blame and negligence
Cassandra Flick and Kimberly Schweitzer. 2021 · 2021
Later among the works it cites.
Are managers susceptible to framing effects? an experimental study of professional judgment of performance metrics
Javier Fuenzalida, Gregg G. Van Ryzin, and Asmus Leth Olsen. 2021 · 2021
Later among the works it cites.
Risk compensation during covid-19: The impact of face mask usage on social distancing
Ashley Luckman, Hossam Zeitoun, Andrea Isoni, Graham Loomes, Ivo Vlaev, Nattavudh Powdthavee, and Daniel Read. 2021 · 2021
Later among the works it cites.
Anchoring bias affects mental model formation and user reliance in explainable ai systems
Mahsan Nourani, Chiradeep Roy, Jeremy E Block, Donald R Honeycutt, Tahrima Rahman, Eric Ragan, and Vibhav Gogate. 2021 · 2021
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 · 2022
Later among the works it cites.
Bandwagon effect revisited: A systematic review to develop future research agenda
Sunali Bindra, Deepika Sharma, Nakul Parameswar, Sanjay Dhir, and Justin Paul. 2022 · 2022
Later among the works it cites.
Ingroup favoritism overrides fairness when resources are limited
Jihwan Chae, Kunil Kim, Yuri Kim, Gahyun Lim, Daeeun Kim, and Hackjin Kim. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He. 2022 · 2022
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022 · 2022
Later among the works it cites.
ZeroGen: Efficient zero-shot learning via dataset generation
Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022 · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Later among the works it cites.
Cognitive biases in natural language: Automatically detecting, differentiating, and measuring bias in text
Kyrtin Atreides and David J Kelley. 2023 · 2023
Later among the works it cites.
A closer look into using large language models for automatic evaluation
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Later among the works it cites.
Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions
John Chung, Ece Kamar, and Saleema Amershi. 2023 · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. 2023 · 2023
Later among the works it cites.
Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt
Thilo Hagendorff, Sarah Fabi, and Michal Kosinski. 2023 · 2023
Later among the works it cites.
Annollm: Making large language models to be better crowdsourced annotators
Xingwei He, Zhenghao Lin, Yeyun Gong, Alex Jin, Hang Zhang, Chen Lin, Jian Jiao, Siu Ming Yiu, Nan Duan, Weizhu Chen, et al. 2023 · 2023
Later among the works it cites.
Mahammed Kamruzzaman, Md Minul Islam Shovon, and Gene Louis Kim. 2023 · 2023
Later among the works it cites.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023 · 2023
Later among the works it cites.
Making large language models better data creators
Dong-Ho Lee, Jay Pujara, Mohit Sewak, Ryen W White, and Sujay Kumar Jauhar. 2023 · 2023
Later among the works it cites.
Evidence for Anchoring Bias During Physician Decision-Making
Dan P. Ly, Paul G. Shekelle, and Zirui Song. 2023 · 2023
Later among the works it cites.
Don’t blame the annotator: Bias already starts in the annotation instructions
Mihir Parmar, Swaroop Mishra, Mor Geva, and Chitta Baral. 2023 · 2023
Later among the works it cites.
Exploring conversational agents as an effective tool for measuring cognitive biases in decision-making
Stephen Pilli. 2023 · 2023
Later among the works it cites.
Predicting loss aversion behavior with machine-learning methods
Ömür Saltık, Wasim Rehman, Rıdvan Söyü, Suleyman Degirmen, and Ahmet Sengonul. 2023 · 2023
Later among the works it cites.
Mental accounting and decision making: a systematic literature review
Emmanuel Marques Silva, Rafael de Lacerda Moreira, and Patricia Maria Bortolon. 2023 · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023 · 2023
Later among the works it cites.
Alaina N Talboy and Elizabeth Fuller. 2023 · 2023
Later among the works it cites.
Theory of mind in large language models: Examining performance of 11 state-of-the-art models vs. children aged 7-10 on advanced tests
Max van Duijn, Bram van Dijk, Tom Kouwenhoven, Werner de Valk, Marco Spruit, and Peter van der Putten. 2023 · 2023
Later among the works it cites.
"kelly is a warm person, joseph is a role model": Gender biases in llm-generated reference letters
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023 · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023 · 2023
Later among the works it cites.
Siren’s song in the ai ocean: a survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. 2023 · 2023
Later among the works it cites.
Large language models are not robust multiple choice selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2023 · 2023
Later among the works it cites.
Instruction-following evaluation for large language models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023 · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. 2024 · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic. 2024 · 2024
Closest in time.
Measuring and mitigating racial bias in large language model mortgage underwriting
Donald E Bowen III, S McKay Price, Luke CD Stein, and Ke Yang. 2024 · 2024
Closest in time.
Humans or llms as the judge? a study on judgement biases
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024 · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Closest in time.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, et al. 2024 · 2024
Closest in time.
Cognitive bias in high-stakes decision-making with llms
Jessica Echterhoff, Yao Liu, Abeer Alessa, Julian McAuley, and Zexue He. 2024 · 2024
Closest in time.
Determinants of llm-assisted decision-making
Eva Eigner and Thorsten Händler. 2024 · 2024
Closest in time.
Bias and fairness in large language models: A survey
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024 · 2024
Closest in time.
Can large language models understand real-world complex instructions?
Qianyu He, Jie Zeng, Wenhao Huang, Lina Chen, Jin Xiao, Qianxi He, Xunzhe Zhou, Jiaqing Liang, and Yanghua Xiao. 2024 · 2024
Closest in time.
Instructed to bias: Instruction-tuned language models exhibit emergent cognitive bias
Itay Itzhak, Gabriel Stanovsky, Nir Rosenfeld, and Yonatan Belinkov. 2024 · 2024
Closest in time.
Persuading across diverse domains: a dataset and persuasion large language model
Chuhao Jin, Kening Ren, Lingzhen Kong, Xiting Wang, Ruihua Song, and Huan Chen. 2024 · 2024
Closest in time.
Controllable text generation for large language models: A survey
Xun Liang, Hanyu Wang, Yezhaohui Wang, Shichao Song, Jiawei Yang, Simin Niu, Jie Hu, Dan Liu, Shunyu Yao, Feiyu Xiong, et al. 2024 · 2024
Closest in time.
On llms-driven synthetic data generation, curation, and evaluation: A survey
Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, and Haobo Wang. 2024 · 2024
Closest in time.
(ir) rationality and cognitive biases in large language models
Olivia Macmillan-Scott and Mirco Musolesi. 2024 · 2024
Closest in time.
Global industry classification standard (gics)
MSCI and S&P Global. 2023 · 2024
Closest in time.
Do language models exhibit the same cognitive biases in problem solving as human learners?
Andreas Opedal, Alessandro Stolfo, Haruki Shirakami, Ying Jiao, Ryan Cotterell, Bernhard Schölkopf, Abulhair Saparov, and Mrinmaya Sachan. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024 · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. 2024 · 2024
Closest in time.
The political preferences of llms
David Rozado. 2024 · 2024
Closest in time.
Addressing cognitive bias in medical language models
Samuel Schmidgall, Carl Harris, Ime Essien, Daniel Olshvang, Tawsifur Rahman, Ji Woong Kim, Rojin Ziaei, Jason Eshraghian, Peter Abadir, and Rama Chellappa. 2024 · 2024
Closest in time.
Putting gpt-4o to the sword: A comprehensive evaluation of language, vision, speech, and multimodal proficiency
Sakib Shahriar, Brady D Lund, Nishith Reddy Mannuru, Muhammad Arbab Arshad, Kadhim Hayawi, Ravi Varma Kumar Bevara, Aashrith Mannuru, and Laiba Batool. 2024 · 2024
Closest in time.
The good, the bad, and the greedy: Evaluation of llms should not ignore non-determinism
Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. 2024 · 2024
Closest in time.
Large language models are inconsistent and biased evaluators
Rickard Stureborg, Dimitris Alikaniotis, and Yoshi Suhara. 2024 · 2024
Closest in time.
Large language models for data annotation: A survey
Zhen Tan, Alimohammad Beigi, Song Wang, Ruocheng Guo, Amrita Bhattacharjee, Bohan Jiang, Mansooreh Karami, Jundong Li, Lu Cheng, and Huan Liu. 2024 · 2024
Closest in time.
Do llms exhibit human-like response biases? a case study in survey design
Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkwar, and Graham Neubig. 2024 · 2024
Closest in time.
Metaphor understanding challenge dataset for LLMs
Xiaoyu Tong, Rochelle Choenni, Martha Lewis, and Ekaterina Shutova. 2024 · 2024
Closest in time.
Mindscope: Exploring cognitive biases in large language models through multi-agent systems
Zhentao Xie, Jiabao Zhao, Yilei Wang, Jinxin Shi, Yanhong Bai, Xingjiao Wu, and Liang He. 2024 · 2024
Closest in time.
Justice or prejudice? quantifying biases in llm-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al. 2024 · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. 2024 · 2024
Closest in time.
Large language model as attributed training data generator: A tale of diversity and bias
Yue Yu, Yuchen Zhuang, Jieyu Zhang, Yu Meng, Alexander J Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang. 2024 · 2024
Closest in time.
Determinants of social desirability bias in sensitive surveys: a literature review
Ivar Krumpal. 2013 · 2047
Closest in time.
Does narrative information bias individual’s decision making? a systematic review
Anna Winterbottom, Hilary L Bekker, Mark Conner, and Andrew Mooney. 2008 · 2088
Closest in time.