Fetching the paper…
Reading the bibliography…
Large language models (LLMs) match and sometimes exceeding human performance in many domains.
“Verification of Forecasts Expressed in Terms of Probability”
Glenn Brier · 1950
Earlier work this paper cites.
“Do those who know more also know more about how much they know?”
Sarah Lichtenstein and Baruch Fischhoff · 1977
Earlier work this paper cites.
“The trouble with overconfidence.”
Don Moore and Paul Healy · 2008
Earlier work this paper cites.
“The chess master and the computer”
Garry Kasparov · 2010
Earlier work this paper cites.
“A Continuum of Learning: From Rote Memorization to Meaningful Learning in Organic Chemistry”
Nathaniel Grove and Stacey Bretz · 2012
Earlier work this paper cites.
“When is a Crowd Wise?”
Clintin. Davis-Stober, David. Budescu, Jason Dana and Stephen. Broomell · 2014
Earlier work this paper cites.
“The Wisdom of Select Crowds”
Albert. Mannes, Jack. Soll and Richard. Larrick · 2014
Earlier work this paper cites.
“Forecasting Tournaments: Tools for Increasing Transparency and Improving the Quality of Debate”
Philip. Tetlock, Barbara Mellers, Nick Rohrbaugh and Eva Chen · 2014
Earlier work this paper cites.
“Identifying Expertise to Extract the Wisdom of Crowds”
David Budescu and Eva Chen · 2015
Earlier work this paper cites.
“Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions”
Barbara Mellers et al · 2015
Earlier work this paper cites.
“The psychology of intelligence analysis: Drivers of prediction accuracy in world politics.”
Barbara Mellers et al · 2015
Earlier work this paper cites.
“Irrational exuberance”
Robert Shiller · 2015
Earlier work this paper cites.
“Developing expert political judgment: The impact of training and practice on judgmental accuracy in geopolitical forecasting tournaments”
Welton Chang, Eva Chen, Barbara Mellers and Philip Tetlock · 2016
Earlier work this paper cites.
“Superforecasting: How to upgrade your company’s judgment”
Paul Schoemaker and Philip Tetlock · 2016
Earlier work this paper cites.
“Superforecasting: The Art and Science of Prediction”
Philip. Tetlock and Dan Gardner · 2016
Earlier work this paper cites.
“Distilling the wisdom of crowds: Prediction markets vs. prediction polls”
Pavel Atanasov et al · 2017
Earlier work this paper cites.
“Bringing Probability Judgments into Policy Debates via Forecasting Tournaments”
Philip. Tetlock, Barbara Mellers and J Scoblic · 2017
Earlier work this paper cites.
“Attention is All You Need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“An analysis of bitcoin’s price dynamics”
Frode Kjrland et al · 2018
Earlier work this paper cites.
“OpenAI Charter”
OpenAI · 2018
Earlier work this paper cites.
“Fine-tuning Language Models from Human Preferences”
Daniel Ziegler et al · 2019
Earlier work this paper cites.
“On the Dangers of Stochastic Parrots: Can Language Models be too Big?”
Emily. Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
“Improving judgments of existential risk: Better forecasts, questions, explanations, policies”
Ezra Karger, Pavel. Atanasov and Philip Tetlock · 2022
Earlier work this paper cites.
“What do Forecasting Rationales Reveal about Thinking Patterns of Top Geopolitical Forecasters?”
Christopher Karvetski et al · 2022
Earlier work this paper cites.
“Data Contamination: From Memorization to Exploitation”
Inbal Magar and Roy Schwartz · 2022
Earlier work this paper cites.
“Chimeric Forecasting: Combining Probabilistic Predictions from Computational Models and Human Judgment”
Thomas McAndrew et al · 2022
Earlier work this paper cites.
“Early Human Judgment Forecasts of Human Monkeypox, May 2022”
Thomas McAndrew et al · 2022
Earlier work this paper cites.
“The evolution of cognitive biases in human learning”
Peter. Park · 2022
Earlier work this paper cites.
“Forecasting: Theory and Practice”
Fotios Petropoulos et al · 2022
Earlier work this paper cites.
“Emergent abilities of large language models”
Jason Wei et al · 2022
Earlier work this paper cites.
“Perils and opportunities in using large language models in psychological research”
Suhaib Abdurahman et al · 2023
Earlier work this paper cites.
“Harms of AI”
Daron Acemoğlu · 2023
Earlier work this paper cites.
“Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology”, Working Paper Series 31422, 2023
Nikhil Agarwal, Alex Moehring, Pranav Rajpurkar and Tobias Salz · 2023
Earlier work this paper cites.
“ID. 8: Co-Creating Visual Stories with Generative AI”
Victor Antony and Chien-Ming Huang · 2023
Earlier work this paper cites.
“A Theory for Emergence of Complex Skills in Language Models”
Sanjeev Arora and Anirudh Goyal · 2023
Earlier work this paper cites.
“Which humans?”
Mohammad Atari et al · 2023
Cited alongside, same era.
“Hybrid Forecasting of Geopolitical Events”
Daniel. Benjamin et al · 2023
Cited alongside, same era.
“Emergent and Predictable Memorization in Large Language Models”, 2023
Stella Biderman et al · 2023
Cited alongside, same era.
“Generative AI at Work”, Working Paper Series 31161, 2023
Erik Brynjolfsson, Danielle Li and Lindsey Raymond · 2023
Cited alongside, same era.
“Sparks of Artificial General Intelligence: Early Experiments with GPT-4”, 2023
Sébastien Bubeck et al · 2023
Cited alongside, same era.
“Quantifying Memorization Across Neural Language Models”
Nicholas Carlini et al · 2023
“ChatGPT applications in medical, dental, pharmacy, and public health education: A descriptive study highlighting the advantages and limitations”
Malik Sallam, Nesreen Salim, Muna Barakat and Alaa Al-Tammemi · 2023
Later among the works it cites.
Philipp Schoenegger and Peter. Park · 2023
Later among the works it cites.
“SlimPajama-DC: Understanding Data Combinations for LLM Training”
Zhiqiang Shen et al · 2023
Later among the works it cites.
“Utilizing Machine Learning Algorithms Trained on AI-generated Synthetic Participant Recent Music-Listening Activity in Predicting Big Five Personality Traits”
Siddharth Solaiyappan et al · 2023
Later among the works it cites.
“Three challenges for AI-assisted decision-making”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Asking Better Questions: The Art and Science of Forecasting”
Emily Dardaman and Abhishek Gupta · 2023
Cited alongside, same era.
“Super Mario Meets AI: Experimental Effects of Automation and Skills on Team Performance and Coordination”
Fabrizio Dell’Acqua, Bruce Kogut and Patryk Perkowski · 2023
Cited alongside, same era.
“Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality”
Fabrizio Dell’Acqua et al · 2023
Cited alongside, same era.
“Ideal technologies, ideal women: AI and gender imaginaries in Redditors’ discussions on the Replika bot girlfriend”
Iliana Depounti, Paula Saukko and Simone Natale · 2023
Cited alongside, same era.
“Generative artificial intelligence enhances creativity”
Anil Doshi and Oliver Hauser · 2023
Cited alongside, same era.
“Data quality in online human-subjects research: Comparisons between MTurk, Prolific, CloudResearch, Qualtrics, and SONA”
Benjamin Douglas, Patrick Ewell and Markus Brauer · 2023
Cited alongside, same era.
Mark Steyvers and Aakriti Kumar · 2023
Later among the works it cites.
“Larry Summers on who could be replaced by AI [Interviewed by Bloomberg TV’s David Westin]”, 2023
Lawrence Summers and Steve Rattner · 2023
Later among the works it cites.
“AI succession [Youtube video of talk]”
Rich Sutton · 2023
Later among the works it cites.
“Chatgpt for robotics: Design principles and model abilities”
Sai Vemprala, Rogerio Bonatti, Arthur Bucker and Ashish Kapoor · 2023
Later among the works it cites.
“Can ChatGPT Pass High School Exams on English Language Comprehension?”
Joost C.. de Winter · 2023
Later among the works it cites.
“Incentive-compatible forecasting competitions”
Jens Witkowski et al · 2023
Later among the works it cites.
“The rise and potential of large language model based agents: A survey”
Zhiheng Xi et al · 2023
Later among the works it cites.
“ExpertPrompting: Instructing Large Language Models to be Distinguished Experts”, 2023
Benfeng Xu et al · 2023
Later among the works it cites.
“Forecasting Long-Run Causal Effects”
David Bernard and Philipp Schoenegger · 2024
Closest in time.
““It would work for me too”: How Online Communities Shape Software Developers’ Trust in AI-Powered Code Generation Tools”
Ruijia Cheng, Ruotong Wang, Thomas Zimmermann and Denae Ford · 2024
Closest in time.
“Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference”, 2024
Wei-Lin Chiang et al · 2024
Closest in time.
“AI Assistance in Legal Analysis: An Empirical Study” Forthcoming
Jonathan Choi and Daniel Schwarcz · 2024
Closest in time.
“A Taxonomy for Human-LLM Interaction Modes: An Initial Exploration”
Jie Gao et al · 2024
Closest in time.
“Talk2data: A natural language interface for exploratory visual analysis via question decomposition”
Yi Guo et al · 2024
Closest in time.
“Approaching Human-Level Forecasting with Language Models”
Danny Halawi, Fred Zhang, Chen Yueh-Han and Jacob Steinhardt · 2024
Closest in time.
“The Forecasting Proficiency Test: A Practical Forecaster Evaluation Tool”
Mark Himmelstein et al · 2024
Closest in time.
“People cannot distinguish GPT-4 from a human in a Turing test”
Cameron Jones and Benjamin Bergen · 2024
Closest in time.
Thomas McAndrew et al · 2024
Closest in time.
“A Reasoning and Value Alignment Test to Assess Advanced GPT Reasoning”
Timothy McIntosh et al · 2024
Closest in time.
“AI Engine” WordPress Plugin,
Jordy Meow · 2024
Closest in time.
“Models - OpenAI API” Accessed on July 25, 2024,
OpenAI · 2024
Closest in time.
“Diminished diversity-of-thought in a standard large language model”
Peter. Park, Philipp Schoenegger and Chongyang Zhu · 2024
Closest in time.
“Is temperature the creativity parameter of large language models?”
Max Peeperkorn, Tom Kouwenhoven, Dan Brown and Anna Jordanous · 2024
Closest in time.
Philipp Schoenegger et al · 2024
Closest in time.
“Wisdom of the silicon crowd: Llm ensemble prediction capabilities match human crowd accuracy”
Philipp Schoenegger, Indre Tuminauskaite, Peter Park and Philip Tetlock · 2024
Closest in time.
“Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents” Accessed: 2024-07-24,
Vectara · 2024
Closest in time.
“Task supportive and personalized human-large language model interaction: A user study”
Ben Wang, Jiqun Liu, Jamshed Karimnazarov and Nicolas Thompson · 2024
Closest in time.
“Mmlu-pro: A more robust and challenging multi-task language understanding benchmark”
Yubo Wang et al · 2024
Closest in time.
“From Automation to Augmentation: Large Language Models Elevating Essay Scoring Landscape”
Changrong Xiao et al · 2024
Closest in time.
“Human-AI Interaction in the Age of Large Language Models”
Diyi Yang · 2024
Closest in time.