Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have recently been applied to forecasting tasks, with some works claiming these systems match or exceed human performance.
Capital asset prices: A theory of market equilibrium under conditions of risk
William F Sharpe · 1964
Earlier work this paper cites.
Elicitation of personal probabilities and expectations
Leonard J Savage · 1971
Earlier work this paper cites.
Survivorship bias and mutual fund performance
Elton M Gruber and C Blake · 1996
Earlier work this paper cites.
Risk-adjusted performance of mutual funds
Katerina Simons · 1998
Earlier work this paper cites.
The probability of backtest overfitting
David H Bailey, Jonathan Borwein, Marcos Lopez de Prado, and Qiji Jim Zhu · 2015
Earlier work this paper cites.
A backtesting protocol in the era of machine learning
Robert D Arnott, Campbell R Harvey, and Harry Markowitz · 2018
Earlier work this paper cites.
Long-horizon predictability: a cautionary tale
Jacob Boudoukh, Ronen Israel, and Matthew Richardson · 2019
Earlier work this paper cites.
Alignment problems with current forecasting platforms
Nuño Sempere and Alex Lawsen · 2021
Earlier work this paper cites.
Forecasting future world events with neural networks
Andy Zou, Tristan Xiao, Ryan Jia, Joe Kwon, Mantas Mazeika, Richard Li, Dawn Song, Jacob Steinhardt, Owain Evans, and Dan Hendrycks · 2022
Earlier work this paper cites.
White House planning face-to-face meeting with Biden, Xi, 2023
Reuters · 2023
Earlier work this paper cites.
Biden, Xi talks in san francisco, 2023
Japan Times · 2023
Earlier work this paper cites.
Contra papers claiming superhuman AI forecasting, 2024
Nikos Bosse, Peter Mühlbacher, Lawrence Phillips, and Dan Schwarz · 2024
Cited alongside, same era.
Knowledge cutoff issues of GPT-4o regarding Phan et al. (2024), 2024
Danny Halawi · 2024
Cited alongside, same era.
Approaching human-level forecasting with language models, 2024
Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt · 2024
Cited alongside, same era.
Reasoning and tools for human-level forecasting
Elvis Hsieh, Preston Fu, and Jonathan Chen · 2024
Cited alongside, same era.
Questionable practices in machine learning
Gavin Leech, Juan J Vazquez, Niclas Kupper, Misha Yagudin, and Laurence Aitchison · 2024
Cited alongside, same era.
Point-in-time vs. lagged fundamentals: This time i(t)’s different?
Ernest Breitschwerdt · 2025
Closest in time.
Are LLMs prescient? A continuous evaluation using daily news as the oracle
Hui Dai, Ryan Teehan, and Mengye Ren · 2025
Closest in time.
Polymarket settles a market incorrectly – again, 2024
Chris Gerlacher · 2025
Closest in time.
Introducing the SalemCSPi forecasting tournament, 2022
Richard Hanania · 2025
Closest in time.
The emerging science of machine learning benchmarks
Moritz Hardt · 2025
Closest in time.
asgeirtj/system_prompts_leaks/claude.txt, 2025
Asgeir Thor Johnson · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Paleka, Abhimanyu Pallavi Sudhir, Alejandro Alvarez, Vineeth Bhat, Adam Shen, Evan Wang, and Florian Tramèr · 2024
Cited alongside, same era.
LLMs are superhuman forecasters, 2024
Long Phan, Adam Khoja, Mantas Mazeika, and Dan Hendrycks · 2024
Cited alongside, same era.
Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy
Philipp Schoenegger, Indre Tuminauskaite, Peter S Park, Rafael Valdece Sousa Bastos, and Philip E Tetlock · 2024
Cited alongside, same era.
Continual learning for large language models: A survey, 2024
Tongtong Wu, Linhao Luo, Yuan-Fang Li, Shirui Pan, Thuy-Trang Vu, and Gholamreza Haffari · 2024
Cited alongside, same era.
Who predicted 2022?, 2023
Scott Alexander · 2025
Cited alongside, same era.
AI forecasting bots incoming: comment section, 2024
Gwern Branwen · 2025
Cited alongside, same era.
ForecastBench: A dynamic benchmark of AI forecasting capabilities, 2024a
Ezra Karger, Houtan Bastani, Chen Yueh-Han, Zachary Jacobs, Danny Halawi, Fred Zhang, and Philip E. Tetlock
Cited in the paper.
Acx2025 tournament, 2025
Metaculus · 2025
Closest in time.
Against calibration, 2023
Eigil Fjeldgren Rischel · 2025
Closest in time.
PROPHET: An inferable future forecasting benchmark with causal intervened likelihood estimation
Zhengwei Tao, Zhi Jin, Bincheng Li, Xiaoying Bai, Haiyan Zhao, Chengfeng Dou, Xiancai Chen, Jia Li, Linyu Li, and Chongyang Tao · 2025
Closest in time.
Haooowang/llm-knowledge-cutoff-dates, 2025
Hao Wang · 2025
Closest in time.
Out-of-sample tests of forecasting accuracy: An analysis and review
Leonard J. Tashman · 2070
Closest in time.