Fetching the paper…
Reading the bibliography…
Forecasting is a task that is difficult to evaluate: the ground truth can only be known in the future.
Utility Theory for Decision Making
Peter C. Fishburn · 1970
Earlier work this paper cites.
Elicitation of Personal Probabilities and Expectations
Leonard J. Savage · 1971
Earlier work this paper cites.
Metamorphic testing: a new approach for generating next test cases
Tsong Y Chen, Shing C Cheung, and Shiu Ming Yiu · 1998
Earlier work this paper cites.
Economic preferences or attitude expressions?: an analysis of dollar responses to public issues
Daniel Kahneman, Ilana Ritov, David Schkade, Steven J Sherman, and Hal R Varian · 2000
Earlier work this paper cites.
Logarithmic Market Scoring Rules for Modular Combinatorial Information Aggregation
Robin Hanson · 2002
Earlier work this paper cites.
The Promise of Prediction Markets
Kenneth J. Arrow, Robert Forsythe, Michael Gorham, Robert Hahn, Robin Hanson, John O. Ledyard, Saul Levmore, Robert Litan, Paul Milgrom, Forrest D. Nelson, George R. Neumann, Marco Ottaviani, Thomas C. Schelling, Robert J. Shiller, Vernon L. Smith, Erik Snowberg, Cass R. Sunstein, Paul C. Tetlock, Philip E. Tetlock, Hal R. Varian, Justin Wolfers, and Eric Zitzewitz · 2008
Earlier work this paper cites.
Hanson’s automated market maker
Henry Berg and Todd A Proebsting · 2009
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2019
Earlier work this paper cites.
A logic-driven framework for consistency of neural models
Tao Li, Vivek Gupta, Maitrey Mehta, and Vivek Srikumar · 2019
Earlier work this paper cites.
How Feasible Is Long-range Forecasting?, October 2019
Muehlhauser, Luke · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Testing monotonicity of machine learning models, 2020
Arnab Sharma and Heike Wehrheim · 2020
Cited alongside, same era.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg · 2021
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2021
Cited alongside, same era.
Discovering Latent Knowledge in Language Models Without Supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Cited alongside, same era.
Specifying and testing k k -safety properties for machine-learning models
Maria Christakis, Hasan Ferit Eniser, Jörg Hoffmann, Adish Singla, and Valentin Wüstholz · 2022
Cited alongside, same era.
Large Language Model Prediction Capabilities: Evidence from a Real-World Forecasting Tournament, October 2023
Philipp Schoenegger and Peter S. Park · 2023
Later among the works it cites.
Autocast++: Enhancing world event prediction with zero-shot ranking-based context retrieval
Qi Yan, Raihan Seraj, Jiawei He, Lili Meng, and Tristan Sylvain · 2023
Later among the works it cites.
Evaluating superhuman models with consistency checks
Lukas Fluri, Daniel Paleka, and Florian Tramèr · 2024
Closest in time.
FUTURESEARCH: Manifold markets trading bot, 2024
FutureSearch · 2024
Closest in time.
Scorable Functions: A Format for Algorithmic Forecasting, May 2024
Ozzie Gooen · 2024
Closest in time.
Approaching Human-Level Forecasting with Language Models, February 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li-Cheng Lan, Huan Zhang, Ti-Rong Wu, Meng-Yu Tsai, I Wu, Cho-Jui Hsieh, et al · 2022
Cited alongside, same era.
Dutch Book Arguments
Susan Vineberg · 2022
Cited alongside, same era.
Consistency analysis of ChatGPT
Myeongjun Jang and Thomas Lukasiewicz · 2023
Cited alongside, same era.
Searching for a model’s concepts by their shape – a theoretical framework, February 2023
Kaarel, gekaklam, Walter Laurito, Kay Kozaronek, AlexMennen, and June Ku · 2023
Cited alongside, same era.
Goodhart’s Law in Reinforcement Learning
Jacek Karwowski, Oliver Hayman, Xingjian Bai, Klaus Kiendlhofer, Charlie Griffin, and Joar Max Viktor Skalse · 2023
Cited alongside, same era.
Benchmarking and improving generator-validator consistency of language models
Xiang Lisa Li, Vaishnavi Shrivastava, Siyan Li, Tatsunori Hashimoto, and Percy Liang · 2023
Cited alongside, same era.
Semantic consistency for assuring reliability of large language models
Harsh Raj, Vipul Gupta, Domenic Rosati, and Subhabrata Majumdar · 2023
Cited alongside, same era.
Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt · 2024
Closest in time.
Reasoning and tools for human-level forecasting
Elvis Hsieh, Preston Fu, and Jonathan Chen · 2024
Closest in time.
Forecastbench: A dynamic benchmark of ai forecasting capabilities
Ezra Karger, Houtan Bastani, Chen Yueh-Han, Zachary Jacobs, Danny Halawi, Fred Zhang, and Philip E Tetlock · 2024
Closest in time.
Instructor: Structured LLM Outputs, May 2024
Jason Liu · 2024
Closest in time.
LLMs are superhuman forecasters, 2024
Long Phan, Adam Khoja, Mantas Mazeika, and Dan Hendrycks · 2024
Closest in time.
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy, May 2024
Philipp Schoenegger, Indre Tuminauskaite, Peter S. Park, and Philip E. Tetlock · 2024
Closest in time.
Long-range subjective-probability forecasts of slow-motion variables in world politics: Exploring limits on expert judgment
Philip E Tetlock, Christopher Karvetski, Ville A Satopää, and Kevin Chen · 2024
Closest in time.