Fetching the paper…
Reading the bibliography…
We present an outcome-driven fine-tuning framework that enhances the forecasting capabilities of large language models (LLMs) without relying on human-curated reasoning samples.
Controlling the false discovery rate: a practical and powerful approach to multiple testing
Y. Benjamini and Y. Hochberg · 1995
Earlier work this paper cites.
Superforecasting: The art and science of prediction
P. E. Tetlock and D. Gardner · 2016
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, and D. Hassabis · 2017
Earlier work this paper cites.
Structured analytic techniques for intelligence analysis
R. H. Pherson and R. J. Heuer · 2019
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, and J. Steinhardt · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models, 2021
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, and A. Odena · 2021
Earlier work this paper cites.
Autocast++: Enhancing world event prediction with zero-shot ranking-based context retrieval, 2023
Q. Yan, R. Seraj, J. He, L. Meng, and T. Sylvain · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark, 2023
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, and S. R. Bowman · 2023
Earlier work this paper cites.
Nash learning from human feedback, 2023
R. Munos, M. Valko, D. Calandriello, M. G. Azar, M. Rowland, Z. D. Guo, Y. Tang, M. Geist, T. Mesnard, C. Fiegel, A. Michi, M. Selvi, S. Girgin, N. Momchev, O. Bachem, D. J. Mankowitz, D. Precup, and B. Piot · 2023
Earlier work this paper cites.
Expertprompting: Instructing large language models to be distinguished experts, 2023
B. Xu, A. Yang, J. Lin, Q. Wang, C. Zhou, Y. Zhang, and Z. Mao · 2023
Cited alongside, same era.
Forecastbench: A dynamic benchmark of ai forecasting capabilities, 2024
E. Karger, H. Bastani, C. Yueh-Han, Z. Jacobs, D. Halawi, F. Zhang, and P. E. Tetlock · 2024
Cited alongside, same era.
Financial statement analysis with large language models, 2024
A. Kim, M. Muhn, and V. Nikolaev · 2024
Cited alongside, same era.
X. Wang, M. Feng, J. Qiu, J. Gu, and J. Zhao · 2024
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Later among the works it cites.
Is dpo superior to ppo for llm alignment? a comprehensive study, 2024
S. Xu, W. Fu, J. Gao, W. Ye, W. Liu, Z. Mei, and Y. Wu · 2024
Later among the works it cites.
M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, and Y. Zhang · 2024
Later among the works it cites.
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, and I. Kivlichan · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Cao, J. Zhuang, and Q. He · 2024
Cited alongside, same era.
Wisdom of the silicon crowd: Llm ensemble prediction capabilities rival human crowd accuracy, 2024
P. Schoenegger, I. Tuminauskaite, P. S. Park, and P. E. Tetlock · 2024
Cited alongside, same era.
Approaching human-level forecasting with language models, 2024
D. Halawi, F. Zhang, C. Yueh-Han, and J. Steinhardt · 2024
Cited alongside, same era.
Calibrating large language models with sample consistency, 2024
Q. Lyu, K. Shridhar, C. Malaviya, L. Zhang, Y. Elazar, N. Tandon, and C. Callison-Burch · 2024
Cited alongside, same era.
Self-play fine-tuning converts weak language models to strong language models, 2024
Z. Chen, Y. Deng, H. Yuan, K. Ji, and Q. Gu · 2024
Cited alongside, same era.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, and Z. Qiu · 2024
Later among the works it cites.
Exploring the impact of quantization on llm performance
O. Zem · 2024
Later among the works it cites.
An empirical study of llama3 quantization: From llms to mllms
W. Huang, X. Zheng, X. Ma, H. Qin, C. Lv, H. Chen, and M. Magno · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, and Y. He · 2025
Closest in time.